Skip to content
-
Subscribe to our newsletter & never miss our best posts. Subscribe Now!
PHDPedia PHDPedia PHDPedia
PHDPedia PHDPedia PHDPedia
  • Home
  • Sitemap
  • Home
  • Sitemap
Close

Search

  • https://www.facebook.com/
  • https://twitter.com/
  • https://t.me/
  • https://www.instagram.com/
  • https://youtube.com/
Subscribe
Academic Publishing & Open Access

TRACE: Trusted Retrieval & Attribution for Content Ecosystems Launched to Navigate AI’s Impact on Scholarly Communication

By Ammar Sabilarrohman
October 9, 2026 14 Min Read
Comments Off on TRACE: Trusted Retrieval & Attribution for Content Ecosystems Launched to Navigate AI’s Impact on Scholarly Communication

As the scholarly publishing landscape grapples with a transformative shift driven by artificial intelligence, a new collaborative initiative, TRACE: Trusted Retrieval & Attribution for Content Ecosystems, has been launched to build essential infrastructure for AI-mediated research communication. This pivotal announcement, made at the STM Conference preceding the Frankfurt Book Fair, marks a concerted effort by key industry players to ensure the integrity, trustworthiness, and sustainability of scholarly content in the age of generative AI.

The New Digital Frontier in Scholarly Communication

The journey towards digital scholarly communication began in earnest in the 1990s and early 2000s, fundamentally altering how research was disseminated and accessed. This period was characterized by a rapid evolution from print-based distribution to electronic access, necessitating the creation of entirely new infrastructures and standards. Key community initiatives during this era laid the groundwork for the digital ecosystem we now rely upon. The Digital Object Identifier (DOI) pilot project, initiated in the mid-1990s, swiftly led to the formation of Crossref, an organization that has become indispensable for persistent identification and linking of scholarly content. By providing unique, permanent identifiers for online articles, books, and other research outputs, DOIs revolutionized citation practices and facilitated robust interlinking across diverse platforms. This innovation was crucial in establishing a stable digital reference system where none had previously existed, effectively solving the problem of "link rot" and ensuring scholarly persistence.

Concurrently, the Digital Library Federation’s E-Resource Management Initiative (ERMI) catalyzed significant advancements in library systems. ERMI’s efforts spurred the development of Electronic Resource Management (ERM) systems, indexed discovery systems, and critical data exchange standards like KBART (Knowledge Bases And Related Tools) for managing library holdings, and COUNTER (Counting Online Usage of Networked Electronic Resources) for standardizing usage data. These initiatives brought much-needed order and efficiency to the acquisition, management, and assessment of digital resources within academic libraries, enabling institutions to navigate the complexities of electronic licensing and access. The early 2000s also witnessed the burgeoning momentum of the Open Access movement, notably galvanized by the meeting organized by the Open Society in Budapest in 2001. This landmark event, culminating in the Budapest Open Access Initiative (BOAI), articulated a vision for free, immediate online access to scholarly literature, challenging traditional subscription models and advocating for the democratization of knowledge. While the transition from print to electronic publishing was a complex, multi-decade endeavor, considerable progress was achieved, establishing a resilient framework for digital scholarly exchange.

Today, the scholarly communication ecosystem finds itself at a similarly profound inflection point, driven by the rapid proliferation and sophistication of artificial intelligence tools. Just as electronic content distribution upended the established structures of print-based content sharing, the pervasive role of AI systems demands a comprehensive reconsideration of virtually every aspect of research communication. This includes fundamental concepts such as content access, citation practices, usage tracking, and the financial models that underpin the entire scholarly marketplace. The emergence of "agentic access" – where autonomous AI agents discover and interact with content – alongside widespread machine consumption and the AI-driven generation of new content, necessitates a swift and strategic response. While not every existing system will require redevelopment, a multitude of new standards, protocols, and technical infrastructures must be rapidly conceived and implemented. The pace of adaptation required for this AI-driven transformation is anticipated to be significantly faster than the three decades it took to fully transition from print to electronic. The impacts are expected to be equally profound, even as their full scope remains somewhat indistinct.

TRACE Initiative: A Collaborative Response to AI’s Challenge

Recognizing the urgent need for coordinated action, the scholarly community has once again convened to accelerate progress on AI systems and their integration into research workflows. This week, at the STM Conference held in advance of the Frankfurt Book Fair, a new community initiative, TRACE: Trusted Retrieval & Attribution for Content Ecosystems, was officially launched. The announcement featured a distinguished panel comprising Priya Madina from Springer Nature, Tasha Mellins-Cohen from COUNTER, Todd Toler from Ithaka S+R, and Todd Carpenter from NISO. The timing and venue of this announcement—a major international gathering of publishers, technology providers, and academic leaders—underscored the gravity and widespread relevance of the initiative.

Priya Madina, representing Springer Nature, emphasized the industry’s commitment to fostering a responsible AI ecosystem, stating, "The rapid evolution of AI presents both immense opportunities and significant challenges for research integrity. TRACE is crucial for ensuring that as we embrace these new technologies, we do so in a way that upholds the fundamental principles of attribution, trust, and verifiability that underpin scholarly publishing." Tasha Mellins-Cohen of COUNTER highlighted the necessity of adapting measurement standards, noting, "Understanding how AI interacts with content is paramount for sustainable publishing models and for researchers to demonstrate impact. Our work within TRACE will provide the essential data points needed to navigate this new landscape." Todd Toler from Ithaka S+R added, "Our research will provide foundational insights into how AI agents operate within the scholarly content environment, informing the development of practical, implementable solutions that benefit all stakeholders." Todd Carpenter, representing NISO, further articulated the initiative’s broad scope: "TRACE is designed to be an inclusive, collaborative forum, bringing together diverse voices from across the scholarly and AI communities. Its mission is to accelerate the development of the technical infrastructure required to enhance the trustworthiness of AI systems in scholarly communication."

TRACE aims to achieve this by actively supporting the development and adoption of robust content exchange standards, sophisticated provenance tracking mechanisms, and harmonized usage reporting protocols. Beyond directly funding and incubating new work, TRACE also serves as a crucial coordinating forum. This function is vital for aligning and amplifying complementary efforts already underway across the fragmented scholarly communications ecosystem, preventing duplication of effort and fostering synergistic progress. The initiative has commenced with initial funding allocated to three parallel projects, each addressing critical challenges in AI-mediated scholarly communication. These projects focus on how AI agents discover and access content, how provenance and attribution are preserved throughout the content lifecycle, and how AI usage can be measured and reported consistently across various platforms and applications. Collectively, these infrastructure components are designed to ensure that authors and content creators receive appropriate recognition for their work, irrespective of specific publisher business models. Furthermore, they seek to reduce friction and computational overhead in content discovery and access for AI systems. By establishing a coherent and comprehensive framework for trusted retrieval, accurate attribution, and reliable measurement, TRACE intends to empower AI developers to build more dependable products and services. In turn, this enhanced reliability will strengthen confidence in AI-assisted research workflows, fostering greater trust among researchers, publishers, technology providers, and the wider scholarly communications community.

Pioneering Projects Under the TRACE Banner

The TRACE initiative is immediately supporting three critical projects, each spearheaded by a leading organization in scholarly communication infrastructure, demonstrating a multi-faceted approach to the AI challenge.

Standardizing AI Usage Metrics: The COUNTER Approach
COUNTER, the authoritative body for standardizing usage statistics, began proactive work on AI usage metrics in 2025. This foresight recognized that traditional metrics, designed for human consumption of content on publisher platforms, would be insufficient for capturing the complex interactions of AI systems. Phase one of this effort, concluded in early 2026, established initial guidance, introducing new concepts such as an "Access Method" specific to AI, an "Agent" category for AI entities, and new AI-specific metrics. This early work aimed to align AI usage metrics with existing traditional metrics, enabling a clear demonstration of evolving usage patterns and their budget implications for subscribing institutions. However, feedback on phase one highlighted a concern that the initial approach was too technology-dependent. Consequently, phase two, which commenced in the spring of 2026 with funding from TRACE, is adopting a more technology-agnostic methodology. Critically, phase two will incorporate reporting about "inference-time usage" – tracking content when an AI system retrieves, grounds, or cites publisher material in response to user prompts, even when these interactions occur outside of publishers’ own platforms. This is a significant challenge, as it requires tracking usage across a decentralized and often opaque landscape.

Participants in the TRACE initiative were deeply engaged in phase one of COUNTER’s work and continue to contribute to phase two. TRACE funding will specifically support the continued development of the COUNTER API and JSON Schema to seamlessly integrate the new AI metrics advanced earlier this year. Furthermore, COUNTER is actively exploring usage models developed by the SPUR Coalition, a parallel initiative focused on developing broader AI usage metrics. This collaboration ensures alignment and mutual benefit, with plans to reference SPUR telemetry in COUNTER’s phase two guidelines. This will involve mapping citation and presentation telemetry events to existing AI COUNTER metrics and utilizing retrieval telemetry events for a new, holistic "AI retrievals" metric, providing a more comprehensive view of AI interaction with scholarly content.

Forging Trust Through Provenance: NISO’s Pilot Project
Like COUNTER, NISO (National Information Standards Organization) has been closely monitoring AI developments, exploring potential standards work, and providing educational resources on AI tools well before the formalization of TRACE. NISO’s journey towards provenance tracking gained significant traction following a prioritization discussion hosted in the spring of 2025. These discussions highlighted the urgent need for mechanisms to verify the origin and evolution of content within AI systems. The topic was further explored during the NISO Plus Conference in February 2026, with sessions dedicated to AI applications, usage, and discussions on "How to keep the robots in line." The extensive ideas generated at the conference led to a subsequent workshop in Cambridge in May 2026. This workshop, detailed in a previous report, was instrumental in shaping the community’s approach. A concrete outcome of the Cambridge Workshop was the decision to launch work on a minimum metadata package capable of describing content provenance through the inference process. This critical project is now being directly supported by the TRACE initiative. NISO’s work aims to address the fundamental problem of AI systems generating fluent but often untraceable or even fabricated information, thereby eroding trust in AI-assisted research.

TRACE Project Launches to Advance AI interoperability

Unpacking AI Agent Interactions: The Ithaka S+R and STM Study
Complementing the efforts of COUNTER and NISO, STM, the international association of scientific, technical, and medical publishers, has also announced a third foundational component of the TRACE project. This initiative involves a focused, three-month study led by Ithaka S+R, a research and consulting organization specializing in higher education and libraries. The study will meticulously investigate how AI agents discover scholarly content, how they authenticate through a user’s entitlements, how they receive provenance-bearing information, and how they subsequently create a credible record of content use. This research is designed to bring together a diverse array of stakeholders, including publishers, open repositories, critical infrastructure providers, and AI developers. By engaging these varied perspectives, the study will explore where current approaches are converging, identify persistent challenges, and delineate the requirements for future implementation to support trusted scholarly communications. The findings of this research will be crucial in informing the development of practical, interoperable solutions for managing AI access and attribution.

Deep Dive: NISO’s Provenance Pilot and the Quest for Verifiability

The core challenge NISO’s provenance pilot addresses is the well-documented propensity of large language models (LLMs) to generate confident, fluent text that frequently includes fabricated or misattributed citations. Research has consistently shown hallucinated references at rates ranging from approximately 20% to 40% in general LLM outputs. Even when genuine sources are cited, many AI systems currently lack the capability to transparently demonstrate which specific training data, retrieved document, or detailed model parameter contributed to a particular claim. This inherent lack of transparency creates significant legal, ethical, and factual gaps, directly threatening trust in AI-assisted research, publishing, and even enterprise decision-making. The fundamental question remains: how can we trust a research outcome if its derivation cannot be verifiably traced back to its original sources?

Addressing the Hallucination Crisis in AI
Researchers and other users of scholarly content are increasingly leveraging generative and agentic AI systems to efficiently retrieve, summarize, compare, and synthesize authoritative content. The primary risk is not merely the production of fabricated or inaccurate citations, but the profound absence of a verifiable pathway from an AI-generated output back to the precise source object, component, passage, figure, dataset, or standard that supports the claim. This critical lack of connectivity between AI-driven research generation and its source material poses multifaceted challenges. It undermines verifiability, complicates proper attribution, hinders recognition for content creators, obfuscates usage assessment, complicates version control, and impedes awareness of retractions. These challenges collectively erode trust, not only in the generated outputs themselves but also in the integrity and trustworthiness of the entire scholarly record. A community-wide effort is therefore imperative to address these systemic issues comprehensively.

Defining Inference-Time Provenance: Scope and Stakeholders
In response to these interconnected issues, NISO, COUNTER, and Cambridge University Press co-hosted a pivotal workshop in May 2026. Technical experts from publishing, the AI tool developer community, and systems librarianship convened in Cambridge to collaboratively scope potential community efforts. Building on the insights gained from this workshop and the ongoing work by COUNTER to enhance usage metrics for AI systems, NISO is launching a focused pilot project. This pilot aims to discern a minimum set of metadata sufficient to describe the source of a piece of content used specifically in the inference process of AI systems. The pilot’s focus is strategically placed on the layer where the research publishing and repository communities possess practical agency. This includes publisher/repository-controlled source delivery mechanisms, repository and content APIs, retrieval pipelines, agentic tool calls, metadata packages, signed content objects, and verification services. Crucially, the pilot explicitly will not attempt to solve the full spectrum of problems around AI use of content, such as deterministic source attribution within foundation-model weights, complex enforcement issues related to copyright, or nuanced access-control management. These broader issues, while critical, are beyond the scope of this initial, focused effort.

The NISO pilot is slated for a twelve-month duration, though an expedited timeline is hoped for. Its core objective is to define, test, evaluate, compare, and refine a minimum viable metadata model specifically for provenance and attribution tracking in AI systems utilized for research and business R&D applications. The pilot’s concentration on inference-time provenance is key: this refers to the information that should accompany scholarly and professional source materials when they are retrieved, synthesized, cited, or transformed by generative and agentic AI systems after the initial training of the underlying LLMs has occurred. The pilot will convene members from five critical stakeholder categories: content creators, publishers, repositories, AI tool developers, and end-users. This diverse representation within a coordinated framework is designed to produce actionable, evidence-based guidance for future standards development. While specific technological approaches will be determined by willing pilot participants, the overarching goal is to study how provenance functions and is supported in various AI architectures, including Retrieval-Augmented Generation (RAG) with citation grounding, inference functions/training data attribution (TDA), output watermarking, statistical provenance testing, and knowledge graph documentation. The aim is to ensure that provenance data can be consistently preserved and communicated across all these diverse system designs.

Bridging Gaps in Existing Provenance Standards
The rationale for this pilot stems from the recognition that while existing provenance standards provide important foundations, they do not yet fully specify the attribution context required for AI systems that retrieve and synthesize scholarly and professional content. Standards like W3C PROV offer robust frameworks for describing entities, activities, and agents; C2PA (Coalition for Content Provenance and Authenticity) supports signed content credentials and manifests; DOI and related Persistent Identifier (PID) systems effectively identify scholarly objects; and usage reporting frameworks like COUNTER can describe measurable access events. However, the critical missing layer is a practical, interoperable descriptive profile that explicitly connects these rich resources to AI-generated outputs at the precise point of retrieval, synthesis, and display, through their source metadata. This new metadata set must be sufficiently comprehensive to describe any inference content and adaptable enough to function across a wide range of AI system architectures. The ultimate goal of this NISO initiative is to rigorously test various approaches through an iterative process, thereby discerning the absolute minimum amount of data required to accurately describe an object’s source within an AI-mediated workflow.

It is crucial to distinguish between the metadata payload itself and the approach a service provider (e.g., publisher or repository) may deploy to transmit or verify this provenance metadata, including the specific content-delivery endpoints and communication protocols. The latter constitutes a separate, though equally vital, project. This next phase of work, which will involve specifying, developing, and testing the mechanisms to carry the provenance data, could adapt the results of the research work led by Ithaka S+R. This could then lead to the launch of subsequent work by TRACE in coordination with NISO or other relevant technical groups, depending on its ultimate scope.

A Unified Front: Engaging the Broader Ecosystem

The various elements of TRACE, while distinct projects, are deeply interconnected and collectively represent a broader, ecosystem-wide effort. The three current TRACE projects – Ithaka S+R’s research, NISO’s provenance metadata pilot, and COUNTER’s usage metrics initiatives – are not operating in isolation. They actively involve members of the TRACE community but also extend their reach to incorporate a wider array of community voices. Ithaka S+R’s research, for instance, will be conducted independently, drawing on the extensive experience of publishers, open repositories, platform providers, and AI developers. Its findings are committed to being published openly, ensuring transparency and widespread benefit. COUNTER’s work includes not only publishers and AI developers but also representatives from the library community, whose perspectives on budget implications and access are invaluable. NISO’s provenance initiative similarly involves TRACE-participating publishers, alongside AI tool developers, content repositories, and other content businesses. For example, NISO is actively engaging with Creative Commons to explore how a provenance metadata package could be integrated into the CC Signals effort, potentially facilitating wide adoption across diverse media types and content licenses.

This broad, inclusive view is essential because the challenges posed by AI extend far beyond the confines of research publishing. Organizations across the entire content world are grappling with similar issues. Open repositories, for instance, are increasingly needing to manage their traffic to prevent outages caused by AI bots. Authors and creators across every media type, from scientific papers to artistic works, are deeply concerned about provenance and ensuring appropriate recognition for their intellectual contributions. Usage tracking, beyond providing simple cost-per-use metrics, serves as a crucial signal to funders, administrators, and stakeholders, regardless of whether a site operates on a subscription or open access model. These same motivations that drive our community’s interest in usage metrics are also driving similar work within advertising-driven media and other commercial sectors.

There is significant strategic value in aligning the work of the research community with broader commercial and web interests. This ensures that scholarly communication remains engaged with related conversations in the wider tech landscape and maintains relevance to the companies at the forefront of AI development. The era when the research community was sufficiently large and self-contained to build tools and approaches that suited only its own ecosystem, distinctly different from others, is long past. We must understand the inherent value that building for our community, centered around our core interests in verifiability, attribution, and citation, holds outside of our immediate sphere. By leveraging these shared needs, we can collectively push for adoption and interoperability in ways that might not be achievable if we acted alone.

Looking Ahead: Shaping the Future of Research Integrity in the AI Age

The launch of TRACE represents a proactive and collaborative step towards building a robust and trustworthy future for scholarly communication in the age of AI. By addressing critical infrastructure gaps in usage measurement, provenance tracking, and content access for AI agents, the initiative seeks to safeguard research integrity, ensure fair attribution for creators, and provide reliable tools for researchers. The parallel projects by COUNTER, NISO, and Ithaka S+R/STM, all operating under the TRACE umbrella, illustrate a comprehensive approach to tackling the multifaceted challenges presented by generative AI. As these projects evolve, their findings and standards will be critical in shaping how AI interacts with the vast body of human knowledge, influencing everything from academic publishing models to the very nature of scientific discovery. The emphasis on broad community engagement and alignment with wider industry trends underscores a recognition that this transformation requires a unified, globally coordinated effort to ensure that AI serves to enhance, rather than undermine, the pursuit and dissemination of knowledge. The success of TRACE will be measured not just in technical specifications, but in its ability to foster sustained trust in the scholarly record amidst an unprecedented technological revolution.

Tags:

Academic PublishingattributioncommunicationcontentecosystemsimpactJournalslaunchednavigateOpen AccessPeer Reviewretrievalscholarlytracetrusted
Author

Ammar Sabilarrohman

Follow Me
Other Articles
Previous

Francis Halzen Awarded 2026 Nobel Prize in Physics for Pioneering Neutrino Astronomy

Next

The State of Julia Development March 2026 Monthly Internal Update and Ecosystem Progress Report

Recent Posts

Docker Agent: Revolutionizing AI Agent Development with Containerization PrinciplesNSF’s Astronomical Sciences Core Research Program Unveils Funding Priorities for Foundational Astronomy and AstrophysicsDeWalt 20V MAX Cordless Drill and Impact Driver Combo Set Available at Significant Discount During October Prime Day EventThe Ethics of the Mind Navigating the Integration of Artificial Intelligence into Human Neural Implants
Docker Agent: Revolutionizing AI Agent Development with Containerization PrinciplesNSF’s Astronomical Sciences Core Research Program Unveils Funding Priorities for Foundational Astronomy and AstrophysicsDeWalt 20V MAX Cordless Drill and Impact Driver Combo Set Available at Significant Discount During October Prime Day EventThe Ethics of the Mind Navigating the Integration of Artificial Intelligence into Human Neural Implants
  • Docker Agent: Revolutionizing AI Agent Development with Containerization Principles
  • NSF’s Astronomical Sciences Core Research Program Unveils Funding Priorities for Foundational Astronomy and Astrophysics
  • DeWalt 20V MAX Cordless Drill and Impact Driver Combo Set Available at Significant Discount During October Prime Day Event
  • The Ethics of the Mind Navigating the Integration of Artificial Intelligence into Human Neural Implants
  • The Evolution of Knitr Spin Enhanced Multi-Language Support and Structural Integrity in the 2026 Backlog Sprint

Archives

  • October 2026
  • September 2026
  • August 2026
  • July 2026
  • May 2026
  • April 2026

Categories

  • Academic Productivity & Tools
  • Academic Publishing & Open Access
  • Data Science & Statistics for Researchers
  • Funding, Grants & Fellowships
  • Higher Education News
  • Humanities & Social Sciences Research
  • Pedagogy & Teaching in Higher Ed
  • PhD Life & Mental Health
  • Post-PhD Careers & Alt-Ac
  • Research Methods & Methodology
  • Science Communication (SciComm)
  • Thesis & Academic Writing
Copyright 2026 — PHDPedia. All rights reserved. Blogsy WordPress Theme