
Training Data Under Fire: Why AI Transparency and Copyright Risks Are Reaching a Breaking Point

Introduction: What We Cannot See May Matter Most
When we interact with generative artificial intelligence (GenAI), our attention naturally turns to what appears on the screen. We ask whether an AI-generated image resembles an existing artwork, whether generated text reproduces protected expression, or whether an AI-created output can attract copyright protection. Perhaps we should also be asking what happened before the output appeared.
Every generative AI system depends upon training. Large models learn patterns from enormous quantities of text, images, audio, code, and other material. Some of this material is openly available. Some may be licensed. Some may be protected by copyright. And, in many cases, those whose works contributed to the informational environment from which AI models learn may have little practical ability to determine whether, when, or how their works were used. Outputs of GenAI tools can change when similar prompts are submitted, posing reproducibility issues. Incorrect citations or citations that cannot be verified may pose challenges tracing where information came from. Improved transparency around provenance is important not just for rights holders but also for researchers looking to evaluate the trustworthiness and reproducibility of AI-generated information.
This creates a fundamental problem for copyright law: how can rights be meaningfully exercised when the relevant use is difficult to see? I would argue that training data transparency should therefore be understood not simply as a technical feature of responsible AI, but increasingly as part of the legal infrastructure required to make copyright accountability possible.

The Copyright Question Begins Before the Prompt
Copyright discussions surrounding GenAI have often concentrated on outputs and authorship. But training itself may involve legally significant acts. Building datasets and preparing material for machine learning can involve copying copyright-protected works. Whether those activities require permission depends on the applicable copyright regime, the nature of the use, and available exceptions or limitations. Permission can also be granted through licensing agreements. AI creators could license copyright-protected works directly from rightsholders or they could enter into contracts with publishers and other providers of content who own rights in relevant works. These examples show that AI training does not have to rely on infringing, unlicensed uses: agreements can license protected content for use in training sets.
An openly accessible scholarly article remains a copyright work unless copyright has expired, been waived, or some other legal rule applies. Open licenses can grant extensive reuse permissions, but they operate within copyright law rather than outside it. Creative Commons itself has emphasized that applying CC licenses to AI training is legally complex because copyright exceptions and limitations vary between jurisdictions. Open access therefore cannot resolve the AI training problem simply by making more content available. Accessibility and legal permission are related concepts, but they are not identical.
Transparency as a Condition of Accountability
Suppose an author believes that their works were incorporated into the training data of a commercial generative AI model. Before considering infringement, licensing, exceptions, or remedies, there is a more elementary difficulty: how would the author know?
Developers understandably have legitimate interests in protecting commercially sensitive information, proprietary methods, and system security. But complete opacity carries its own costs. Rights holders cannot meaningfully assess potential uses of their works. Regulators face difficulties evaluating compliance. Researchers cannot easily scrutinise the provenance and composition of datasets. Particularly, researchers who want to study GenAI models themselves may also face constraints when trying to understand where their training data came from and what it consists of.
Meaningful transparency should also apply, where technically possible, to provenance at time of use. Information that has been generated by AI and presented in research contexts should include verifiable citations or source information that allows users to trace claims back to their source. Absence of that ability reveals that transparency at the training stage becomes untethered from accountability at the output stage. This is something that Europe has begun implementing.

Europe Moves from Principle to Obligation
The EU Artificial Intelligence Act represents an important shift because it connects AI transparency directly with copyright compliance. Providers of general-purpose AI models are required to maintain a policy designed to comply with EU copyright law and to make publicly available a sufficiently detailed summary of the content used to train their models. These obligations began applying to relevant providers from August 2, 2025.
The European Commission has subsequently developed a mandatory template for these public summaries. Importantly, the framework does not demand publication of every individual training item. Instead, it seeks structured information concerning categories and sources of training content, including publicly available datasets, private datasets, material scraped from online sources, user data, and synthetic data. Transparency does not determine whether copyright infringement has occurred. Nor does a training-data summary automatically establish that every use was lawful. It does something more basic: it creates conditions in which those legal questions can actually be asked. That distinction is important. Transparency should not become a substitute for substantive copyright analysis. It is better understood as an enabling mechanism for it. This also brings up the question of how other countries are handling these processes.
Australia: A Different Stage of the Conversation
Australia has not adopted an equivalent training-content disclosure regime. Instead, copyright and AI questions continue to develop through policy consultation and stakeholder engagement, particularly through the Australian Government’s Copyright and Artificial Intelligence Reference Group (CAIRG). Significantly, the Group’s initial work focused directly on the use of copyright material as inputs for AI systems and copyright-related AI transparency.
Current policy priorities include examining fair and legal avenues for the use of copyright material in AI through licensing arrangements, improving certainty about the application of copyright law to AI-generated material, and exploring more accessible enforcement mechanisms. The Government has also indicated that it is not considering the introduction of a text and data mining exception into Australian copyright law.
Before debating whether particular uses of copyright material for AI training should be permitted, licensed, restricted, or otherwise regulated, policymakers require sufficient information about how that material enters the training process in the first place. The Australian debate should therefore ask not only: Should AI developers be permitted to train models on copyright works? It should also ask: What should developers be required to tell us about how those works enter the training process?

The Open Access Dilemma
The open access community faces a risk that concerns about AI training will cause creators and institutions to shy away from openness. Researchers may become reluctant to make works openly available if they believe openness simply creates an unrestricted resource for commercial AI development. Publishers or repositories may respond with increasingly restrictive technological measures. Such reactions could unintentionally weaken the enormous gains achieved by open access.
Open access is based upon making knowledge more accessible while maintaining appropriate structures for attribution, licensing, integrity, and reuse. Training data transparency can complement these principles. It can help creators understand how openly available materials travel through AI ecosystems while allowing developers to demonstrate more responsible approaches to data acquisition and copyright compliance.
If researchers trust that the use of open materials within AI systems is visible and governed, they may have greater confidence in continuing to share their work. Attribution and reciprocity may afford further protection. Should the lineage of training data be traceable, rights holders can be properly attributed, and developers/researchers may have a clearer metric with which to evaluate data quality entering a model. Provenance can thus act as a mechanism for both upholding copyright responsibility and detecting low-quality/spoofed source documents from which to exclude future training.
Conclusion: Transparency Is Not the Enemy of Innovation
One concern is that mandating additional disclosures about training-data might burden AI developers excessively or reveal commercially sensitive information. Reasonable openness about training datasets should include appropriately detailed information about dataset origin and composition, major contributors or sources, kinds of content included, how it was collected, under what license (if any) it was assembled, and how rights reservations will be respected.
The European approach has emerged as a notable attempt at useful traceability and human auditability. Keeping detailed enough records of training material provenance to be able to retrieve the source data, if necessary, can help show responsible data stewardship, catch potential copyright issues earlier, aid in licensing determinations, and allow questionable or low-quality sources to be tracked down and rectified. Transparency, in this context, is not just transparency for transparency’s sake; it allows for better accountability and better-quality AI.
The copyright debate surrounding generative AI is entering a new phase. The first wave of discussion understandably concentrated on what AI produces: Who authored the output? Who owns it? Does it reproduce protected material? The next phase must also examine the systems behind those outputs. Training data transparency will not resolve every copyright question raised by GenAI. It cannot determine by itself whether a particular act is infringing, whether an exception applies, or what licensing model should prevail. For open access, the stakes are particularly high. The more sustainable path lies in building governance mechanisms that allow openness and accountability to coexist. If generative AI is to learn from the world’s knowledge, we should at least be able to understand, at a meaningful level, what it is learning from. That is not an obstacle to innovation. It may be one of the conditions for trusting it.

Discover open access articles about AI Data Transparency and Copyright on the AGOSR database:
Copyright and Ai in the Uk: Opting-In or Opting-Out?
Varieties of Transparency: Exploring Agency Within Ai Systems
The Evolving Role of Copyright Law in the Age of Ai-Generated Works
Bridging the Ai Ethics Gap a Tripartite Framework for Accountability, Implementation, and Governance






