How could the presumption of use of cultural content by AI providers rebalance the burden of proof?
Introduction
Generative artificial intelligence systems are trained on vast quantities of text, images, music and audiovisual content. For authors, artists, publishers, producers and collective management organisations, however, one central difficulty remains: how can they prove that a protected work was actually used to develop or deploy an AI model when the training data remain largely opaque?
A bill dated December 12, 2025, introduced by French Senator Laure Darcos, seeks to address this imbalance by establishing a presumption of use of cultural content by artificial intelligence providers. Adopted by the Senate on 8 April 2026, the bill was then transmitted to the National Assembly. It does not create a new intellectual property right: it changes the rules of evidence so that right holders can enforce their rights more effectively.
Why is it so difficult to prove the use of works by AI?
Model training relies on data that are rarely accessible
To establish copyright infringement, the right holder must, in principle, identify the work concerned, establish its rights and characterise the alleged acts of reproduction or exploitation. This becomes particularly difficult when content has been absorbed into vast datasets, assembled by several service providers and used to train a model whose internal workings are not public.
To learn more about the issues arising from the use of protected works to train AI systems, we invite you to read our previously published article.
Right holders may sometimes identify similarities in a generated output, prompt the reproduction of an element resembling their work or identify a reference to content in technical documentation. These elements do not, however, always make it possible to establish the origin of the training data with certainty.
The opt-out mechanism does not by itself resolve the evidentiary difficulty
Text and data mining consists of automatically analysing large quantities of digital content in order to extract information, trends or correlations. Article 4 of Directive (EU) 2019/790 authorises, subject to certain conditions, the reproductions and extractions necessary for such analysis where they concern content that has been lawfully accessed. For mining carried out for any purpose, however, this exception applies only if the right holders have not expressly reserved their rights. This ability to object to text and data mining is commonly referred to as an “opt-out”: the right holder expressly indicates, in particular by machine-readable means, that they do not consent to their content being used for this purpose.
Under French law, these rules are set out in Article L. 122-5-3 of the French Intellectual Property Code.
In practice, a reservation of rights is not always sufficient to protect the right holder. The right holder may object to the use of its work without being able to determine whether it was incorporated into a training dataset or whether its objection was respected. Without access to the technical information held by the provider, it may therefore remain difficult to prove unauthorised use.
How would the presumption of use of cultural content operate?
The right holder would have to provide an indication making the use plausible
The bill provides that subject matter protected by copyright or related rights would be presumed to have been used by an AI provider where an indication relating to the development or deployment of the system, or to the output it generates, makes that use plausible.
It would therefore not be an automatic presumption applicable to every work available online. The claimant would have to provide a sufficiently specific prima facie showing. A mere abstract assertion that a model must necessarily have been trained on cultural content should not be sufficient.
- The right holder submits one or more credible indications.
- The provider, which holds the technical information, may produce evidence to the contrary.
The mechanism thus places the burden of proof more closely on the party that actually possesses the information needed to determine the origin of the data and the conditions under which they were used.
The presumption would remain rebuttable
The AI provider could rebut the presumption by demonstrating, for example, that the work was not incorporated into the corpus, that it was used under a licence, that it came from a source covered by an authorisation or that the use validly fell within an exception.
The proposal does not therefore mean that every provider would automatically be considered an infringer. It would create an evidentiary tool, not strict liability.
In a favourable opinion issued on March 19, 2026, the French Conseil d’État considered that the mechanism could be reconciled with European Union law, subject in particular to using the more neutral term “use”, protecting trade secrets and limiting the mechanism to civil matters.
The presumption would encourage the use of licensing
The presumption would relate only to whether content was used; it would not, by itself, establish that the use was unlawful. A provider could therefore show that the work concerned was covered by a licence or another authorisation. This mechanism would encourage providers to document their sources and enter into agreements with right holders in order to secure the training of their models. The presumption would thus serve as a lever for accountability and negotiation, potentially fostering the development of individual or collective licensing arrangements.
What evidence could trigger the presumption?
The bill does not set out an exhaustive list. The assessment would therefore have to be carried out by the court on a case-by-case basis.
Evidence derived from generated outputs
- the repeated generation of elements substantially similar to a work;
- the reproduction of a passage, image, composition or details unlikely to result from mere coincidence;
- the appearance of signatures, watermarks, copyright notices or metadata associated with the original content;
- the system’s ability to reproduce a very precisely identified creative universe.
An isolated similarity would not necessarily constitute sufficient evidence. The analysis would notably have to distinguish the reproduction of protectable elements from the reproduction of a style, an idea, a genre or commonplace characteristics.
Evidence relating to the development of the model
- documentation published by the provider;
- a description of the datasets used;
- transparency reports;
- statements by service providers or researchers;
- licences acquired for certain categories of content;
- information disclosed during an expert investigation or a court-ordered evidentiary measure.
How would the presumption interact with European law?
It would not abolish the text and data mining exception
The presumption would not directly alter the conditions of the text and data mining exception. It would operate upstream of the legal analysis in order to determine whether the content was used. Once that use had been established or presumed, it would remain necessary to determine whether it was:
- authorised by a licence;
- covered by the text and data mining exception;
- carried out despite a valid reservation of rights;
- or constituted copyright infringement
- or an infringement of related rights.
The bill would therefore not automatically guarantee compensation. Its main effect would be to prevent a claim from failing before any examination of its merits because of a lack of access to evidence.
It would complement the transparency obligations under the AI Act
The AI Act requires providers of general-purpose AI models to put in place a policy designed to comply with European Union copyright law, including reservations of rights, and to publish a sufficiently detailed summary of the content used for training. Those transparency requirements do not, however, necessarily provide access to an exhaustive, work-by-work list of all the data incorporated.
The French presumption would thus serve a distinct function: facilitating the resolution of a civil dispute where the available information makes use plausible but does not yet make it possible to prove it directly.
Conclusion
The presumption of use of cultural content by artificial intelligence providers would not resolve every conflict between creation and AI. It would nevertheless provide a targeted response to one of the main obstacles encountered by right holders: the inability to prove technical facts under the exclusive control of their opponent.
Its effectiveness will depend on the definition of sufficient indications, the protection of trade secrets, its interaction with European law and the powers of the court to obtain reliable information. Cultural businesses already have an interest in formalising their reservations of rights and structuring the collection of evidence.
Dreyfus & Associés assists its clients in managing complex intellectual property cases, offering personalised advice and comprehensive operational support for the complete protection of intellectual property.
FAQ
Can an AI system freely use every work available online?
The fact that content is available online does not mean that it is in the public domain. Its use may be covered by a licence or a statutory exception, or may require the right holder’s authorisation.
Is similarity between a generated image and a work sufficient?
Not necessarily. The similarity must be assessed in light of the original elements reproduced, the circumstances of generation and the other available evidence.
Can trade secrets prevent any disclosure of training data?
Trade secrets must be protected, but they do not necessarily preclude every evidentiary measure. Confidentiality mechanisms may allow a court or an expert to access certain information without making it public.
Why must AI providers strengthen the traceability of their data?
To rebut the presumption, they would have to be able to document the provenance of the data, the licences, filtering operations, the handling of opt-outs and the role of their subcontractors. Rigorous governance of training data would therefore make it easier to demonstrate lawful use or the absence of use of the work concerned.
What should a business do if its content appears to have been used?
It should preserve reproducible evidence (outputs, prompts, settings, date and model version), identify the works and rights concerned, review the relevant licences and reservations of rights and, with specialist counsel, determine the appropriate evidentiary measures and courses of action.


















