Marking AI-generated content: how it works in practice (C2PA, watermarks, metadata)

Compliance7 min read·
K

Kees van der Vlies

Partner | IT Auditor

Also available in:Nederlands

The EU AI Act requires providers of generative AI systems to mark synthetic audio, image, video and text in a machine-readable way, so it is detectable as artificially generated. For new systems this obligation applies from 2 August 2026; for systems placed on the market earlier, from 2 December 2026, provided the postponement in the Digital Omnibus is formally completed. The law does not say how to mark, however. In this article we walk through the techniques used in practice, what they can and cannot do, and how to arrive at a defensible implementation.

What exactly does the law require?

Article 50(2) of the AI Act requires that the output of generative AI systems is marked in a machine-readable format and detectable as artificially generated or manipulated. The technical solution must be effective, interoperable, robust and reliable, as far as technically feasible, taking into account the state of the art and the costs. That wording is deliberately open: the legislator prescribes no standard and acknowledges that no technique is perfect. It also means that as a provider you must be able to explain the trade-offs you made. A visible label alone ("made with AI" in the corner of an image) is in any case insufficient: it is not machine-readable.

In practice there are three families of techniques, which complement each other: provenance metadata, watermarks and after-the-fact detection.

Technique 1: provenance metadata with C2PA

C2PA (Coalition for Content Provenance and Authenticity) is the most common open standard for recording the provenance of content. Major parties such as Adobe, Microsoft, Google and OpenAI participate. C2PA works with a manifest attached to the file: a set of cryptographically signed claims about provenance, for example that an image was generated by a particular AI model, when, and which edits were applied afterwards. The signature makes the manifest verifiable: anyone checking the file can establish that the claims originate from the signing party and have not been tampered with.

The strength of C2PA is verifiability and interoperability: it is exactly the type of machine-readable marking the AI Act asks for, and more and more tooling can read manifests. The weakness is the fragility of metadata. Take a screenshot and the manifest is gone. Many social media platforms strip metadata on upload, although this is slowly improving. So C2PA proves that content carrying a manifest is AI-generated, but the absence of a manifest proves nothing.

Technique 2: watermarks in the content itself

A watermark sits not next to but inside the content: a statistical pattern embedded in the pixels, the audio wave or the word choice of generated text, invisible to humans but detectable with the right detector. Well-known examples are the watermarking techniques that major model providers have built into their image and audio generation. Watermarks survive editing better than metadata: compressing, cropping or screenshotting does not easily remove a good watermark.

There are limitations in return. Detection usually requires the provider's own detector, which limits interoperability: there is no universal detector for all watermarks. For text, watermarking is fundamentally harder than for image and audio: paraphrasing or translating often removes the signal, and short texts contain too little material for reliable detection. And targeted attacks to remove watermarks remain a cat-and-mouse game.

Technique 3: after-the-fact detection

Detection classifiers try to recognize whether content is AI-generated without any marking, based on statistical characteristics. For a provider's own compliance question they are not a solution (they are not a marking), but they do play a role in the chain, for example at platforms assessing uploaded content. Be cautious with them in policy and communication: detectors produce false positives and false negatives, and "the detector says it is AI" is not proof. For text in particular, AI detectors are notoriously unreliable.

So what is a defensible approach?

Because no single technique is robust and interoperable and reliable on its own, a defensible implementation in practice comes down to a combination, which is also the prevailing recommendation in the European Commission's guidelines and the Code of Practice being developed for AI content marking. Concretely: attach C2PA manifests to generated content for verifiable, interoperable provenance, use the watermarking capability of the underlying model where available for robustness, and document the assessment: which techniques were applied, what is the state of the art, where are the limits of technical feasibility (short text, for example).

If you build on the API of an external foundation model, first check what the model provider already does. Many large providers already deliver generated images with a C2PA manifest and watermark. Your concern as provider of the derived system is then mainly: do not destroy that marking in your own processing steps (think of recompressing or re-rendering images), add marking where the provider delivers nothing, and record contractually what you rely on.

What does an auditor test?

If this subject comes up in an audit or gap analysis, expect the following test questions. Is there an inventory of systems that generate synthetic content? Has it been determined per system whether the marking obligation applies and who the provider is? Is the marking technically implemented and demonstrably working (sample test: generate content and verify the manifest or watermark)? Does the marking survive your own processing chain? Is the technical feasibility assessment documented? And is ongoing management assigned: who monitors whether model or pipeline updates break the marking? That last point is the most underestimated one in practice: a marking that worked at go-live and silently disappeared three model versions later is a realistic scenario.

Conclusion

Machine-readable marking of AI content is technically achievable, but requires a deliberate combination of C2PA provenance and watermarks, attention to your own processing chain, and documented trade-offs where the technology falls short. Start with the inventory and role determination, rely where possible on what model providers already deliver, and test the operation periodically. Want your marking approach reviewed, or this subject included in a broader AI Act gap analysis? Feel free to contact us.

Frequently asked questions

Which marking standard does the EU AI Act prescribe?+

None. The law requires marking to be machine-readable, effective, interoperable, robust and reliable as far as technically feasible. In practice, C2PA is the most common standard for provenance metadata, usually combined with watermarks embedded in the content itself.

Is a visible label such as 'made with AI' sufficient?+

No. A visible label is not machine-readable and therefore does not satisfy Article 50(2). It can be a useful addition to machine-readable marking, for example for deepfakes, which also carry a disclosure obligation towards the public.

What is the difference between C2PA and a watermark?+

C2PA attaches a cryptographically signed manifest with provenance information to the file; it is verifiable and interoperable, but is lost with screenshots or metadata stripping. A watermark is embedded in the content itself and survives editing better, but usually requires the provider's detector. That is why they are combined.

I build on the API of a large AI model. Do I need to mark content myself?+

First check what the model provider already delivers: many large providers already ship generated images with a C2PA manifest and watermark. Make sure your own processing does not remove that marking, add marking where the provider delivers nothing, and record contractually what you rely on.

Does marking also work for AI-generated text?+

To a limited extent. Text watermarks are vulnerable to paraphrasing and translation, and reliable detection is barely possible for short texts. The law acknowledges this with the 'as far as technically feasible' standard: document that assessment and use metadata at publication where possible.

Need help with compliance?

Need to comply with ISO 27001, ISO 42001, NEN 7510, NIS2 or DORA, or do you need a SOC 2 report? We guide you through the entire process: from gap analysis to implementation.

Explore Compliance Services

About the author

K
Kees van der Vlies

Partner | IT Auditor

Back to knowledge base

Have a question?

Get in touch for advice on IT audit, compliance and information security.

Contact us