SAN FRANCISCO — AI safety startup Anthropic has confirmed it will begin embedding machine-readable watermarks into text generated by its Claude AI models and attaching digital provenance metadata to generated media files. The move is designed to enhance digital content transparency and align with strict legal obligations under the European Union’s AI Act.

The new watermarking system applies globally to supported Claude models across various deployment channels, including the Claude Platform API, Claude Code, Claude Cowork, and Claude Tag.

"As AI-generated content becomes commonplace, greater transparency and signals about where content comes from can give people useful context about the information they consume," Anthropic stated in an official support disclosure detailing its compliance with the EU AI Act’s Transparency Code.

How the Watermarking System Operates

Anthropic’s watermarking mechanism operates across two primary formats:

  1. Embedded Text Watermarks: For text generation, supported Claude models weave an imperceptible, machine-readable signal directly into the text output during generation. Anthropic emphasized that the watermark is completely invisible to human readers and does not alter the quality, tone, or meaning of the response. Because the signal is embedded into the text structure itself, it stays intact when copied, pasted, or transferred across different applications.

  2. Signed Provenance Metadata for Files: For visual and exported file formats—including PNG, JPG, and SVG graphics—Claude attaches digitally signed provenance metadata adhering to the Coalition for Content Provenance and Authenticity (C2PA) open standard. C2PA is widely adopted across the tech industry to track content origins and history.

The company noted that text watermarking is implemented directly at the foundational model level, ensuring that any application powered by a supported Claude model will automatically include the trace.

EU AI Act Timeline and Detection Tools

Claude models released in the European Union carry the machine-readable watermarking system. For models launched prior to this date, European legislation permits a transitional grace period during which Anthropic will retrofit marking support into legacy systems.

To make the system actionable, Anthropic revealed it is building detection mechanisms that will allow users and third-party platforms to scan text or media files to verify if they were processed by Claude. Technical documentation for these detection tools is scheduled to be released in the coming weeks.

Key Technical Limitations

Despite the breakthrough, Anthropic explicitly cautioned that digital watermarking is not a silver bullet for content authenticity.

Detecting a watermark indicates that content was processed by Claude, but it does not guarantee full provenance. For instance, users frequently rely on Claude to edit, summarize, translate, or proofread existing text; in such cases, the underlying ideas or original draft may belong to a human creator even if a Claude mark is present.

Conversely, the absence of a watermark does not guarantee a piece of content is purely human-made. Anthropic outlined several scenarios where watermarks may be unreadable or absent:

  • Content generated using legacy models released prior to watermark support.

  • Text that has undergone extensive human editing, paraphrasing, or mixing with non-AI content.

  • Very short snippets or sentences that lack sufficient word volume to embed a stable signal.

  • Files whose metadata was stripped during format conversions, compression, re-saving, or screen captures.