Search

September 14, 2026

AI Content Tracking

The European Union, through the AI Act (EU 2024/1689), enforces transparency rules for AI-generated content starting August 2, 2026. AI providers wishing to comply for the services they offer within the European Union must ensure that generated content from their LLMs (such as text, images, audio, and video) includes machine-readable marking. This will allow systems to detect whether content has been produced or modified by AI. 

My technical focus centered on text watermarking, as adding metadata to files is relatively easy to understand. Google developed SynthID, an algorithm capable of adding watermarks to both files and text created by AI. The watermark is not an invisible ASCII character, as one might imagine. While generating text, the system slightly adjusts the probabilities of selecting specific tokens, creating a statistical word-selection pattern that can be easily detected later. In essence, the text itself becomes the watermark. 

Anthropic announced that it is integrating this technology (specifically Google DeepMind’s SynthID-Text) into all its new models and gradually rolling it out to older ones. Anthropic has also announced that it will offer a detection API in the future. This will make it possible to check whether a text contains this watermark, confirming it was generated or edited by Anthropic LLMs. 

 

 

OpenAI has similarly researched and developed comparable text watermarking technology. However, it has not publicly activated it within ChatGPT. For other AI providers, there does not yet appear to be the same universal adoption of this approach. 

Many questions arise: 

  • Will Anthropic actually make the watermark detection mechanism public, or will it merely offer an API that returns a simple “watermarked / not watermarked” status? 
  • If the exact mechanism remains proprietary, who will have the capacity to prove that a piece of text was created by AI? 
  • How effective can such a watermark be if it can potentially be removed or destroyed through simple translation, editing, or paraphrasing? 
  • If every AI-generated output carries an invisible tracking mechanism in the future, what is its ultimate purpose? 
  • Can the application of watermarking impact the quality of the generated text? 
  • Will token selection during watermarking affect token generation performance? 
  • How can watermarking be applied to code generation (perhaps only within code comments)? 

The EU emphasizes transparency, combating disinformation, and protecting against manipulation. However, the technology required for this transparency brings along a mechanism for origin tracing and potential digital content trackability. 

The central question is not whether we will be able to tell if a text was written by AI, but rather who will know it, who will be able to prove it, and ultimately who will control this information. 

 

 

 

 

Scroll to Top