OpenAI is preparing to quietly change the way ChatGPT and Codex write text for users in the European Union.
The change will not add a visible label, hidden Unicode characters or unusual formatting. Instead, eligible AI-generated text will carry an invisible statistical watermark embedded in the model’s word choices, giving approved detectors a way to estimate whether a passage was generated or processed by an OpenAI system.
OpenAI calls the technology textGrain.
The company says the EU rollout will begin over the coming weeks, while API customers worldwide can already opt in to watermarking for supported models. OpenAI is not making the feature a global default yet.
The move is designed in part to comply with Article 50 of the EU AI Act, which requires providers of generative AI systems to make synthetic text, images, audio and video machine-readable and detectable as artificially generated or manipulated where technically feasible.
That sounds like the beginning of a universal AI detector.
It is not.
OpenAI’s own testing shows why.
The watermark lives inside the wording itself
TextGrain works by subtly influencing which words or word pieces the model selects while generating a response.
Instead of changing what a reader sees, it changes the statistical pattern behind the prose. Across a sufficiently long passage, those choices can produce a signature that a detector looks for.
That gives watermarking one important advantage over metadata.
Metadata can disappear when someone copies text into another application, exports a file or strips document properties. A watermark embedded in the wording can potentially survive ordinary copying and some editing.
OpenAI says this is why provenance needs multiple layers rather than one universal mechanism. Its broader approach combines machine-readable metadata, Content Credentials, watermarks and verification services depending on the media type.
For text, however, the problem is much harder.
Words are extremely easy to change.
Change a quarter of the words and detection collapses
OpenAI published unusually candid numbers with its announcement.
In testing on 400-token passages, its detector identified the watermark about 92% of the time under the original conditions. Replacing only 10% of the words with synonyms reduced detection to approximately 66%.
Replace 25% of the words and detection fell to just 17%.
That is a significant limitation.
Someone does not necessarily need sophisticated watermark-removal software. Ordinary rewriting can degrade the signal.
A person could:
- revise wording manually;
- use another AI model to paraphrase the passage;
- translate it into another language and back;
- shorten it substantially;
- combine AI text with human writing.
The result may still have originated largely from ChatGPT while no longer producing a strong watermark signal.
This is why OpenAI warns that failure to detect a watermark cannot prove that a person wrote the text without AI assistance.
That distinction is likely to become extremely important.
Longer text is easier to detect
Detection also depends heavily on length.
At a targeted 1% false-positive rate, OpenAI says its detector identified the watermark in roughly 80% of 200-token psychology responses and about 95% of comparable 400-token responses.
Mathematical text performed considerably worse because a model has less freedom to choose among different words when expressing equations or tightly constrained technical concepts.
That makes intuitive sense.
Watermarking needs room for the model to make stylistic choices.
If there are dozens of reasonable ways to write a sentence, textGrain can subtly influence those choices.
If there is essentially one correct mathematical expression, there is far less room to encode a statistical pattern.
Short emails, code snippets and brief answers therefore pose a fundamentally different detection problem from essays or long reports.
This is not an AI plagiarism detector
This may be the most important point for employers, schools and publishers.
A detected watermark does not tell you who generated the text.
OpenAI says textGrain does not reveal the person’s identity, account, original prompt or ChatGPT conversation. It also cannot determine how much of the final text was written by a person versus generated or edited by AI.
It also does not establish plagiarism.
An employee could legitimately use ChatGPT to refine a memo and leave the watermark intact.
A student could heavily rewrite AI-generated material until the watermark disappears.
A journalist could use AI for copyediting but still be responsible for every factual statement.
The watermark tells you something about provenance.
It does not tell you whether that use was permitted, dishonest or appropriate.
OpenAI puts the limitation directly: watermarking does not establish ownership or responsibility and does not verify whether the underlying text is accurate.
That should prevent organizations from making a common mistake with AI detectors: treating a probabilistic signal as dispositive evidence.
OpenAI is restricting access to the detector
Another unusual aspect of the rollout is that OpenAI is not making its text detector publicly available at launch.
Approved researchers and expert organizations can apply for access, but OpenAI says it wants to evaluate reliability and responsible use before wider distribution.
That decision makes sense given the false-positive and false-negative problem.
Imagine an employer using a public detector to accuse an employee of misconduct.
Or a university automatically disciplining a student because a detector produced a positive result.
A false positive could have serious consequences.
A false negative could also create false confidence.
OpenAI appears to be trying to avoid positioning textGrain as a forensic truth machine.
The EU AI Act is driving the change
The regulatory context is just as interesting as the technical implementation.
Article 50 of the EU AI Act requires providers of AI systems that generate synthetic text, audio, images or video to ensure that outputs are marked in a machine-readable format and can be detected as artificially generated or manipulated, subject to technical feasibility and certain exceptions.
Those obligations became applicable on Aug. 2, 2026, although the European Commission has established a limited transition period for some systems placed on the market earlier.
The Commission says the rules are intended to reduce risks involving impersonation, fraud, misinformation and manipulation by helping people and platforms understand when AI-generated content is involved.
OpenAI also signed the EU Code of Practice on Transparency of AI-Generated Content, which the European Commission has recognized as a mechanism that can assist companies in demonstrating compliance, although adherence does not automatically prove compliance with the law.
That means watermarking is no longer merely a research project.
It is becoming a regulatory control.
The penalties make provenance more than an academic issue
Article 50 violations can potentially expose organizations to significant penalties.
The European Commission says violations of these transparency requirements may result in fines of up to €15 million or 3% of worldwide annual turnover from the previous financial year, depending on the circumstances and applicable proportionality rules.
For large AI providers, that changes the economics.
A provenance mechanism that once might have been viewed as an optional trust feature can become part of a formal compliance architecture.
And it will not affect OpenAI alone.
Any provider offering qualifying generative AI systems in Europe needs to consider how its outputs can be identified in a machine-readable way.
Watermarking may become part of enterprise AI governance
Enterprises should pay attention as well.
If a company deploys an AI system that generates customer-facing content, regulatory notices, product descriptions, news material or public communications, it may eventually need to understand whether provenance signals survive the entire content pipeline.
For example:
AI generates text
↓
Employee edits text
↓
CMS reformats it
↓
Translation system rewrites it
↓
Publisher posts it
Did the watermark survive?
Did the company have an obligation to preserve it?
Was additional disclosure required?
Those questions are likely to become part of AI governance programs.
The EU rules also distinguish between provider obligations and deployer obligations. Providers may need to embed machine-readable marks, while deployers have separate duties in certain contexts, including disclosure requirements for deepfakes and AI-generated text involving matters of public interest when there has not been appropriate human review or editorial control.
Simply receiving watermarked output from an AI provider therefore may not satisfy every downstream transparency obligation.
Quality does not appear to suffer much
One obvious question is whether forcing the model to choose words according to a watermark pattern makes the writing worse.
OpenAI says its current evaluations do not show a meaningful performance decline on GPT-6 Astra.
Across several internal and external benchmarks, the watermarked and unwatermarked versions produced broadly similar results. Some benchmarks slightly favored the watermarked version, while others slightly favored the ordinary output.
If that result holds in real-world use, watermarking could become relatively inexpensive from a product-quality perspective.
The harder problem is durability.
A watermark that disappears after moderate rewriting provides limited value as definitive evidence.
The future is probably layered provenance, not one watermark
That may be the biggest lesson from OpenAI’s announcement.
There probably will not be one magical “AI detector.”
Instead, provenance is likely to work through several overlapping mechanisms:
- statistical text watermarks;
- embedded media watermarks;
- cryptographic Content Credentials;
- model-provider verification services;
- audit logs;
- platform disclosure requirements;
- visible labeling in certain circumstances.
Each solves a different part of the problem.
Metadata can carry detailed provenance but can be stripped.
Watermarks can survive some transformations but can weaken after editing.
Logs can prove what happened inside an organization but may not travel with the content.
Visible labels help humans but can be removed.
The compliance system works best when several mechanisms reinforce each other.
OpenAI itself now says no single provenance technique is sufficient on its own.
The invisible watermark is important precisely because it is imperfect
The most interesting part of textGrain is not that OpenAI has solved AI detection.
It has not.
The interesting part is that one of the world’s largest AI providers is now modifying the statistical construction of generated language specifically to satisfy a regulatory transparency requirement.
That is a significant shift.
AI regulation is beginning to shape not just what companies disclose in their policies, but how models generate words at the technical level.
The EU AI Act is effectively reaching inside the generation process.
For users, the experience may look identical.
ChatGPT answers a question.
The words appear normally.
Nothing announces that they contain a watermark.
But underneath those words, the model has made slightly different choices so that another machine may later recognize where they came from.
That is a remarkably subtle form of regulation.
And if Article 50 becomes the template other jurisdictions follow, invisible provenance may become a standard feature of generative AI long before most users ever realize it is there.