Published on · 8 min read

Claude will sign everything it writes

Anthropic is slipping an invisible signature into everything Claude writes, and subscribers revolted before a single marked sentence existed.

On August 11, Anthropic posted a help page the way you file paperwork, with no announcement and no press release, a few paragraphs explaining that Claude would start marking what it writes[1]. Three days later the quiet approach had backfired, and the company published a full article answering subscribers who were calling the mark a scarlet letter and posting screenshots of their canceled plans on X[2].

Black line drawing on a mauve background, a hand holding a quill pen laid across two overlapping sheets of white paper.
The hand and the quill, the official illustration from the announcement (illustration Anthropic).

Three days from help page to bonfire

The help page reads like a memo. Models released from August 2 onward carry the mark at launch, older ones get it over the coming months, and the marking covers everywhere Claude speaks, the consumer app, the API, Claude Code, Cowork, and Claude Tag, including when it runs on Amazon, Google, or Microsoft infrastructure. For the images and SVG files it produces, Anthropic also attaches provenance metadata in the C2PA format, the open standard that shows whether a file has been altered since it was made.

One sentence lit the fire, the one noting that the mark travels with the text when you paste it elsewhere and sometimes survives editing. Dozens of subscribers read that as betrayal. One called it a conspiracy against innocent Claude users, another compared the mark to a scarlet letter that would stay stuck to the text long after it left the app. Most of the complaints came from people who said they only use Claude to proofread their own writing, plus developers worried that a cryptographic signature would degrade their code[3]. On Reddit, where the same flood was expected, most people approved, and one commenter put it with a bluntness worth quoting, that the only reason to refuse the mark is wanting to lie to somebody.

The calendar explains most of it. On August 2, Article 50 of the EU AI Act took effect and now requires generative systems to mark their output in a machine-readable format, an obligation I wrote about here as it landed. Anthropic is one of roughly 190 signatories to the code of practice drafted by the European AI Office, alongside Google, Meta, Microsoft, OpenAI, Mistral, and Cohere[4]. So the watermark is no house eccentricity, and the competitors who haven't announced theirs will get there.

A coin bent one percent

The method has a name, SynthID-Text, out of Google DeepMind and published in Nature in 2024[5]. Nothing gets added to the text, no hidden character, no extra token, and the bill doesn't move by a cent. The watermark lives in the model's hesitations. When Claude writes that the weather was cold and reaches for the next word, "overcast" and "gray" are equally good, so a secret key can tilt the draw between them and nobody sees a thing.

One choice proves nothing, the way a coin bent one percent tells you nothing on a single toss and everything after a thousand. The detector doesn't read the text, it counts, comparing how often the favored words turn up against what chance would produce. Human writing lands in the right buckets about half the time, and marked writing lands there a little too often for luck.

Google has run this inside Gemini since 2024 and measured the effect across some twenty million responses, with users reporting no difference in quality, creativity, or speed[6]. Anthropic promises the same neutrality and adds the detail that should have defused half the revolt, that the mark carries no identifying information and traces back to no person, no company, and no conversation. It says a model came through. It never says which of its users put that model to work.

No thin space ever ratted anyone out

The stubborn legend of the invisible characters refuses to die. For two years, whole threads have catalogued zero-width spaces and narrow no-break spaces found in model output, presented as proof of covert marking. They're typographic habits picked up during training, and their fragility alone disqualifies them, because a find-and-replace wipes them in a second[7]. A watermark you delete by accident while pasting into a form is not a watermark.

The case against the em dash rests on the same confusion, when all it reveals is a writing habit. This blog is an accidental counterexample, since French typography demands a thin space before a semicolon, an exclamation point, a question mark, and inside its guillemets, which means every one of my French articles carries dozens of them. A detector tuned to that character would convict correct French typesetting and clear anybody who sets type carelessly. Real watermarks hide elsewhere, in the statistics, where the eye doesn't go.

What the mark will never tell you

Anthropic lists the limits honestly, and they're severe. The watermark doesn't confirm that a human wrote a text, doesn't recognize output from rival models, vanishes under a full rewrite, and weakens with every edit. It works poorly on short passages and on factual ones, where too few interchangeable words exist to hold a preference, which settles the question for code and should calm the worried developers. The help page also grants that a detected mark is a signal and never a conclusion, and that its absence proves nothing about whether a machine was involved.

What worries me is what people will do with that signal. The signature attests that a model was involved somewhere, without separating the text Claude wrote from the text Claude merely fixed. A teacher or a recruiter who gets the result will read it as a verdict on authorship anyway, because that is the verdict they went looking for. Anybody writing in a second language who runs a draft past a machine ends up with a fully marked text carrying an entirely original thought, which promises some memorable injustices.

The detector doesn't exist yet. Anthropic has announced a verification API, and announced is the operative word, since nothing has shipped[8]. No school, no employer, and no platform can read the mark, which makes for an odd object, a signature whose only reader is the party stamping it. The accusations won't wait, of course, and they will lean on the homemade detectors whose accuracy we already know.

The diligent cheater sleeps fine

The asymmetry shows up the moment you try to beat the thing, because beating it takes almost no work. Run the marked text through a model that doesn't mark, DeepSeek, Qwen, a Llama running on your own machine, and the statistics dissolve. The student running an industrial cheating operation knows that trick. The one who asked Claude to reorder a paragraph doesn't, and that's the one the mark will finger. A rule that catches only the honest and absent-minded deserves a second look.

OpenAI decided the other way, and that decision remains the most instructive part of the story. The company built its own text watermark, measured it at 99.9 percent reliability in internal testing, then left it in a drawer for two years. The stated reasons included the risk of false accusations and the damage to non-native English speakers, but one number outweighed the rest, since an internal survey found that nearly 30 percent of users would use ChatGPT less if the product carried a mark its competitors didn't[9]. Transparency was ready, and competition put it away.

That is exactly what the European regulation repairs, by requiring the mark of everybody at once rather than of whoever is brave enough to go first. The paradox is worth savoring, since a continent that builds none of the models in question has gotten a California company to sign its output worldwide, including for American users who never voted for any of it and are saying so at length.

I'm in favor of the watermark, and not only because Brussels demands it. Knowing whether a text came out of a machine belongs to the same basic right as knowing who you are talking to, and I'll take an imperfect mechanism over the blindness that preceded it. But a mark only its issuer can read is a promise rather than a proof, and everything turns on the day that verification API opens to schools, employers, and newsrooms.

The best part is the timing. No marked model has shipped as I write this, since the current generation predates August 2 by several months[7:1]. People canceled their subscriptions to protest a signature that never appeared in a single one of their conversations, which is, you have to admit, a fairly original way to argue for authenticity.


  1. Anthropic, How Claude marks AI-generated content, help page published August 11, 2026. ↩︎

  2. Anthropic, How Claude's text watermarking works, August 14, 2026. ↩︎

  3. Amanda Silberling, Some Claude users are mad that Anthropic's new watermarks will catch them cheating at their jobs, classes, TechCrunch, August 12, 2026. Gizmodo collected the cancellation screenshots posted on X in Anthropic Explains Its Watermark System as Some Claude Users Loudly Revolt. ↩︎

  4. The code of practice on transparency of AI-generated content was published on June 10, 2026 and had gathered around 190 signatories by the end of July. ↩︎

  5. Sumanth Dathathri et al., Scalable watermarking for identifying large language model outputs, Nature, October 23, 2024. Google DeepMind open-sourced the code the same week. ↩︎

  6. Rhiannon Williams, Google DeepMind is making its AI text watermark open source, MIT Technology Review, October 23, 2024. ↩︎

  7. The clearest piece I have read on the subject, including the invisible-character legend and the real state of the rollout, is Pain in the Agent's AI text watermarks. BleepingComputer separately confirms that no model released before August 2 carries the mark yet, in How Anthropic plans to watermark Claude's AI-generated text. ↩︎ ↩︎

  8. Anthropic mentions a detection API "coming soon" in its August 14 article, with no date and no access terms. ↩︎

  9. Deepa Seetharaman and Matt Barnum, There's a Tool to Catch Students Cheating With ChatGPT. OpenAI Hasn't Released It, The Wall Street Journal, August 4, 2024. Simon Willison relayed the figures outside the paywall. ↩︎

  1. Europe regulates the AI it doesn't build

  2. Fable 5 Is Back, the Government in the Loop

  3. The Biter Bit, Anthropic Accuses Alibaba of Copying Claude