How is text watermark work?

Introduction

Anthropic recently announced that future models will add watermarks to text using SynthID-Text: How Claude’s text watermark works🔗. The official statement says:

  • We use a method of watermarking that does not have any practical impact on the quality or content of Claude’s outputs;
  • The difference between watermarked and un-watermarked text will not be distinguishable to readers;
  • Nothing is added to the text and there are no hidden characters;
  • Watermarking doesn’t require extra tokens, and will not be more expensive;
  • Watermarking carries no identifying information and can’t be traced to a specific person, organization, or chat;
  • Watermarking won’t be specific to Claude. As of August 2, the EU requires AI providers serving its market to mark AI-generated content. Other major model developers have signed the same Code of Practice and will be implementing their own watermarks.

So I wanted to look into how LLM text watermarking is implemented and what impact it has on users.

Greenlist and Redlist

How Is a Watermark Added?

An LLM is a next-token machine: every token in its response is determined probabilistically. To infer whether a watermark has been added from the generated text, statistical features can be embedded during the generation process. According to A Watermark for Large Language Models🔗, one method embeds a watermark by changing the model’s “preference probability for which words to choose,” by classifying the token list at each generation step into two groups:

  • 🟢 Greenlist: tokens the model is slightly encouraged to choose, increasing their probability
  • 🔴 Redlist: tokens the model is not specifically encouraged to choose, reducing their probability

A watermark does not mean certain words are always green or red. Instead, the lists are generated dynamically based on the context at the time, so that to someone without the key it looks random, while someone with the key can reproduce the rule.

(secret key + surrounding words + pseudorandom generation function = seed), and the seed is used to influence the probability of the next word, thereby writing the watermark into the resulting content

By setting a secret key known only to the model and detector, when preparing to generate the next token, the “preceding text” and “key” are passed to a pseudorandom function to generate a random seed. This seed is used to decide which tokens are “green” and which are “red” for that token selection step.

User prompt

LLM generates the next token

Determine the greenlist based on the key and preceding text

Slightly increase the probability of greenlist tokens

Model selects a token

Append the token to the text as the new preceding text

Final watermarked text is generated

Soft Watermarks and Hard Watermarks

There are two ways to implement a watermark, differing in how strictly redlist tokens are constrained:

Hard Watermark

  • Mechanism: Completely prohibits the model from choosing tokens on the redlist
  • Pros: In the generated text, the greenlist hit rate is theoretically 100%, making detection very easy
  • Cons: Severely degrades text quality. Low-entropy (highly predictable) sequences may be forced to choose the wrong word. For example, “The capital of France is” is almost certainly followed by “Paris,” but if “Paris” happens to be on the redlist, the model can only output an unnatural result

Soft Watermark

  • Mechanism: Does not prohibit the redlist; instead, it makes the model “slightly prefer” greenlist tokens
  • Pros: Also uses a Z-test. The greenlist hit rate will be higher than 50% but lower than 100%, requiring more text to reach statistical significance. At high-entropy positions (where multiple candidate words have similar probabilities), the watermark takes effect; at low-entropy positions (where one word clearly dominates), the bias has almost no effect, thus preserving the naturalness of the text
  • Cons: Suitable parameters must be tuned manually
Comparison ItemHard WatermarkSoft Watermark
Redlist TokensCompletely prohibitedAllowed
Detection StrengthStrongestLower (requires longer text)
Minimum Detectable Text LengthVery shortLonger (depends on text entropy)
Impact on Text QualitySevere degradationMinimal impact
Low-Entropy HandlingForces unnatural alternativesAdaptive, with little watermark impact
Practical DeployabilityPoorGood

How Do We Determine Whether a Watermark Exists?

  • The article has T tokens in total
  • Among them, S tokens fall on the greenlist
  • The proportion of the vocabulary that belongs to the greenlist is γ (typical value: 0.5)

If there is no watermark, we can expect the greenlist hit rate to be approximately γ. For example, with an article length of 1,000 tokens and a greenlist proportion of 50%, the probability that normally randomly generated text would show a 65% or 75% greenlist ratio is very low, so we can conclude that it may have been watermarked.

When the observed greenlist hit rate deviates from expectation, a statistical test is needed to determine whether this is random fluctuation or evidence of a watermark.

  • Z-score (test statistic) represents how far the observed value is from the mean, measured in units of standard deviation:
  • P-value (significance p-value) represents, assuming there is no watermark, the probability of observing the current result or a more extreme one. The value ranges from 0 to 1; the smaller it is, the more statistically significant the result.

Using an article length of 1,000 tokens and γ = 0.5 as an example:

ScenarioGreenlist Hits (S)Hit RateZ-scoreP-valueJudgment
No watermark (expected)50050%0.001.00Normal
Actual detection65065%+3.16< 0.001Suspicious
Actual detection75075%+6.32< 0.0001Extremely suspicious

SynthID-Text

How Is a Watermark Added?

Similarly, preferences are embedded through statistical features during text generation. However, compared with forcibly changing the selection probability of the next word according to seed-based classification, SynthID-Text breaks the process into two main steps:

  • Selecting the roster (handled by the AI): The AI picks a batch of the most fluent and reasonable candidate words, deciding who is eligible to enter the tournament (the AI’s probabilities determine the number of “entry tickets”). Suppose the tournament has 10 total slots:
    • A (70%) takes 7 slots
    • B (20%) takes 2 slots
    • C (10%) takes 1 slot
  • Determining the winner (handled by the g-value): Once in the tournament, the AI’s original probabilities no longer matter. The g-value assigns each participating word a random score between 0 and 1 based on the combination of “preceding text + secret key.” During the match, it is purely a comparison of whose g-value is higher; that token advances.

The core Tournament Sampling algorithm works as follows:

  1. Generate a random seed: Generate a random seed through a hash function based on the “previous several tokens” and the “secret key”
  2. Compute g-values: Use m pseudorandom functions g₁, g₂, …, gₘ to assign scores to each candidate token in the vocabulary
  3. Tournament sampling:
    • Sample 2m candidate tokens from the LLM distribution (duplicates are possible)
    • Round 1: Pair up candidate tokens; the one with the higher g₁ score wins
    • Round 2: Pair the winners again; the one with the higher g₂ score wins
    • Repeat until round m; the final winning token becomes the next generated token
  4. Repeat the process: Repeat the above steps for every generated token

SynthID-Text does not alter output quality because it does not forcibly interfere with the AI’s selection result. Instead, it runs a tournament after the AI has made its decision, hiding the watermark within the process.

Detection Mechanism

SynthID-Text detection only requires the text, the key, and the random seed generator. The detector computes the g-values of all tokens in the text and uses a scoring function to determine whether the text contains a watermark. The detector provides three states:

StateMeaning
WatermarkedDetermines with high confidence that the text contains a watermark
Not WatermarkedDetermines with high confidence that the text does not contain a watermark
UncertainCannot determine; more text or human judgment is needed

Limitations and Challenges

Detection effectiveness depends on the “degrees of freedom” in the text. When the model generates in low-entropy scenarios (such as factual answers or code), the word selection space is constrained, making watermark embedding weaker. In addition, translation or substantial rewriting also reduces detection confidence.

In terms of detection scope, each AI provider uses different keys and algorithms. Claude’s detection API cannot detect outputs from GPT or Gemini. A watermark can only answer “whether this text was generated by Claude”; it cannot confirm whether the text was created entirely by a human.

Summary

The SynthID-Text method has always reminded me of a story about the Florentine Republic’s “lottery by leather bags”: only the pre-arranged “high-probability words / insiders” can obtain tickets to enter the tournament. The g-value scoring PK looks random, but in reality the distribution still follows the statistics provided by the AI.

Text watermarks have regulatory reasons behind them and can be useful for verification in certain contexts. I don’t think creators need to be too anxious about their works being watermarked. The resistance comes from the fact that AI-generated content is often associated with “cutting corners, cheapness, and lack of creativity,” but similar issues have occurred in contexts such as computer-aided graphics, photography, industrialization, and so on. Should we really resist AI output, or should we spend our time focusing on the quality of the work itself? And is there even a single metric for measuring quality?

Further Reading