The European Union now requires providers of generative AI systems to make synthetic text and other outputs detectable in a machine-readable form. Anthropic has said that new Claude models will mark generated content as part of its response to those transparency obligations.

The objective is reasonable. As synthetic content becomes harder to distinguish from human work, provenance matters.

But provenance creates a second problem when institutions ask the signal to establish more than it can.

A watermark may be evidence that an AI system participated in producing or transforming text. It is not, by itself, a finding about who authored the work, whether a policy was breached, or whether anyone committed misconduct.

The difficult question is not whether the mark can be detected. It is what the detection is permitted to mean.

Start with the proposition

A detector result can enter a decision process at several different levels:

  1. A machine-readable signal was detected.
  2. The signal is consistent with processing by a particular system or class of systems.
  3. The system played a particular role in producing the work.
  4. The person did not author the work, breached a policy, or committed misconduct.

These are not interchangeable conclusions. Each step moves further from the technical signal and requires additional evidence.

The first may be a direct detector output. The second depends on the detector's scope, reliability, and operating conditions. The third requires evidence about the workflow. The fourth requires a governing rule, proof that the rule applied, and an accountable evaluation of the person's conduct.

The watermark is an artifact. Authorship, policy breach, and misconduct are findings.

That distinction is easy to lose when a technical tool returns a confident label.

The same signal can describe different work

Consider two writers.

The first enters a short prompt and asks an AI system to produce an entire article.

The second spends several days researching and writing an article, then asks the same system to improve the flow, correct the grammar, and tighten several sentences.

Both resulting documents may contain a detectable mark. Yet the human contribution in the two cases is plainly different.

If the detector reports the same result for both, it has not distinguished generation from editing. It has identified a technological event without reconstructing the creative process.

In an earlier article, I argued that asking whether AI was used is too crude a test for authorship. Watermarking makes the institutional consequence of that problem clearer. Once a detectable signal enters a university, workplace, publisher, professional body, or court, a statement about system involvement can quickly become a conclusion about a person.

That transition is where provenance overreach occurs.

Detector evidence has operating limits

Anthropic has disclosed the watermark's general properties. It describes an imperceptible mark woven directly into generated text at the model level. The mark travels with copied text and may persist through some editing. This establishes that the watermark is embedded in the generated language rather than attached as ordinary file metadata.

The company has not yet published the technical detection documentation on which a consequential use would depend. Its public guidance does not specify the detector, decision threshold, minimum reliable text length, or measured false-positive and false-negative rates. It would therefore be premature to attribute the mark to a more specific mechanism, such as statistically biased word selection governed by a secret key, unless Anthropic confirms that design.

Anthropic also states the central limitation directly: a detected mark indicates that content may have been processed by Claude, but does not establish its full provenance. The company expressly identifies proofreading, translation, summarization, and file conversion as workflows in which marked output may contain ideas, text, or data originating elsewhere.

Before a detector result is used in a consequential decision, an institution should be able to answer at least the following questions:

  • Which models, versions, languages, and output surfaces does the detector cover?
  • What minimum quantity of text is required for a reliable result?
  • How do editing, translation, quotation, formatting, and mixed human-AI text affect detection?
  • What are the measured false-positive and false-negative rates under relevant conditions?
  • What threshold produced the result, and what does that threshold actually support?
  • Which detector version was used, and were its output and settings preserved?
  • Has the method been independently evaluated against human writing and competing systems?

These questions are not objections to watermarking. They define the boundary of the evidence.

A detector may be highly reliable at identifying a signal and still be incapable of determining the role AI played in the work. Technical accuracy does not expand the proposition being tested.

The wording of the result therefore matters. A statement such as "the document contains a signal consistent with processing by a marked AI system" stays closer to the evidence than "the document was AI-authored."

The second statement contains an attribution the detector did not make.

Institutions can promote the artifact into a finding

Suppose a university prohibits students from submitting work substantially generated by AI without disclosure. A paper produces a positive watermark result. The institution then tells the student that the detector proves the paper was AI-generated.

Several questions remain unanswered.

What conduct did the policy prohibit? Did it distinguish generation from editing, translation, accessibility support, or permitted assistance? Was the student given notice of that distinction? What does the detector establish about this particular workflow? What competing explanations were examined? What additional records support the conclusion?

Without those steps, the institution has not merely detected a mark. It has promoted an artifact into a finding.

The same problem can arise when an employer alleges policy breach, a publisher rejects an authorship claim, or an investigator treats marked text as proof that a person attempted deception.

The possible harm does not come from the watermark alone. It comes from the authority assigned to it inside a decision process.

What evidence can close the gap?

If the degree and nature of AI assistance matter, the inquiry should reconstruct how the work developed.

Relevant evidence may include:

  • original drafts and document version histories;
  • research notes and source records;
  • prompts and AI account activity;
  • timestamps and file metadata;
  • tracked editorial changes;
  • testimony or explanation from the author;
  • the applicable policy, including permitted and prohibited uses; and
  • the detector output, version, settings, confidence measure, and known limitations.

No single record will always resolve authorship. Together, these records can distinguish a short prompt followed by wholesale generation from a human-developed work that received limited editorial assistance.

They also allow the institution to test alternatives rather than treating the detector's label as self-executing.

This article is a worked example

I wrote the original draft of this article. The argument, examples, structure, and forensic reasoning were mine. I decided what I wanted to say and what conclusions I was prepared to defend.

I then gave the draft to Claude and asked it to improve the flow, correct the grammar, tighten the sentences, and preserve my voice. I reviewed the result, rejected or changed wording where necessary, and remained accountable for the final text.

If this version contains a watermark, the mark may support a conclusion that Claude participated in producing some of its language. That would be useful provenance information.

It would not establish that Claude conceived the argument, performed the research, selected the evidence, exercised final editorial control, or authored the work.

The challenge does not invalidate the signal. It reveals the boundary of what the signal establishes.

A defensible governance rule

Institutions do not need to ignore watermark evidence. They need to stop it at the correct evidentiary boundary.

A workable control would be:

Watermark detection must not be treated as proof of AI authorship, misconduct, or policy breach. It may initiate a review. Any consequential determination requires a documented interpretation of the signal, disclosure of the detector's material limitations, examination of relevant alternatives, corroborating process evidence, and accountable human judgment against a clearly stated rule.

This approach preserves the value of provenance without turning it into automatic attribution.

It also separates three questions that institutions too often collapse:

  • Did an AI system interact with the work?
  • What role did the system play?
  • What conclusion does the evidence justify under the applicable rule?

A watermark may help answer the first question. It cannot answer all three.

The boundary matters

Machine-readable marking can improve transparency. It can help platforms trace synthetic content, support investigations, and identify material that warrants closer review.

Its value depends on keeping its meaning precise.

A watermark is evidence about technological provenance. It is not a substitute for reconstructing authorship. It is not a policy. It is not a finding of misconduct. And it should never become one merely because its output is easy to read and difficult to challenge.

The question is not whether AI touched the work.

It is what the evidence proves about that involvement.

Sources