Anthropic's new watermark does not solve the AI detection problem in schools
Before generative AI arrived, law schools had assessment methods they trusted, or trusted enough. Take-home exams, seminar papers, and written assignments were never perfectly secure. A student could always consult a friend who had taken the course or possibly a practitioner. But the risk of getting caught was real, the pool of people willing to help was small, and the odds that the helper was actually good enough to move a grade were smaller still. The system sort of worked not because cheating was impossible but because effective cheating was rare. Plus, these sorts of assignments often more accurately captured the skills that practicing lawyers would employ.
AI ended that happy equilibrium. Every student now has access to an assistant that never sleeps, never rats you out, got a perfect score on the bar exam, has all the tools and connectors described in this blog, and writes at a level that puts a very high floor under anyone's attainment. The time-tested methods of separating students by ability suddenly do not separate much of anything. And methods that no longer separate can induce less effort to learn. It is understandable that many law professors find this annoying.
That is why Anthropic's recent announcement looks so intriguing. The company says that new Claude models will embed an invisible watermark into the text they generate, a move tied to the EU AI Act's Article 50 transparency rules that took effect August 2, 2026. Read quickly, it sounds like the restoration project has begun: the machines will now mark their own work, and the take-home exam can come back from the dead. It is worth understanding how the technology probably works because the details point the other way.
Note: Anthropic has not disclosed the underlying algorithm, but the mechanism almost certainly resembles the approach Kirchenbauer and colleagues published in 2023, since it remains the standard technique in the field.

To see how the watermark works, start with how a language model writes at all. At each step, the model does not "know" the next word. Based on its elaborate training process, the model computes a score for every word in its vocabulary, tens of thousands of candidates, reflecting how well each would continue the text, and then picks among the high scorers with a controlled dose of randomness. The crucial fact is that at most positions, many words score nearly the same. After "The court held that the statute was," the words "unconstitutional," "invalid," "preempted," and "void" might all be live candidates with similar scores. The probability distribution of words often has what information theorists call high entropy. Ordinary generation just picks one. That surplus of nearly interchangeable choices is the slack the watermark exploits: the model can be steered among words it would plausibly have chosen anyway, without the text reading any differently.
The steering works like this. Just before the model picks each word, the scheme takes the word(s) it generated immediately before and feeds it through a fixed mathematical recipe, what computer scientists call a hash function. (The same idea is used in Bitcoin.) All that matters here is two properties: the recipe turns any word into a number in a way that looks random but is perfectly repeatable, and the same input always yields the same output. That number is then used to deal the entire vocabulary into two piles, a "green" pile and a "red" pile, roughly half and half. Because the deal depends on the preceding word, the piles are reshuffled at every position: "unconstitutional" might be green after the word "was" but red after the word "deemed." The model then gets a small thumb on the scale in favor of green-pile words before it makes its pick. Isn’t that ingenious!
No single choice betrays anything. Picking "invalid" over "void" is exactly what a human might do. But a human writer, knowing nothing about the piles, lands on green words about half the time by pure chance. Watermarked text lands on green words markedly more often, and because the deal is repeatable, anyone who knows the recipe can walk back through a finished document, reconstruct which words were green at each position, and count. The count either looks like a coin flip or it doesn't, and the difference between those two outcomes is measurable with the same kind of statistics used to test whether a coin is loaded. Those knowledgeable in statistics can just feel the word “p-value” forming in their throat.
Here is the detail that changes the analysis: there are two ways to deploy this. In "public mode," the seeding rule that generates the green/red split is simply published, and anyone can reconstruct it and run the detector themselves, no secret required. In "private mode," the split is generated using a secret key, so the green list at any position is cryptographically unpredictable to anyone who doesn't hold that key. Watermarked text is then statistically indistinguishable from ordinary writing to everyone except the key-holder, and detection only happens if that party runs it for you, typically through a hosted detection service, probably at a price.

Anthropic has not said which mode it is using, but a private-key model with a hosted detector is the more likely choice for a company protecting its own verification service, and it is consistent with what little Anthropic has disclosed publicly: that the mark travels through copying and pasting and may survive some editing, without any detail on how a third party could check it independently. If that is the design, a teacher cannot verify a paper's provenance on their own. They would need to submit it to Anthropic and wait for a probabilistic answer, not run a detector themselves.
That limitation alone is significant, but it is not the largest one. The technique can only detect text produced by Claude, using Claude's own key. A student who uses ChatGPT, Gemini, or any other tool leaves no green-list bias in the writing at all, because Claude never touched it. A negative result from Anthropic's detector would be accurate, that particular text did not come from Claude, but a teacher unfamiliar with the mechanism might easily misread a negative as proof the student wrote the essay unaided. For watermarking to function as general AI detection across a classroom, every major provider would need compatible watermarking, and someone would need infrastructure to check submitted text against all of them. Neither exists today, and there is no indication of coordination toward it. What we are seeing is Anthropic ticking off a checkbox for Brussels, not building a classroom integrity tool. Treating it as the latter is a very bad mistake.

Even confined to Claude's own scope, the watermark is fragile in ways that matter for enforcement. Detection depends on accumulating enough word choices for the statistical signal to clear a threshold, so short passages, a paragraph or two, often will not carry a reliable signature. It depends on the bias surviving the final text, and any paraphrasing, by hand or by a second AI pass, word choices off the original green list and degrades the signal. None of this requires sophistication to defeat. An ordinary editing pass, the kind any writer does regardless of whether AI was involved, would probably be enough. I can see the Reddit posts now on how to evade the Claude watermarking technology. Maybe even an app!
There is a second implication worth naming, less about detection and more about business model. A private-key watermark that only the provider can check is, in effect, a verification service Anthropic controls and could charge for or restrict access to. Substack has already partnered with a third-party detection company for exactly this purpose. It would not be surprising to see Anthropic, or its competitors, build out paid or gated verification APIs rather than open detectors, which would deepen rather than close the gap between what a school actually needs, an independent check anyone can run, and what the market is likely to offer, a proprietary judgment call from whichever company happens to hold the key.
The practical result for schools is a technology that will produce both false positives and false negatives at rates too high to justify disciplinary use. A false accusation, based on a probabilistic signal from a company that has not disclosed how its detector works or how reliable it is at the passage lengths students actually submit, is not something a school should stake a student's standing on. A false negative, a paper that used AI extensively but escaped detection because the student used a different tool or lightly edited the output, will be common enough that any confidence built on clean results is misplaced from the outset.

The right conclusion is not to wait for better detection technology. It is that reliable detection of AI-assisted writing is not currently available, and may not become available given how easily the signal degrades. Institutions should design their teaching, assessment, and integrity policies as though detection will never arrive. That means shifting toward assessment less dependent on take-home written products a machine can produce invisibly: in-class writing, oral examination, drafting with visible revision history, or assignments that ask students to engage with AI tools openly rather than pretending (looking at you, Berkeley) the tools do not exist. You can see policies emerging at schools like the University of Texas and Columbia that are far more sensible. It also means being honest with students and colleagues about what these tools can and cannot do, rather than letting headlines about invisible watermarks create false confidence in an enforcement mechanism that does not work as advertised.