Breaking News

Anthropic Reveals What The Watermark Is And How It Can Be Defeated via @sejournal, @martinibuster

All Paid Media PPC News Social MediaAdvertising Video Advertising Columns Ask A PPC ExpertNEW PPC Pulse Webinar Google Local Services Ads Are Moving To PMax: What To Check First Prepare for Google's LSA move into Performance Max with a before-and-after audit checklist from CallRail. Digital Marketing All Digital Marketing Analytics & Data Ecommerce Lead Generation Content Marketing Social Platforms Google YouTube Reddit LinkedIn TikTokNEW WordPress Other CMS Platforms Webinar Google Local Services Ads Are Moving To PMax: What To Check First Prepare for Google's LSA move into Performance Max with a before-and-after audit checklist from CallRail. SEJ Live Webinars Rundowns PodcastNEW Ebooks All Resources SEJ Pro SEJ Pro: compare AI search notes in private Ask the SEJ experts you already read what AI, agents, and algorithm changes mean for your traffic. Private, human to human. $97/mo. Anthropic provides more details about its text watermark and reveals how it can be defeated. SEJ STAFF Roger Montti 6 hours ago ⋅ 6 min read SEJ STAFF Roger Montti Owner - Martinibuster.com at Martinibuster.com Bio Follow Anthropic announced how its watermark works, confirming virtually all of the details previously reported about a similar watermarking method called MirrorMark. Similar to MirrorMark, the watermark is the randomness pattern itself which mirrors the randomness of the LLM when it generates text. Contrary to what some AI influencers say, there are no Unicode characters that are embedded into the text. So it’s not something that you can copy and paste into a text file to remove or to identify. Also, it’s not about em dash use and neither is it about patterns that LLMs tend to use, like “It’s not this, it’s that” style of writing. It’s not looking for the likelihood that something was written by an AI. What it’s looking for is a specific watermark pattern. LLMs generate the next likely text in a sequence but with randomness built in. It doesn’t always pick the most likely next word; there is an element of randomness to the word that’s chosen. SynthID uses that randomness to set a pattern that’s dictated by a watermark key plus the context of preceding words. Because a SynthID-style watermark subtly alters the word choice randomness, the text that’s generated is indistinguishable from regular generated text. Users can’t identify the watermark without the watermark key. “That pattern is undetectable to the reader, but is detectable to anyone who has a key that encodes it. When watermarking is used, choices are still made at random, but the source of the randomness is different. Instead of using an arbitrary random number generator to pick the next word, watermarking uses the key and a few words that come before to settle what word the model should pick. That is, the words that Claude picks are still random, but now, one can check the sequence of words and see if it’s consistent with the choices Claude would make if it was using the key.” The announcement said that the new watermark is a version of SynthID-Text which was developed by Google DeepMind in 2024. It’s not SynthID, it’s a version of it. The state of the art for this kind of watermarking has significantly improved in the intervening two years. “Claude’s text watermark is a version of the SynthID-Text approach published by Google DeepMind in a Nature paper in 2024. It belongs to a family of approaches that go back to a proposal by Scott Aaronson in 2022, all of which share the same design principle that we described above—the watermark only changes the source of the randomness used to pick among words.” Yes, it can be defeated through paraphrasing. According to Anthropic, light editing probably won’t defeat it. “Can’t someone just edit the text to get around the watermarking? To some extent, yes. Light editing probably won’t remove the watermark completely; a complete rewrite where every word is replaced will. In the latter case, of course, it’s arguable whether the text can any longer be described as AI-generated.” SynthID looks for the watermark word pattern that was inserted at the time the text was generated. So if you paraphrase or edit enough of the document it’s going to erase the words that act as a watermark. SynthID was developed in 2024 and the state of the art has moved on over the past two years. A recent version of SynthID, called MirrorMark, extends SynthID by spreading the watermark across the generated text and using the surrounding words as context for determining where each part is placed, which makes it more resistant to editing. SynthID is a zero-bit watermark, which means it’s detecting watermark or no watermark. MirrorMark can encode multiple bits of information, essentially spreading the watermark across the generated text. I am not saying that MirrorMark is what Anthropic is using. But I am saying that before you put all your eggs into the SynthID basket, which is two years old, it may be useful to see what a 2026 version of SynthID can do. Read Full Bio SEJ STAFF Roger Montti Owner - Martinibuster.com at Martinibuster.com I have 25 years hands-on experience in SEO, evolving along with the search engines by keeping up with the latest ... Learn how to connect search, AI, and PPC into one unstoppable strategy. WP Engine Vs Automattic: Judge Inclined To Grant Preliminary Injunction Labeled: A New Wave Of AI Content Labeling Efforts Agentic AI In SEO: AI Agents & Workflows For Ideation (Part 1) Join 75,000+ Digital Leaders. Learn how to connect search, AI, and PPC into one unstoppable strategy. Learn how to connect search, AI, and PPC into one unstoppable strategy. In a world ruled by algorithms, SEJ brings timely, relevant information for SEOs, marketers, and entrepreneurs to optimize and grow their businesses -- and careers.


Source: Search Engine Journal

This article has been carefully curated and reformatted for educational and informational purposes. Full credit goes to the original publisher.


📚 Visit more helpful articles on Joab Peters Blog

No comments