Plagiarism

Every generation of cheap content invents its own watermark. Every watermark eventually gets a workaround. Here’s the twenty-year pattern hiding inside this week’s AI news.

In this post:

I read a thread this week from the CTO of GPTZero explaining how Anthropic, Google, and OpenAI actually watermark AI text. Not the marketing version. The real mechanism.

I read it twice. Once for the engineering. Once because I’d already seen this exact movie three times before, just with a different cast.

The trick behind the watermark

Before the model generates your next word, it uses a secret key to split the vocabulary into a “green” list and a “red” list. Then it quietly tilts the odds toward green.

Do that token after token, and you get a statistical fingerprint. Invisible to a human eye. Obvious to anyone holding the key.

“Selecting a randomized set of ‘green’ tokens before a word is generated, and then softly promoting use of green tokens during sampling” — with “negligible impact on text quality.”
— Kirchenbauer, Geiping, Wen, Katz, Miers & Goldstein, “A Watermark for Large Language Models,” arXiv

Elegant. Quietly terrifying if you think about it for more than ten seconds.

Every generation of cheap content invents its own watermark, and every generation of watermark eventually gets a workaround.

I’ve watched that loop run on repeat for twenty years — first building growth engines inside SaaS companies, now writing checks for the founders building the next ones.

Four eras, one pattern

Lay the last two decades side by side and the shape is unmistakable. A shortcut shows up, it works for a while, a detector arrives, and someone pays the price.

The earliest documented reference to this whole family of tricks — “spamdexing” — showed up in print in 1996. This is not a new fight. It’s the same fight with better math each time.

The content farm I watched up close

Demand Media is the cleanest case study I know. By 2008, the company had produced roughly 340,000 articles and 135,000 videos. A year later it was publishing close to a million items a month — about four English-language Wikipedias’ worth of content, annually.

Weeks after that IPO, Google shipped Panda specifically to catch this pattern. Amit Singhal and Matt Cutts wrote in the official announcement that it was designed “to reduce rankings for low-quality sites — sites which are low-value add for users, copy content from other websites or sites that are just not very useful.”

Cutts came back to this years later with a line I think about often:

“With Panda, Google took a big enough revenue hit via some partners that Google actually needed to disclose Panda as a material impact on an earnings call. But I believe it was the right decision to launch Panda, both for the long-term trust of our users and for a better ecosystem for publishers.”
— Matt Cutts, former head of Google Webspam, via Wikipedia: Google Panda

Read that twice. The company building the detector took a hit shipping the correction — and shipped it anyway. That tells you how little the industry’s mispriced advantage of volume over quality was ever worth long-term.

Demand Media spent the next decade rebranding. It was eventually sold off to Graham Holdings for a fraction of its early promise. A year after Panda, Google Penguin did the same thing to link farms. Different lever, identical arc.

We’re inside the same loop right now

I already wrote about the flood of AI slop hitting outbound and content marketing this year. NewsGuard’s tracking makes the scale concrete.

“Strong evidence that the content is being published without significant human oversight” — and “top brands are unintentionally supporting these sites” through programmatic ad spend.
— NewsGuard, AI Tracking Center

Watermarking is this decade’s Panda. A channel got flooded with content optimized for a machine’s approval instead of a person’s actual need, and the platform sitting on top of that channel has every incentive to find it and discount it.

The GPTZero thread even gets into the counter-move: use a statistical model instead of a fixed hash, and the watermark survives paraphrasing better. That’s the modern version of the SEO cat-and-mouse game I ran in 2012 with keyword density and backlink profiles. Better math, same trap.

Why smart people keep walking into it

None of this happens because marketers are careless. It happens because the shortcut genuinely works for whoever takes it first.

Demand Media scaling to 340,000 articles before Panda existed wasn’t a mistake. For that window, it was a real advantage. The problem is the advantage is borrowed against a correction that’s already being built — usually by people with more resources than you, and a direct incentive to close the gap.

This is a structural problem, not a superficial one, which is exactly why polishing the tactic never fixes it. A year spent getting incrementally better at producing volume evaporates the moment detection catches up, because you were optimizing the wrong variable the whole time.

Quality’s payoff is also slower to show up on a dashboard. Nobody gets applauded in week one for the article that took real research instead of a prompt. Volume is visible immediately. Trust compounds quietly and shows up a year later — a bad trade if you’re measured quarterly, a great trade if you’re actually building something.

What survives every version of this

Here’s the reassuring part: the content that survived Panda, survived Penguin, and will survive watermark detection is the same content every time. The stuff that was never built to game the signal in the first place.

“People-first content means content that’s created primarily for people, and not to manipulate search engine rankings.” Of the four factors in Google’s E-E-A-T framework, “trust is most important.”
— Google Search Central, Creating helpful, reliable, people-first content

Watermark detection is answering a narrower version of the exact same question: was a real person actually standing behind this. If the honest answer is yes, you were never the target — no matter how many times the enforcement layer gets smarter.

I said this about the AI-native world generally in Double Down on Humans, and it applies here without edits: when the cost of digital creation drops to zero, the value of human connection scales to infinity.

The tools for cutting corners keep getting better. So does the effort spent catching them. The only strategy that’s never once had to be rebuilt from scratch, across four decades of this exact same game, is the boring one.

Build for the person reading it. Not for the green tokens.