Imagine an 18th-century botanical treatise that has survived wars, fires, and three hundred years of history. Only three copies exist worldwide. One morning, an artificial intelligence company buys it through an anonymous service, runs it through a machine that slices off its binding for a quick scan, then shreds it. This isn’t dystopian fiction: it’s a very real industrial practice, brought to light in late July 2026 by several converging investigations. The destruction of rare books by AI raises a brutal question: how far can the race for training data go?
An industrial, anonymous modus operandi
The mechanism is as simple as it is shocking. AI companies go through ISBNdb, a service specializing in bulk book procurement. This platform lets you order up to a million books at a time, all while preserving buyer anonymity. Even better: it offers non-disclosure agreements (NDAs) as a commercial selling point and advises its clients to use the phrase “digital preservation” rather than destruction.
Once the books are delivered, they go through high-speed scanners. To speed up the process, the binding is literally sliced off so the pages can be scanned flat, with no manual handling. A bookseller interviewed by 404 Media describes this as a one-way trip: after scanning, the originals are shredded or pulped. All that remains is a digital file on a server.
The goal is clear: build massive training datasets without worrying about contamination from synthetic content.
Why pre-2022 books are worth gold
Since the explosion of large language models, a growing share of text available online is AI-generated. Blogs, articles, product descriptions, Wikipedia pages: algorithmic “slop” is everywhere. For companies training their models, this synthetic content is a major problem.
A model fed text generated by other AIs tends to degrade in performance, a phenomenon sometimes called “model collapse.” The workaround is to seek out sources predating 2022, the year synthetic content began flooding the web.
Old and out-of-print books fit the bill perfectly:
- They predate the era of mainstream LLMs.
- Their text has been edited, reviewed, and fact-checked, unlike the vast majority of web content.
- They are often unavailable in digital format, making them an exclusive resource.
The irony: these same companies that reject synthetic content for training their AIs are among the first to produce it at industrial scale. An article on Nonograph sums up this contradiction with a biting line: “If AI companies avoid slop, shouldn’t we do the same?”
A practice validated by the courts
Perhaps the most disturbing fact in this story is its legality. A US federal judge ruled that this practice falls under fair use. The reasoning: since only one copy exists at any given time (the digital file replacing the destroyed physical original), there is no unlawful duplication.
This decision opens a legal highway. A company can now acquire, digitize, and destroy books, including extremely rare editions, without fear of copyright infringement lawsuits, as long as the physical copy is eliminated.
Anthropic, the only company named in the controversy, has notably hired the former head of partnerships at Google Books with an explicit mandate: acquire “every book in the world.” An ambition that, paired with recent case law, suggests the phenomenon is about to accelerate.
What collective memory loses
Rare books with very few surviving copies have already gone through this pipeline, according to the bookseller’s testimony relayed by 404 Media. We’re not talking about dime-store novels or obsolete manuals, but potentially unique editions, old scientific texts, historical treatises.
A rare book is more than an information carrier. It’s a material object, a witness to its era, sometimes annotated, hand-bound, carrying its own history. Once shredded, it’s gone forever. You can reprint a bestseller or put a website back online. You can’t recreate the last three copies of an 18th-century botanical text.
The wording on ISBNdb’s own website is chilling in its unabashed cynicism: “‘An AI company destroys two million books’ is not a sympathy-inducing headline.” Yet they built an entire business model on exactly that, betting on discretion.
Key takeaways
- AI companies buy, scan, and destroy rare books through ISBNdb, a service that guarantees anonymity and offers NDAs.
- Pre-2022 works are targeted because they are free of AI-generated content, making them “clean” training data.
- A US federal judge validated this practice under fair use, setting a precedent with serious implications.
- Extremely rare editions, sometimes with only a handful of copies left, are disappearing irreversibly.
- No company, including Anthropic (the only one named), has officially commented on the matter to date.
This story deserves close attention. If you have an opinion on the issue or information to share, don’t hesitate to weigh in via the comments or contact me directly. And if topics at the intersection of tech, law, and ethics interest you, consider subscribing to the blog.
Sources
- HedgieMarkets on X, thread from July 27, 2026, initial source of the controversy.
- 404 Media, mention of the bookseller’s testimony on the destruction of rare editions.
- The Atlantic, coverage on AI companies recruiting professors (related context).
- Dallas Express, article “The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions” documenting the modus operandi.
- Yahoo News, article detailing the judicial validation of the practice via a federal court ruling.
