Project Panama: The AI Company That Destroyed Millions of Books (and What It Says About Us)
What was Anthropic's Project Panama?
It was an internal Anthropic program, revealed in court documents, to buy millions of used physical books, scan them destructively (cutting off the spines) and use the text to train the Claude AI. The goal was to digitize between 500 thousand and 2 million books in six months, recycling the paper after scanning.
Picture an industrial warehouse where trucks unload books by the millions. Hydraulic machines slice off the spines, high-speed scanners photograph every page, and the leftover paper goes straight to recycling. No ceremony, no physical archive, no way back.
This is not dystopian fiction. It is "Project Panama", an internal program at Anthropic (the company behind Claude, one of the world's most advanced AIs) revealed in court documents and dissected by the international press in recent weeks. The goal was ambitious: digitize between 500 thousand and 2 million books in six months, building a central library of "all the books in the world" to keep "forever". In digital form, of course. The paper became scrap.
Almost everyone's first reaction is a knot in the stomach: destroying books evokes some of history's worst images. But this story deserves more than instant outrage. It is a lesson in what AI truly values, what the law allows, and choices that will shape how human knowledge circulates from now on.
What happened, in facts
- The bulk purchase: Anthropic spent tens of millions of dollars buying used physical books in large lots, hiring specialized companies for industrial scanning.
- The right team: in February 2024, it hired Tom Turvey, a former executive of the Google Books project, precisely the person who spent years digitizing books for Google.
- The destructive process: cutting the spine and scanning loose pages is far faster and cheaper than the delicate scanning that preserves the book. Destruction was not an accident; it was the method.
- The dark side: before buying books legally, the company had downloaded millions of works from pirate libraries like LibGen. That became a lawsuit.
- The bill arrived: in June 2025, Judge William Alsup ruled that training AI on legally purchased books is fair use, but keeping a pirate library is not. The result: a $1.5 billion settlement with authors, the largest in copyright history, with final approval in July 2026.
The question that matters: is this good or bad?
It depends on which value you place on the scale. Let us weigh both sides honestly.
The bad side: what is lost when paper becomes scrap
The symbolism is terrible, and symbols matter. A civilization that normalizes shredding books by the millions, even legally, is saying something about its priorities. A physical book carries reading marks, specific editions, material history. None of that survives the scanner.
The knowledge was locked up, not set free. Here lies the crucial difference from Google Books, the great digitization project of the 2000s: Google scanned while preserving the copies (many from partner libraries) and made excerpts publicly available. Project Panama digitized for one company's exclusive commercial use. The "library of all the books in the world" exists, but you cannot walk in.
The precedent is worrying. If every AI company decides to build its private library by buying up and destroying the world's stock of used books, the cumulative effect on used bookstores, community libraries and cheap access to reading is not trivial.
The good side: what this story reveals that is positive
Books beat the internet. Think about what it means for a frontier tech company to spend fortunes on paper: after training models on much of the internet, the conclusion was that edited, reviewed, deep text (in other words, books) is the most valuable fuel there is. In an era of infinite shallow content, the market just put the highest price on the most careful form of human knowledge.
No rare works were lost (as far as we know). Independent fact-checks indicate the bulk were used copies of common print runs, with millions of sibling copies alive in libraries and on shelves around the world. The texts survive; what died were objects, not works.
Authors won a billion-dollar precedent. The $1.5 billion settlement carved in stone a distinction every creator cares about: buying and using legally is allowed; pirating costs dearly. That pushes the AI industry toward licensing and compensation, not free-for-all.
And the AI you use got better because of it. When Claude (or any model trained on books) explains a concept to you with depth and nuance, part of that comes from here. There is an honest argument that turning books into reasoning capability accessible to millions of people is a way of multiplying knowledge's reach, not burying it.
Google Books vs Project Panama: same idea, two spirits
| Google Books (2000s) | Project Panama (2024+) | |
|---|---|---|
| Goal | Index and give public access to excerpts | Train a commercial model |
| Method | Scanning that preserves the copy | Destructive scanning |
| Book's fate | Returned to the library | Recycling |
| Access to the result | Partially public | Private |
The comparison stings because it shows another path was possible. More expensive and slower, yes. But possible.
What this story teaches (for those of us who are not Anthropic)
- Data provenance became a business issue. The question "where did the data that trained this AI come from?" moved from philosophy to law. Companies that use AI will live with it, and so will everyone who creates content.
- Well-crafted content just got more valuable. If books are the gold of training, your company's deep, well-written knowledge (documentation, research, teaching material) is an asset too, and deserves to be treated as one.
- Legal and right are not synonyms. Destroying books you bought is legal. The court said so. Whether it is what we want as a norm of civilization is a conversation the law will not have for us.
In the end, Project Panama is a mirror: the world's most advanced AI was built on the oldest technology we have, the book. Whether that strikes you as a tribute or a desecration probably says as much about you as about Anthropic. And maybe that is the best question to take away: what other shortcuts are we willing to accept in exchange for smarter machines?
Frequently asked questions
It was an internal Anthropic program, revealed in court documents, to buy millions of used physical books, scan them destructively (cutting off the spines) and use the text to train the Claude AI. The goal was to digitize between 500 thousand and 2 million books in six months, recycling the paper after scanning.
In the US, yes. In June 2025, Judge William Alsup ruled that legally buying books, digitizing them and training AI on them is protected by fair use. What is not legal is using pirated copies: that is why Anthropic settled with authors for $1.5 billion.
Because, before buying books legally, the company downloaded millions of works from pirate libraries like LibGen. The settlement, the largest in copyright history, finally approved in July 2026, covers those pirated works, not the destruction of purchased books.
According to independent fact-checks, the bulk were used copies of common editions, with many other copies existing worldwide. The works (the texts) were not lost; the physical objects were.
Because books are edited, reviewed, deep and well-structured text, far superior training material to the average of the internet. Models trained on books reason and write better, which is why companies pay fortunes for this kind of data.

Data and AI executive with 20+ years building technology that moves businesses. Microsoft Certified Trainer, with executive education at MIT Sloan. At Data Lover, he trains professionals and leads enterprise AI projects.
See profile and all articles →


