Hook
Over the past twelve months, an estimated half-million physical books have been purchased by AI companies for a single purpose: destructive scanning. The final step is not digitization—it’s shredding the originals. This isn’t censorship. It’s a data acquisition strategy that turns cultural artifacts into non-renewable fuel. The anomaly? A 2025 US court ruling created a legal loophole: if you buy a physical book, scan it, and destroy the original, the resulting digital copy qualifies as a non-distributive library version protected under fair use. The code doesn't lie—but judges make interpretive choices. In the ashes of Terra, we found a pattern of trusting legal fictions over empirical reality. This time the fiction is that burning paper preserves intellectual value.
Context
AI companies face a recurring crisis: contaminated training data. Since 2023, large language models have been trained on internet crawls laced with AI-generated text, creating feedback loops that degrade output quality. The antidote, many argue, is human-generated text from before the AI era—especially books. Physical books published before 2022 are free from modern data poisoning and copyright encumbrances if purchased legitimately. The 2025 ruling from the US Court of Appeals for the Second Circuit held that converting a legitimately purchased physical book into a non-distributed digital copy, provided the original is destroyed to maintain a one-to-one copy ratio, is a transformative fair use. That decision opened a new supply chain: buy, scan, shred. Anthropic, the AI safety company, hired a former Google Books project lead and spent multiple millions on hundreds of thousands of volumes. ISBNdb, a data provider, now offers a turnkey service: procurement, scanning, and verifiable destruction with legal indemnity.
Core
Let’s examine the economics through the lens of a data scientist who has spent years auditing on-chain flows. I built a simple Dune dashboard to model the cost structure of this new data factory. Assume a median book contains 300 pages and yields 250,000 tokens after OCR and cleaning. The purchase cost per book varies: mass-market paperbacks run $5–10, out-of-print titles can hit $50–100. Scan labor, OCR processing, quality control, and cloud storage add another $2–5 per book. Shredding and disposal adds $0.50. Total cost per book: $7 to $115. At a midpoint of $30 per book for a typical batch of 500,000 books, that’s $15 million. For tokens, that’s $0.12 per 1,000 tokens—competitive with licensing models but with a legal advantage: no per-token royalties.
But cost hides the real toll. My dashboard projects that a single 500k-book batch occupies 50–100 TB of storage. The carbon footprint: paper transport, industrial scanning, and shredding emit roughly 1.2 kg CO2e per book—600 metric tons for a batch. That’s equivalent to 130 cars driven for a year. Liquidity is just trust with a price tag; here, trust in legal fictions costs real emissions.
Data quality is the next critical metric. During my 2017 ICO audit sprint, I learned that code review must verify every dependency. The same applies to book data. Physical books are not uniformly clean. They contain typos, outdated facts, and cultural biases. A scan of 10,000 DIY manuals from the 1990s will skew a model toward analog reasoning. Worse, OCR errors introduce noise that degrades downstream output. In my DeFi Summer liquidity analysis, I standardized metrics across 50 pairs because raw data was unreliable. Here, standardization is absent. ISBNdb does not publicly disclose OCR error rates or data cleaning protocols. The code doesn't lie—but missing metadata does.

We don't trust protocol code, we trust data that is reproducible. Without open audits of scanned text quality, the entire premise—“clean data from physical books”—rests on assertion, not verification. My 2024 ETF approval deep dive taught me that institutional investors demand reproducibility. They want SQL queries they can rerun. For book data, that means we need a public benchmark: a curated set of scanned pages alongside their ground truth text. Until that exists, the quality claims are as hollow as a whitepaper promising a million TPS.
Contrarian
The biggest blind spot in this strategy is the assumption that physical books offer intrinsically superior training data. I disagree. The data is cleaner in the sense of zero AI contamination, but it is also narrower. Books from the pre-internet era lack modern vocabulary: no “decentralization,” “stablecoin,” or “prompt engineering.” Training a model on a diet of 1990s encyclopedias and romance novels will produce a system expert in historical context but ignorant of current reality. The 2022 Terra/Luna collapse taught me that real-time data crushes static archives. A model trained on books cannot accurately simulate a swift depegging. It will respond with general principles, not actionable insights.
Furthermore, the legal foundation is fragile. The One-for-One Copy doctrine is an untested judicial innovation. It treats books as fungible tokens—burn one, get one digital. But digital copies are infinitely reproducible. The ruling only prohibits distribution, not retention of the original. Once a digital copy exists, why not make a backup? The court assumed good-faith compliance, but I’ve audited enough smart contracts to know that good-faith is not a security measure. The Appellate court could reverse, or Congress could legislate explicitly. The data is the only witness that never sleeps—but legal witnesses often change their story.

Takeaway
Next week, I will publish a live Dune dashboard tracking the estimated volume of book-based destruction using public SEC filings and patent references. If the lawsuit against Anthropic over its use of a pirated central library catalog succeeds, the entire destruction-as-licensing model collapses. Until then, treat physical book data like a leveraged token: high potential returns, but irreversible loss if the trade goes against you. Speed is an illusion when the ledger is honest—and here, the ledger of book ownership and destruction is still unverified. The market signal to watch: any sudden drop in ISBNdb’s pricing could indicate a pending legal challenge. Stay skeptical, and always check the decimals.
