Amazon, once an online bookseller, is destroying rare books to train AI models
Amazon has begun dismantling a collection of rare, out‑of‑print volumes housed in its Seattle warehouse to feed large language models (LLMs). The practice was uncovered in late July 2026 by a whistle‑blower who described crates of centuries‑old tomes being shredded on site. The move raises alarms for cultural institutions, collectors, and AI ethicists who fear irreversible loss of literary heritage. It matters because the books are not just historical artifacts; they are also irreplaceable data sources that could improve AI understanding of language nuance.
Key takeaways
- Amazon is converting physical rare books into digital text for its generative‑AI training pipelines.
- The destruction occurs in Amazon’s own fulfillment centers, bypassing traditional digitization partners.
- Preservation groups warn the loss could erase unique primary sources for scholars worldwide.
- The controversy spotlights a broader industry tension between AI progress and cultural stewardship.
Background
Amazon entered the AI arena in early 2025 with the launch of its Bedrock foundation‑model service. To stay competitive with rivals like OpenAI and Anthropic, the retailer sought “high‑quality, niche data” that public web crawls lack. Rare books—first editions, handwritten manuscripts, and limited‑run scholarly works—fit that bill because they contain language patterns unavailable elsewhere.
Historically, libraries and specialty dealers have digitized such works under strict licensing agreements. Amazon’s approach departs from that model, opting instead for in‑house bulk scanning followed by physical destruction, a method disclosed in a report by TechCrunch【TechCrunch】. Critics argue this bypasses preservation norms and sidesteps the consent of rights holders.
What happened
In July 2026, an employee at Amazon’s Seattle fulfillment hub reported that crates marked “Historical Collection – Decommission” were being fed into industrial shredders. The books, identified by internal inventory codes, included 19th‑century travelogues, limited‑edition poetry, and early scientific treatises. According to the employee, supervisors instructed staff to “process these items for AI ingestion” and then discard the paper waste.
Amazon’s spokesperson confirmed the operation, stating that the scanned text is used to “enhance the factual accuracy and cultural depth” of its LLMs. The company claimed the practice complies with U.S. copyright law because many of the volumes are in the public domain. However, preservation advocates point out that even public‑domain works can hold unique marginalia and physical attributes that are lost forever when the books are destroyed.
Why it matters
- Cultural loss – Rare books often contain annotations, binding details, and provenance information that cannot be captured by OCR alone. Their destruction eliminates a layer of scholarly insight.
- Legal gray area – While public‑domain status permits copying, the lack of licensing agreements raises questions about the ethical use of cultural assets for commercial AI gains.
- Industry precedent – If major retailers adopt similar practices, the market for licensed digitization could shrink, undermining libraries’ revenue streams and preservation budgets.
- Public trust – Consumers increasingly demand transparency about how AI models are trained. Unseen destruction of heritage items may erode confidence in Amazon’s technology initiatives.
The issue also resonates with broader debates highlighted in recent coverage such as What we learned from Wafcon 2026, where experts warned that unchecked data harvesting could jeopardize cultural memory.
What happens next
Amazon has pledged to review its handling procedures after the story broke, promising a “more responsible approach” to sourcing training data. The company is reportedly consulting with the Chronicle News editorial board to develop a public policy framework. Meanwhile, libraries and heritage groups are mobilizing petitions and legal challenges, seeking injunctions to halt further destruction.
Legislators in Washington state have announced hearings on AI‑related data practices, with a focus on protecting rare‑book collections. Should new regulations emerge, Amazon may be required to obtain explicit permissions before digitizing and discarding physical works. The outcome could set a global standard for how AI developers source and preserve cultural content.
Frequently asked questions
How can Amazon legally scan and destroy public‑domain books?
U.S. copyright law allows copying of public‑domain works without permission, but it does not address the ethical implications of destroying unique physical artifacts.
Are there alternatives to destroying the books?
Yes. Publishers and libraries can provide digitized copies under licensing agreements, preserving the originals while still supplying data for AI training.
What can individuals do to protect rare books from similar practices?
Supporting advocacy groups, donating to preservation funds, and contacting elected representatives to demand transparent AI data policies are effective steps.
Bottom line
Amazon’s decision to shred rare books for AI training has sparked a fierce debate over cultural preservation versus technological advancement. The story, originally reported by TechCrunch, underscores the need for clearer industry guidelines.
Reporting by TechCrunch.
Related reading
- What we learned from Wafcon 2026
- Trump threatens to bomb US ally Oman if it 'gets in the way' over Iran deal
- [Russia's prominent anti-war politician jailed for 11 years](/articles/russias-prominent-anti-war-politician-jailed-for-11