AI

Why AI Companies Are Buying Books and Destroying Them for Model Training

SAN FRANCISCO — Artificial-intelligence companies are increasingly buying physical books in bulk, cutting off their spines, scanning their pages at industrial speed and discarding the original volumes in an effort to build vast libraries of text for training AI systems. 

The practice has drawn alarm from authors, booksellers, archivists and readers, particularly after an investigation found that shipments of rare books had reached an Amazon facility in Las Vegas where workers said they remove bindings and scan books before the physical copies are destroyed. 

Amazon confirmed to 404 Media that it purchases books through commercial channels “to improve the products and services customers use,” but did not directly describe its scanning procedures, identify the titles involved or specify the AI products that use the material. 

The emerging book-buying system reflects a new phase in the race to develop powerful language models. Companies need enormous amounts of high-quality text to train systems that can write, summarize, code, analyze and converse. Much of the easily available internet has already been harvested, is legally contested or is increasingly contaminated by AI-generated material. 

Books offer something that short web posts and social-media feeds often do not: long-form, human-authored, edited and coherent writing. That makes them valuable as training material. But the industrial method of converting physical books into data has raised a broader question: when companies turn books into private AI datasets, who benefits — and what is lost? 

How destructive scanning works 

The process is technically simple and commercially efficient. 

A company or contractor acquires printed books, often in large quantities through used-book marketplaces, wholesalers, estate sales or commercial sellers. Workers then cut off the books’ bindings, separate or loosen the pages and feed them through high-speed document scanners. Optical character recognition software converts the scanned images into machine-readable text. The paper copies are then recycled, shredded or otherwise discarded. 

Removing a book’s spine allows pages to move quickly through automatic scanning equipment. A bound volume can be scanned more slowly and carefully, but destructive scanning makes it possible to process large quantities of text at a lower cost. 

The method has long been used in some digitization projects, especially when libraries or archives work with duplicate, damaged or low-value copies. The difference in the AI context is scale and purpose. Instead of creating a public digital archive, companies can use the resulting text to train proprietary models whose internal datasets are generally not available to the public. 

404 Media reported that it placed a tracking device in a rare-book shipment it believed might be acquired by an AI company. The shipment ultimately arrived at an Amazon warehouse in Las Vegas. Employees at the location told the publication that their work involved receiving large shipments of books, cutting off bindings and scanning the pages. 

The report did not establish that every book processed at the facility was used to train AI models. But it provided rare visibility into an industry practice that has largely occurred outside public view. 

Why companies want physical books 

AI companies have several incentives to seek printed books. 

First, books are a source of high-quality human writing. A well-edited book can contain sustained arguments, narrative structure, subject-specific vocabulary, historical context, and stylistic variety. For large language models, which learn patterns from vast amounts of text, such material can be more valuable than fragmented or repetitive online content. 

Second, many books are not readily available online. Out-of-print titles, specialized academic works, regional histories, niche technical manuals and older editions may exist only in physical form. These texts can fill gaps in a company’s dataset. 

Third, physical books can be a way to obtain text without relying on pirated digital libraries. The use of unauthorized “shadow libraries” has exposed AI companies to major copyright lawsuits by authors, publishers and visual artists. Acquiring a physical copy through commercial channels may appear less legally risky than downloading an unauthorized digital edition. 

Fourth, printed books are generally written by people. As more online content is generated or heavily modified by AI, developers face concerns about training new models on synthetic material created by earlier models. Researchers sometimes describe this as a risk of “model collapse” or degraded training quality, though the severity depends on how data is selected and filtered. 

Books offer a comparatively stable source of human-authored text, created before the current explosion of generative AI. For companies competing to produce better language models, that makes large physical libraries a valuable resource. 

Amazon’s reported operation 

The Amazon case has drawn particular attention because the company began as an online bookseller before becoming one of the world’s dominant technology, cloud-computing and retail companies. 

According to 404 Media, Amazon’s Las Vegas facility received a shipment of rare books that included a hidden tracking device. The publication reported that books were sent to the facility for destructive scanning. 

Amazon told the publication that it buys books through commercial channels to improve customer products and services. The company did not say whether the material was used to train its Nova AI models, Kindle-related tools, search systems, or other products. 

The lack of detail is a source of concern for authors and booksellers. Companies can lawfully buy books in the marketplace, but that does not answer whether authors should be compensated when their works become raw material for commercial AI systems. 

The issue is especially sensitive when rare or antique copies are involved. A standard used paperback may be easily replaced. A scarce local history, first edition, annotated copy or specialized small-print-run title may not be. Even if a text survives elsewhere, destroying a physical artifact can erase value that exists beyond the words on the page. 

Books are objects as well as containers of text. Their bindings, illustrations, editions, ownership marks, marginal notes, and printing history can matter to scholars, collectors, and communities. A scanned text file cannot fully preserve those features. 

Anthropic’s “Project Panama” 

The practice became more widely known through litigation involving Anthropic, the developer of the Claude AI assistant. 

Court documents in a lawsuit brought by authors revealed that Anthropic had pursued a project code-named “Project Panama.” Under that initiative, the company bought millions of physical books, used contractors to cut off their bindings, scanned the pages, and disposed of the originals. 

Internal documents described an ambition to “destructively scan all the books in the world,” according to reporting on the case. One contractor proposal involved scanning between 500,000 and 2 million books over six months

The litigation drew a crucial legal distinction. A U.S. federal judge found that Anthropic’s training on lawfully acquired and destructively scanned physical books was transformative fair use under copyright law. The court did not extend the same protection to the company’s acquisition and storage of pirated digital books. 

Anthropic later reached a $1.5 billion settlement with authors and publishers over claims involving pirated works, according to reports on the litigation. The settlement did not eliminate the broader controversy over training AI systems on copyrighted books acquired through legal purchases. 

The case provides a template for other companies. Purchasing physical books, scanning them internally, and destroying the originals may reduce certain copyright risks compared with downloading pirated files. But it does not resolve ethical questions about consent, compensation, and the preservation of cultural material. 

The legal argument, and its limits 

Under the first-sale doctrine in U.S. copyright law, a person who legally buys a physical book generally has the right to sell, lend, give away or destroy that particular copy. Ownership of the object does not mean ownership of the copyright, but it does give the purchaser broad control over the physical item. 

Companies have argued that scanning legally purchased books for internal AI training is a transformative use because the AI model does not provide readers with a replacement copy of the original book. It learns statistical patterns from the text rather than distributing the work in its original form. 

Supporters of that position say the practice resembles established forms of text analysis, including search indexing, plagiarism detection, and research databases. They argue that blocking AI training on lawfully acquired books could limit innovation and place U.S. companies at a disadvantage. 

Critics see a legal loophole. They argue that companies are using ownership of individual physical copies to create massive digital datasets without licensing the underlying works or paying the people who wrote them. 

The distinction becomes sharper when the books are rare, out of print or unavailable in libraries. A company may lawfully buy and destroy one copy, but if many companies follow the same practice, scarce editions could become harder for researchers, readers and libraries to obtain. 

The cultural cost of turning books into data 

The most emotional objection to destructive scanning is not only about copyright. It is about stewardship. 

Booksellers have reported unusual bulk orders for older titles and niche books, leading some to suspect that AI companies are attempting to acquire titles systematically by ISBN number. 

The concern is that rare works may disappear from the used-book market and be converted into proprietary training material. A text that once could be bought, borrowed, studied or preserved may survive only as an unsearchable component inside a commercial model. 

That outcome is particularly troubling for historians, librarians, and archivists. A book’s physical form can carry evidence about its publication, circulation, and use. Marginal notes, inscriptions, library stamps, binding variations, and printing errors may have scholarly significance. 

Advocates for preservation argue that companies seeking large-scale text datasets should prioritize licensed digital collections, public-domain works, voluntary author agreements, non-destructive scanning and partnerships with libraries. They also argue that rare or culturally significant volumes should be assessed before being dismantled. 

Some digitization projects already use non-destructive scanners that photograph pages while preserving bindings. Those systems are slower and more expensive, but they protect the physical object. The decision to use destructive scanning is therefore not unavoidable; it is often a choice driven by speed and cost. 

A dispute about the future of knowledge 

The book-scanning boom illustrates a deeper conflict over who controls the raw material of artificial intelligence. 

AI companies say they need high-quality data to build tools that improve productivity, scientific research, education, and software development. They argue that books are part of the broad body of human knowledge from which new technology should learn. 

Authors, publishers and archivists counter that human knowledge is not simply free industrial input. Books are copyrighted works, cultural artifacts and the product of years of labor. Turning them into commercial AI data without permission or payment risks concentrating the value of that labor in a small number of technology companies. 

The dispute is unlikely to end soon. Courts will continue to define the boundaries of fair use, lawmakers may consider new rules for AI training data, and publishers will pursue licensing agreements where they see commercial opportunity. 

For now, the practice is expanding because it solves a practical problem for AI developers: where to find large amounts of human-created text in an internet increasingly shaped by AI itself. 

The result is a paradox of the digital age. Companies are buying physical books, objects once viewed as relics of an earlier information era, because those books may hold some of the most valuable data for building the next generation of machines. 

We Recommend

The yoopya.com portal presents worldwide news, covering a large spectrum of content categories including Entertainment, Politics, Sports, Health, Education, Science and Technology and more. Top local and global news in the best possible journalistic quality. We connect users via a free webmail service and innovative.
AI

Why AI Companies Are Buying Books and Destroying Them for Model Training

Reading time: 8 min

Discover more from Top Local & Global trusted News | Secure Email Account

Subscribe now to keep reading and get access to the full archive.

Continue reading