Artificial Intelligence

AI Companies Are Destroying Books. That's Not Nearly as Scary as It Sounds.

Destructively scanning books is making information more accessible to the masses.

|


Are AI companies acquiring books, scanning them for training data, and destroying the physical copies? Yes, but the reality is much less dire than popular coverage suggests. 

Last week, the Demand Progress Education Fund sent a coalition letter to the Federal Trade Commission (FTC) asking the agency to investigate whether such bulk book digitization, "including [of] rare, out-of-print titles with few surviving copies," is anticompetitive behavior that illegally protects AI companies from competition. The focal point of the letter is Project Panama, Anthropic's self-described "effort to destructively scan all the books in the world." 

A destructive scan, one in which a physical book is destroyed in the process of being digitized, sounds sinister, but this method is a highly efficient way to convert books from bits to bytes. It also helps ensure compliance with copyright law. As federal Judge William Alsup found in Bartz v. Anthropic (2025), destructive scanning to be fair use because the one-to-one transformation from one form (physical) to another (digital), while neither multiplying nor redistributing copies, "was even more clearly transformative than those [fair use cases] (where the number of copies went up by at least one)." 

The coalition letter tries to sidestep this issue by noting that "a fair-use determination is…not a license to foreclose competition." That's true, but it does little to support their argument that AI companies are violating antitrust laws.

The coalition claims that "pre-AI human-authored text is a finite resource with an inelastic supply," but in many cases publishers are ready, willing, and able to publish more copies. In other words, supply is elastic. In cases concerning truly out-of-print works (inelastic supply), the existing supply is often abundant. For example, even though the Encyclopaedia Britannica 15th edition has been out of print for over a decade, a complete, 32-volume set can be readily found and purchased for a couple hundred dollars online. 

In the case of materials for which there remains only one copy—that is, supply is low and inelastic—there's no reason to believe these inputs are so essential to AI model development that their digitization would result in "an insurmountable systemic moat around AI incumbents." The market for AI development remains wide open because other training data abounds. For instance, Project Gutenberg hosts and freely distributes over 75,000 digitized books. Moreover, data derived from physical books isn't the only kind used to train AI models; developers also use data from synthetic data, publicly available data, licensed data, and user-generated data, among other sources.

The coalition's theory of harm is unable to satisfy the legal standards of the sort of predatory overbuying claim it presents. Per the Supreme Court's opinion in Weyerhaeuser Co. v. Ross-Simmons Hardwood Lumber Co. (2007), such a claim involves "bidding up input prices through the exercise of monopsony power." But there's no evidence that prices in the out-of-print book market have been generally bid up to the point that only incumbent AI companies can afford them. 

Furthermore, bulk book digitization is accompanied by substantial procompetitive effects, including benefits to consumers from improved AI models, many of which are available for free. Insofar as copyright law compels an AI company to unintentionally destroy the last copy of a 50-something-year-old book, the information contained within becomes far more useful and accessible to mankind by being incorporated into an AI model rather than remaining in its physical copy and in the possession of a single person.

Booksellers also benefit: One told 404Media that AI companies' bulk book purchases benefit him "financially…by clearing out old inventory that is otherwise unlikely to sell." In fact, this seller also revealed that the bulk purchases "were of books [that] all had ISBNs," an identification system adopted in 1970. This implies that no antiquarian works are being targeted for digitization and destruction. (This should come as no surprise because such works, when not in the custody of a university or museum, are purchased at auctions, not as part of a bulk order. Moreover, there's no obvious reason for an AI company to train its model on an original Shakespeare folio, which costs millions of dollars, when the text can be accessed for free.)

AI companies digitizing one copy of the plenitude of out-of-print works does not threaten humanity's cultural inheritance. Moreover, the FTC should not waste its limited resources investigating a practice that evinces fierce, dynamic competition between AI developers, not market foreclosure.