Burning the Past to Build the Future

3 hours ago 2

Rommie Analytics

Anthropic’s Project Panama is devouring old books for the human writing its technology is making increasingly hard to find.

At undisclosed locations, in nondescript warehouses, vast troves of print books are scheduled for destruction. Each one awaits its bindings hacked off so that its pages can be scanned into a central library, where they will be used to train artificial intelligence models to enhance their capabilities. The more books consigned to oblivion, the more effective these models become at generating outputs that not merely mimic human-like text and speech but become indistinguishable from the real thing.

Burning the past to build the future seems like a fiction inspired by Ray Bradbury’s Fahrenheit 451, but it is a real undertaking by the artificial intelligence giant, Anthropic, to feed its chatbot, Claude, with physical texts. It’s called Project Panama, and the company is acutely aware of how bad it looks, as shown in internal documents revealed by legal filings reviewed by The Washington Post. “Project Panama is our effort to destructively scan all the books in the world,” the documents state. “We don’t want it to be known that we are working on this.” 

This destructive scanning is a fairly common practice, but Project Panama is distinguished by its world-devouring scale. It’s hard to avoid the feeling that Anthropic, which has yet to be publicly traded but is estimated to have a value of close to $1 trillion, tried to keep it under wraps because it seems to provide more evidence that AI giants have little regard for the human value of creative work, apart from its utility in building a post-human future. This modern-day variation of physical book burning also comes at a time when public opposition to AI accelerationism is becoming fevered.

Despite the opposition of AI labs like Anthropic, Meta, and OpenAI, authors have sued the tech giants over allegations of copyright infringement. Anthropic recently notched a win when U.S. District Judge William Alsup ruled that training Claude on legally purchased books falls under fair use. According to Alsup, the act of destructively scanning texts into an internal archive for training AI models counts as “transformative.” Under Section 107 of the Copyright Act, transformative uses have a better chance of being considered fair use. “Like any reader aspiring to be a writer, Anthropic’s LLMs [large language models] trained upon works not to race ahead and replicate or supplant them—but to turn a hard corner and create something different,” Alsup concluded.

Alsup’s ruling is a significant victory but only a partial one. Project Panama was born after Anthropic, based in San Francisco, compiled a large collection of pirated books and grew wary of potential litigation that, sure enough, came. In late July, U.S. District Judge Araceli Martinez-Olguin signed off on a $1.5 billion settlement specifically over Anthropic’s use of pirated books, which a group of authors argued were used without their permission to train Claude. And while training AI models on purchased instead of pirated books would seem to be a more ethical approach, a review of the facts in the case shows that Anthropic consistently prioritized speed over other considerations, including the preservation of unique and rare texts. 

Between 2021 and 2022, Anthropic obtained a stunning seven million copies of pirated books through different sources, including Pirate Library Mirror, which acknowledged that it had to “deliberately violate the copyright law in most countries” in order to amass its enormous catalogue of books. Dario Amodei, Anthropic’s co-founder and chief executive officer, preferred this method to avoid the “legal/practice/business slog” of negotiating with publishers, legal filings show. Soon, however, the risks of training an AI model entirely on pirated books became too great. In fact, Meta, the parent company of Facebook with a market capitalization of $1.5 trillion, is being sued for the same thing by authors who allege that it scraped together an enormous collection of books to train its AI models without their permission. This decision reportedly came from CEO Mark Zuckerberg himself.

In 2024, Anthropic hired Tom Turvey to build “a central library of all the books in the world.” Turvey was instrumental in the creation of the Google Books project, in which the Mountain View, California-based company Google, with a market capitalization of $4.34 trillion, has scanned and digitized more than 40 million books in over 500 languages despite legal challenges, without destroying physical books. In other words, there was already a model that did not involve a Fahrenheit 451-style destruction of physical texts. But Anthropic seemed to embrace destructive scanning because it was the fastest and cheapest path forward. What other conclusion can be drawn from Turvey, once he left Google for Anthropic, initially contacting major publishers to inquire about licensing books for training Anthropic’s models? Clearly another way was considered and then abandoned—one that would have been less of a public relations nightmare. “Had Turvey kept up those conversations, he might have reached agreements to license copies for AI training from publishers,” legal filings state. “But Turvey let those conversations wither.”

Anthropic apparently proceeded believing that Project Panama could brew into a scandal, opting for secrecy instead of transparency. Selling the public on the largest destructive-scanning endeavor in history indeed seems like it would have been a tall order; taller still when you must explain why there is a premium placed on books published before 2022—the year OpenAI’s large language model ChatGPT became available to the general public.

Books printed before that point are much less likely to be tainted with AI writing, making them valuable to both labs seeking purely human data for training purposes and third-party providers that can discreetly handle the acquisition, such as ISBNdb. The International Standard Book Number, or ISBN, is a numeric code that is used to identify books. Each code contains information about a title, including the publisher, edition, and format. One of ISBNdb’s main offerings is “a vast collection of unique book data searchable by ISBN.” That’s an attractive service for merchants, libraries, and now possibly AI companies. On its website, the firm outlines how “pre-LLM” print books are guaranteed to be free of “contamination” of AI-generated text and what is known as “data poisoning”—methods creators use to sabotage models, ranging from semantic traps to digital camouflage and hidden prompts. AI companies hardly brag that they are  interested in this sort of business and ISBNdb, for its part, boasts an iron-clad “NDA on every engagement.”

Ironically, after the independent news outlet, 404 Media, reported on ISBNdb’s sales pitch, the service took down the page explaining its value to tech companies and issued a statement claiming it had pivoted in a different direction. “The page was a test of market interest; no such service was ever brought to life,” ISBNdb said. Although it repeatedly referenced Anthropic on the since-archived page, ISBNdb subsequently stated that it has never worked with it or any other AI lab. 

ISBNdb isn’t the only firm accused by merchants of attempting to snap up books from around the world for the purpose of training AI models. In the Netherlands, an antiquarian bookseller received an email from 2077AI, a Singapore-based AI lab, that stated it “is participating in a new project that focuses on collecting books in multiple languages, currently mainly in English.” 2077AI added that it had “compiled a very extensive list of editions that we are currently trying to get our hands on and we plan to place a fairly large order.”

The list, as reported on by the Dutch news outlet BNR, includes fairly recent texts, such as Barrett’s Traditional Fairy Tales, a book about Irish folk stories published in 2021. Other second-hand booksellers across Europe have reported similar contacts with third-party middleman services attempting to make huge orders with no discernible coherence in terms of content. One book shop in Germany has repeatedly received a massive number of orders of seemingly random books every night since May, between 3 and 5 a.m., made by a Canadian company called Zoom Books. According to Swiss broadcaster SRF, “Zoom Books bought pallets of out-of-print goods—cookbooks, biographies, fiction,” with special emphasis on older titles. Like ISBNdb, Zoom Books responded to press coverage by denying that these were acquisitions on behalf of AI companies. But these booksellers have expressed skepticism and are extremely wary of selling rare editions just to be scanned and destroyed. 

The backlash against the wholesale destruction of physical books has left some AI advocates baffled, chalking it up to ignorance, outrage at easy targets. In The Atlantic, Alex Reisner examined several of the central claims now circulating online. He wrote that part of the problem is that “the specifics of this situation are obscured by hard-to-trace purchasing histories and the silence of AI companies and other potential buyers,” leaving many unanswered questions. He also noted that among the books being destroyed are technical and academic titles, considered “rare” only in the narrow sense of being out of print. These are things the average person would presumably never read, let alone miss upon their being mulched. (And many, if not most, books end up in landfills or recycling without Big AI’s assistance.) And yet, Reisner still found plenty of cause for concern in the process of taking millions of books, denuding them of their creators, and throwing them into the maw of a chatbot. “It masks the hard work and collaboration that goes into knowledge creation, undermines incentives for authors to write books, prevents experts from finding one another, and gives companies tremendous power over what information people can access,” he concluded.

Those who wave away the outcry as the stuff of  book-obsessives risk missing the essential point Reisner is getting at, one that goes beyond books. It’s the same reason people don’t like the vision of a future in which every American lives under the watchful eye of an AI-powered surveillance network, such as the notions promoted by Garrett Langley, the founder and CEO of Flock, an Atlanta-based technology company whose wares include automated license plate readers. It is the same reason people in Ohio, where I live, protested against a proposal by the Ohio Environmental Protection Agency to make it easier for data centers to dump wastewater into lakes, rivers, and streams without consequence.

It’s not a mystery what links these varied assertions of tech power: There is a growing fear that the developers and proprietors of AI are recklessly scaling what is arguably the most world’s most powerful force at breakneck speed with little regard for the human things that are torn asunder in the wake, whether that’s physical books, art, privacy, or the environment. And whenever something like Project Panama comes to light, those concerns seem increasingly valid.

The post Burning the Past to Build the Future appeared first on Washington Monthly.

Read Entire Article