The LedeDestroying Books to Build a MindAnthropic is trying to “destructively scan all the books in the world.” How worried should we be?By Francesca MancinoSeptember 8, 2026Illustration by Anthony Gerace; Source photographs from GettySave this storySave this storySave this storySave this storyIn early 2026, the owner of James Payne, Books and Prints, based in Brooklyn, began receiving order requests that he later described as “brazen and bizarre.” The books themselves were strange—a so-called legal survival guide from 1998 titled “How to Win in Small Claims Court in New York”; an academic text from 2006 called “Cultures of Glass Architecture”—and so was the order method. Booksellers like Payne typically post their holdings on multiple websites, such as AbeBooks, Alibris, Amazon, and Biblio, with Alibris widely considered to be the sleepiest revenue stream of the bunch. Suddenly, Payne went from selling fewer than ten books annually on Alibris to fulfilling orders of ten books, averaging as much as seventy-five to eighty dollars each, at a time on the site. Other booksellers around the country reported similar happenings.Sylvia Petras, who owns Leaf and Stone Books, in Toronto, was envious as she watched book dealers in the U.S. post about their surging sales. Then the orders started trickling into her store, too, and that envy was quickly replaced by confusion. Petras owns a vast inventory of printed books from the fifteenth to the seventeenth centuries, along with scarce and scholarly books; odd titles, on subjects such as sewage plants, were flying off the shelves. Like other booksellers who were inundated with requests for niche books, Petras noticed that the orders were not being placed under an individual’s name but, rather, by enigmatic L.L.C.s: Green Parrot Project and Red Sparrow Project. One bookseller, who asked to remain anonymous, decided to attach a G.P.S. tracker to one of the purchased books before shipping it out to the mystery buyer. After researching her options, she selected a SmartCard, a device resembling a credit card, slipped it into a small envelope, and sealed it onto the inner rear board of a hardcover book, behind the dust jacket. She then followed along, online, as the book moved from Grandview Heights, Ohio, to an industrial park in Addison, Illinois, where a company called ARC Document Solutions operates an industrial-grade scanning facility. In a promotional video, an ARC spokesperson explains that many of the company’s sites (there are more than a hundred nationwide) run 24/7. He refers to a client who has hired ARC to scan “one football field with boxes stacked four-high every month.”Anthropic, the artificial-intelligence company behind the family of large language models known as Claude, is trying to acquire as many printed books as possible. We know this because of Anthropic’s own legal filings, unsealed in a copyright suit that was brought against the company back in 2024. Having identified books as the “highest quality source of training data” for Anthropic’s L.L.M.s, it initiated a covert program called Project Panama, defined succinctly in court documents as “our effort to destructively scan all the books in the world.” “Destructively scan” means precisely that: Anthropic or its affiliates “stripped the bindings from the print books, cut the pages to workable dimensions, and scanned those pages—discarding each print copy while creating a digital one in its place,” according to the court filings.The process involves the use of a hydraulic-powered cutting machine, also known as a “book guillotine.” It’s a sinister-sounding practice that is, actually, relatively commonplace, with university-preservation departments routinely disbinding books in order to make them easier to scan. Destroying a book isn’t the only way to scan it, of course. The library at Princeton makes a point to retain physical copies of the texts they are digitizing, even if it makes for a less efficient process. “It’s totally possible, though more expensive, to take images that preserve the integrity of the book without flattening it,” Meredith Martin, a professor of English and the faculty director of the Center for Digital Humanities at Princeton, told me.Still, Martin and other experts emphasized that old books are thrown away—and subsequently destroyed—all the time, even for non-academic reasons, by libraries that are simply culling their collections. Donation is obviously more palatable, but it can be difficult to off-load, say, an outdated instruction manual or textbook, or several worn-out copies of the same mass-market paperback. (Some of these books won’t even be accepted by prison libraries, which generally reject hardcovers or texts that are in poor condition.) As a result, these books might end up in a landfill, or, depending on their bindings, get broken down and pulped for recycling, or perhaps shredded. In a Medium post where she makes the case for libraries “weeding” their shelves, Claire Sewell, an academic librarian in Houston, recognizes that “seeing a dumpster full of books can seem completely antithetical to everything libraries are supposed to stand for as repositories of knowledge.” And yet, “old, outdated, damaged, or simply low circulating books have to be weeded on a regular basis in order for us to make space for new books that you’ll actually want to check out.”What made Anthropic’s project controversial, though, wasn’t the fear that it would destroy a water-damaged copy of “The Secret,” but, instead, genuinely rare and valuable books. To satisfy its voracity for billions of pages, Anthropic started out by placing mass orders with wholesalers, before turning to individual and secondhand booksellers to help fill in the gaps. The company’s interest in obscure texts led to headlines about A.I. companies buying and destroying “rare old books,” or “antique books,” which generally calls to mind first editions.There are many legitimate ethical concerns associated with the rise of A.I. companies like Anthropic, but the destruction of valuable books isn’t necessarily one of them. “None of our data acquisition programs buy and destroy rare or antiquarian books,” an Anthropic spokesperson wrote, in a statement. (“Antiquarian” refers to books that are at least a hundred years old, regardless of rarity.) The Antiquarian Booksellers’ Association of America said it has not been notified by any of its members that such books have been dealt to A.I. companies. Of the individual booksellers I spoke with, none has sold any antiquarian books to buyers that seemed to be A.I. companies. Rather, the majority of the books they’ve sold have ISBNs, unique numerical identifiers that were not instituted until 1970.Petras, the bookstore owner in Toronto, said that she was generally comfortable selling to undercover buyers, but that she would never sell them a book with interesting marginalia or an important provenance. Joyce Kosofsky, one of the owners of Brattle Book Shop, in Boston, said that the books she sold all had multiple copies. “We probably had them at the cheapest price,” she guessed. In general, she argued, books are no different from any other saleable good: “Just like when you go into a clothing store, and you buy a pair of jeans—they’re your jeans. You can wear them. You can decorate them. You can give them away. No one follows you around saying, ‘What are you going to do with your jeans?’ ”Given that Anthropic is purchasing books that are “likely not very rare,” Martin, the Princeton professor, said, the company’s use of book guillotines shouldn’t be considered inflammatory. But, if the books aren’t rare, then what are they? The booksellers shared the names of more than six hundred titles that they believed they had sold to A.I. companies, and I sent the list to Melanie Walsh, an Assistant Professor in the Information School at the University of Washington, to process digitally. The texts were obscure in their subject matter and lack of popularity: these were books, sometimes with low-print runs—usually a thousand copies or fewer—to meet realistic market demands. Of the top ten publishers, eight were academic presses—roughly a third of the sample over all. Most of the books were published between the nineteen-seventies and the twenty-tens, and the genres spanned history, biography, fiction, poetry, literary criticism, law, and the social sciences. “Based on this sample, it appears that A.I. companies may be interested in training models on peer-reviewed academic research across a wide range of subjects,” Walsh concluded. This dovetails with what Mycal Tucker, a research scientist at Anthropic, discovered while organizing training data for an A.I. model: “Nonfiction works tend to be more valuable than fiction,” he said in his written testimony, adding that nonfiction-book data helped the model perform well in disciplines as diverse as philosophy and astronomy.Walsh said that when it comes to nonfiction works that are specialized, as is the case with “Utilization of Municipal Wastewater Sludge,” a 1972 booklet that Petras recently sold—“there’s an argument to be made that these obscure academic books may make a bigger impact as part of a Claude model than they would otherwise.” After all, they were already headed toward obsolescence.Anthropic’s attempts to get its hands on “all the books in the world” has attracted legal challenges. In 2025, the company agreed to pay $1.5 billion to settle a class-action lawsuit brought by a group of authors who accused the company of violating their copyrights, namely by using their books for A.I. training without their permission. William Alsup, the judge presiding over the case, reprimanded Anthropic for some of its actions, such as downloading over seven million pirated copies of books and keeping the files “as a permanent, general-purpose resource,” even if they weren’t being used to train Claude. (“Anthropic seems to believe that because some of the works it copied were sometimes used in training L.L.M.s, Anthropic was entitled to take for free all the works in the world and keep them forever with no further accounting,” Alsup wrote in his decision.)But the larger practice of using books to train A.I. models was fair use, according to Alsup. “Authors cannot rightly exclude anyone from using their works for training or learning,” he wrote. “Everyone reads texts, too, then writes new texts. They may need to pay for getting their hands on a text in the first instance. But to make anyone pay specifically for the use of a book each time they read it, each time they recall it from memory, each time they later draw upon it when writing new things in new ways would be unthinkable.”Brandon Butler, a copyright lawyer and the executive director of the Re:Create Coalition, an advocacy group that supports both the creators of proprietary works and their users, described A.I. training as “the fairest use in the history of copyright law.” This is because of how it “takes from existing culture only the stuff that actually belongs to all of us—facts, ideas, elements of language and grammar—and makes it easier for all of us to use,” Butler argued. Copyright could be infringed upon if L.L.M.s regurgitated memorized material verbatim or conveyed protected expression, but they generally convey unprotected information in a way that varies from the original source’s wording. In this way, there is no marketplace competition between A.I. companies and authors, Butler said, because someone who encounters a paraphrased excerpt of a book is still incentivized to buy the book. (This is complicated by how some L.L.M.s, including Meta’s Llama, have occasionally “memorized” entire novels nearly verbatim.)According to Judge Alsup, Anthropic was also within its rights to destroy the print books that it had obtained legally. “Anthropic purchased millions of print copies ‘to build a research library,’ ” he wrote. “It destroyed each print copy while replacing it with a digital copy for use in its library (not for sharing nor sale outside the company).” Since the format change was for convenience—easier storage and search functionality—and did not result in more copies of the original books, it was fair use.Legality aside, though, we are met with difficulties, even just on a linguistic level, if Anthropic’s understanding of the term “research library” is so different from ours that it feels like a misnomer. They’re “not building a library because somebody might, one day, want to read all these books,” Butler explained. “They’re not interested in helping anyone else read the books. What they want, really, is data.” Their digital collection is being composed with the sole intent to enlarge Anthropic’s L.L.M.s’ “memory ” and to hone their writing abilities.The consequence is that, even if these models successfully preserve literary materials that might otherwise be lost to time or foundering interest, the source will always remain in the shadows. Researchers will not have access to the book data, and users cannot control or observe how information is being sourced to them. Mike Furlough, the executive director of HathiTrust, told me that it is our right to ask A.I. companies for more transparency about how their data sets are created, notwithstanding that their mission is not to distribute books and that publicizing this data might sacrifice their competitive edge. Libraries and academic researchers are held to a much more rigorous standard: “an academic researcher who wants to produce a small-scale language model,” he said, “would be expected to disclose the entire list of books that they would have used.” He continued, “That’s just common academic practice because you want to be able to reproduce that, or understand what went into that.” The same is true of basically any reputable tool that a layperson would use for research; even Wikipedia includes citations. Anthropic’s goal to corral “all the books in the world” could almost be exciting, were it not for the fact that this hypothetical repository of written knowledge would exist only in the digital depths of Claude.“Every generation rewrote the book’s epitaph,” the historian Leah Price wrote, in a nonfiction history of the medium, in 2019. “All that changes is whodunnit.” In the nineteenth century, the suspected murderer was libraries: in a pamphlet printed in London, titled “The Truth About Giving Readers Free Access to the Books in a Public Lending Library,” an anonymous writer sought to expose the ghastly reality of free access, which could result in the theft and “misplacement ” of books. More recently, fears coalesced around the advent of the e-book and the Kindle.In “What We Talk About When We Talk About Books,” Price points out that books have always had a chameleonic quality—that they adapt with technological innovation, and that the two things aren’t necessarily always in tension. And yet, what it means to be a reader, a consumer of literature, is a concept that’s beginning to shift, too. In an e-mail to me, Price expressed concerns about the looming threat that “A.I. will replace human readers,” since “destructive scanning sacrifices multiple potential future human readings to a single instance of machine reading.” Books come with baggage, and A.I. shears them of that. Books demand our time, involvement, and deliberation—each one has a unique set of readerly requirements. In a 2006 essay titled “The Rise of Fictionality,” exploring the proliferation of fiction writing in the eighteenth century, the author Catherine Gallagher lists the ways in which early novels commanded the reader’s attention and engagement, asking them “to anticipate problems, make suppositional predictions, and see possible outcomes and alternative interpretations.” If one still believes in reading as an “imaginative, immersive experience,” Price told me, then there is a critical contrast—particularly when it comes to hallmark works of literature—between unmediated human reading versus bots doing the reading for users and extracting worthwhile content. “A novel is not just a bucket of sentences,” Price said. “The order in which those sentences occur matters.” And yet Anthropic’s goal, it seems, is to accumulate and spit out ever more buckets.Matthew Kirschenbaum, a Commonwealth Professor of Artificial Intelligence and English at the University of Virginia, warned of what will happen if the “great virtue of the book”—its existence as a self-contained object, with a human writer, and a human history—is dissolved into data that siphons all context from the original. The boundaries of the book (e.g., its covers and materiality, publication information, and author) disappear when the text is “homogenized” and “loses all of its relationships that were present in the original book.” Thus, the very thing protecting A.I. from copyright infringement—its ability to rephrase things—is also the problem: L.L.M.s transform literature into some hybridized, half-dead thing that is far removed, if not completely untethered, from its original source. It creates a product so drastically new that the text as an individual work is effectively destroyed. This less obvious form of erasure is much more hazardous and worrisome than destructive scanning.Even still, I wonder if obtaining information by way of L.L.M.s will only ever intrigue a percentage of the population. There is a parallel here with the rise of the novel in the eighteenth century. After Daniel Defoe’s “Robinson Crusoe” and Samuel Richardson’s “Pamela” were published, in 1719 and 1740, respectively, multiple print runs were exhausted throughout the remainder of the century. Competitors, noticing the copies being printed to keep up with demand, devised a way to make a quick buck: they published a variety of pirated editions. These second-rate versions included child-friendly copies, abridgements preserving the best-written or juiciest bits, shortened texts that were simply cheaper than the full-length ones, and unauthorized sequels. These pirated books were incredibly popular, too, even as people continued to seek out the original texts. We can divide these eighteenth-century consumers into two camps: those who bought and read the novels written by Defoe and/or Richardson, and those who were partial to altered or imitative copies. There have always been—and there will perhaps always be—people who prefer slop over the primary source. ?