Tag: creators

  • AI Companies Are Destroying Physical Books. Here’s Why Your Business Should Care.

    AI Companies Are Destroying Physical Books. Here’s Why Your Business Should Care.

    Imagine spending years writing a book.

    Then imagine an AI company buying a second-hand copy, slicing off its spine, scanning every page and sending the remains for recycling. The words survive—but now as data inside a private system built to generate commercial products.

    This is not a dystopian thought experiment. Court records show that Anthropic, the company behind Claude, bought and destructively scanned millions of print books while building an internal digital library and training its AI models.

    The easy reaction is outrage: A technology company destroyed books to build a machine that writes.

    But for creators, ecommerce teams and business owners, the more useful question is this:

    If AI companies can treat physical knowledge as a resource to acquire, process and discard, how should you expect them to treat your website, product descriptions, customer conversations and creative work?

    That is the part every business should be thinking about.

    What actually happened?

    According to documents disclosed in the US copyright case Bartz v. Anthropic, Anthropic created a large internal collection of books from two very different sources.

    First, it downloaded more than seven million books from pirate websites. Second, it legally bought millions of printed books and converted them into digital files.

    The physical process was destructive by design. Books had their bindings or spines removed so that loose pages could pass through high-speed scanners. Court filings described industrial cutting equipment, production scanners and recycling of the paper after digitisation.

    In a June 2025 ruling, US District Judge William Alsup treated those two routes differently:

    • Converting legally purchased print books into internal digital replacements was held to be fair use in this case.
    • Acquiring and retaining pirated copies for a general-purpose library was not excused as fair use.
    • Training on the works was also held to be transformative on the record before the court.

    That distinction matters. “The court said AI companies can steal books” is not an accurate summary. The ruling separated lawful purchase and format conversion from the acquisition of pirated material.

    Anthropic later agreed to a $1.5 billion settlement concerning pirated books, without admitting wrongdoing. The settlement did not erase the court’s earlier fair-use ruling on training and the destructive scanning of lawfully purchased copies.

    Were rare books really destroyed?

    This is where the viral version of the story often runs ahead of the evidence.

    It is confirmed that Anthropic destructively scanned millions of purchased books. It is also true that booksellers in several countries have reported strange bulk orders containing obscure, old and out-of-print titles. Some sellers suspect those orders are connected to AI training and that the books may be pulped after scanning.

    However, there is not yet public proof that AI companies are systematically targeting and destroying rare or antiquarian books across the industry.

    Anthropic told The Guardian that its acquisition programmes do not buy and destroy rare or antiquarian books. The identities and intentions of buyers behind many of the unusual bulk orders remain unclear.

    So the responsible conclusion is:

    Mass destructive scanning is documented. The broader destruction of genuinely rare books is a serious concern, but it has not been established at the same level of certainty.

    That nuance does not make the story unimportant. It makes the real story more credible.

    Why would an AI company want physical books?

    Because the open web is no longer enough.

    Modern AI models need enormous quantities of high-quality language. Books are especially valuable because they contain edited, structured, long-form thinking—something the internet does not always provide.

    Physical books also offer three advantages.

    1. They contain material that may not exist online

    Many older, specialist and out-of-print works were never turned into commercial ebooks. Their pages hold information that is effectively invisible to internet-scale data collection.

    2. Older books contain less AI-generated material

    As AI-generated text spreads across the web, training future systems on indiscriminate online data risks feeding models content produced by other models. Pre-generative-AI books are attractive because their human origin is easier to establish.

    3. Buying a physical copy can create a cleaner legal position

    The Anthropic ruling shows why acquisition method matters. Buying a copy, destroying it and keeping one internal digital replacement presented a stronger fair-use argument than downloading an unauthorised digital copy.

    In other words, this was not simply a knowledge project. It was also a data-sourcing and legal-risk strategy.

    The uncomfortable business lesson: your content is an input

    Most businesses still think about AI tools as products they consume.

    You pay for a chatbot, connect an API or add an AI assistant to your workflow. It feels like a normal software relationship: the vendor provides the tool, and you use it.

    But AI platforms are also built around inputs. They need language, images, behaviour, feedback and context. Your business may be a customer on one side of that system and a source of valuable data on the other.

    That does not mean every AI provider trains on every prompt or secretly takes every file. Policies, contracts and product settings differ. Enterprise and API offerings often include stronger data controls than free consumer tools.

    The point is simpler: never assume your content is protected merely because you created it or because it sits inside a tool you pay for. Protection comes from clear terms, technical controls and deliberate choices.

    What this means for ecommerce and marketing teams

    For an ecommerce business, “content” is not just blog posts.

    It includes product descriptions, photography, customer reviews, campaign concepts, brand voice, internal merchandising rules, conversion experiments, support tickets and pricing logic. Individually, these assets may look ordinary. Together, they describe how your company competes.

    If teams paste that material into AI tools without checking the terms, they may expose far more than a few paragraphs of copy.

    Consider four common situations:

    A marketer uploads next quarter’s campaign plan

    The document may contain unreleased offers, audience insights, budgets and positioning. The risk is not only copyright. It is confidentiality.

    A product team feeds an entire catalogue into a writing tool

    Generated descriptions may save time, but the input also reveals assortment strategy, attributes and product data. Who can retain it, and for how long?

    Customer service uses public AI tools to rewrite tickets

    Those tickets may contain names, addresses, order details or health information. Now the issue includes privacy and GDPR—not just content ownership.

    A creator builds a brand on a third-party model

    If the model, price, policy or output quality changes, the creator’s workflow can break overnight. Dependence becomes a platform risk.

    Five practical actions businesses should take now

    You do not need to stop using AI. You need to stop using it casually.

    1. Classify information before it enters an AI tool

    Create three simple categories: public, internal and restricted. Public material may be acceptable in approved tools. Internal content needs controls. Restricted data—such as personal information, credentials, contracts and unreleased financials—should not enter an unapproved system.

    2. Read the terms that matter

    Check whether the provider may retain inputs, use them to improve models, allow human review or share them with subprocessors. Confirm whether training is disabled by default, optional or unavailable for your plan.

    Do not let “enterprise-grade” function as a substitute for reading the contract.

    3. Keep an original source of truth

    Store product copy, research, images, prompts and campaign assets in systems you control. AI output should enter your workflow; your workflow should not live entirely inside one AI platform.

    4. Preserve human provenance

    Keep drafts, timestamps, licences and approval records for important creative assets. This helps demonstrate where work came from, what a human contributed and which material you had permission to use.

    5. Avoid single-model dependence

    Build processes around tasks and standards rather than one vendor’s interface. Where practical, keep prompts portable, retain exports and test a backup provider. The goal is not to switch tools every week. It is to maintain leverage.

    Legal does not automatically mean ethical—or wise

    The court’s decision addressed specific copyright questions under US law. It did not settle every ethical question raised by destroying physical books, nor did it create a universal rule for every AI model, dataset or country.

    A purchased mass-market paperback is not the same thing as a fragile edition with annotations, a distinctive binding or historical provenance. A digital text can preserve words while losing the object’s physical evidence.

    The environmental picture is complicated too. Recycling the paper is better than sending it to landfill, but buying, transporting, cutting and scanning millions of books still consumes material and energy. A company can follow a legally defensible process without proving it chose the most responsible one.

    For businesses, that distinction is essential. Compliance asks, “Are we allowed to do this?” Trust asks, “Will customers, creators and partners believe this is fair?”

    The strongest brands need an answer to both.

    The real story is not about paper

    Physical books make this issue visible because we understand what is being lost. We can picture the blade cutting through the binding. We can see the pages becoming data.

    Digital extraction is easier to ignore. A website can be scraped without an empty shelf. A creator’s style can be absorbed without a damaged cover. A customer conversation can become a data point without anyone hearing the paper shredder.

    That is why this story matters.

    AI is not magic floating above the economy. It is infrastructure built from human work: books, art, code, conversations, decisions and data. Businesses benefiting from these systems should ask where those inputs came from—and apply the same scrutiny to where their own information goes.

    Use AI. Experiment with it. Build with it.

    But do not confuse convenience with control.

    Frequently asked questions

    Are AI companies really destroying physical books?

    Yes, in at least one well-documented case. Court records confirm that Anthropic bought and destructively scanned millions of physical books, removing bindings or spines and recycling the remains after digitisation.

    Are AI companies destroying rare books?

    There are credible reports of unusual purchases involving obscure, old and out-of-print books, and booksellers suspect AI-related buyers. However, systematic destruction of genuinely rare or antiquarian books has not been conclusively established. Anthropic denies buying and destroying rare or antiquarian books through its acquisition programmes.

    Was Anthropic’s scanning ruled legal?

    In June 2025, a US federal judge held that converting lawfully purchased print books into internal digital replacements was fair use in the specific case. The same ruling did not excuse Anthropic’s acquisition and retention of pirated library copies.

    Can an AI company train on my business content?

    It depends on how the content is obtained, the provider’s terms, your product tier, applicable law and the settings or contract governing your account. Businesses should verify these conditions rather than assume all AI tools handle data in the same way.

    Should businesses stop using generative AI?

    No. Businesses should use approved tools, classify sensitive information, understand provider terms, preserve source files and avoid depending entirely on a single model or platform.

    Sources and further reading