Tag: Anthropic

  • AI Companies Are Destroying Physical Books. Here’s Why Your Business Should Care.

    AI Companies Are Destroying Physical Books. Here’s Why Your Business Should Care.

    Imagine spending years writing a book.

    Then imagine an AI company buying a second-hand copy, slicing off its spine, scanning every page and sending the remains for recycling. The words survive—but now as data inside a private system built to generate commercial products.

    This is not a dystopian thought experiment. Court records show that Anthropic, the company behind Claude, bought and destructively scanned millions of print books while building an internal digital library and training its AI models.

    The easy reaction is outrage: A technology company destroyed books to build a machine that writes.

    But for creators, ecommerce teams and business owners, the more useful question is this:

    If AI companies can treat physical knowledge as a resource to acquire, process and discard, how should you expect them to treat your website, product descriptions, customer conversations and creative work?

    That is the part every business should be thinking about.

    What actually happened?

    According to documents disclosed in the US copyright case Bartz v. Anthropic, Anthropic created a large internal collection of books from two very different sources.

    First, it downloaded more than seven million books from pirate websites. Second, it legally bought millions of printed books and converted them into digital files.

    The physical process was destructive by design. Books had their bindings or spines removed so that loose pages could pass through high-speed scanners. Court filings described industrial cutting equipment, production scanners and recycling of the paper after digitisation.

    In a June 2025 ruling, US District Judge William Alsup treated those two routes differently:

    • Converting legally purchased print books into internal digital replacements was held to be fair use in this case.
    • Acquiring and retaining pirated copies for a general-purpose library was not excused as fair use.
    • Training on the works was also held to be transformative on the record before the court.

    That distinction matters. “The court said AI companies can steal books” is not an accurate summary. The ruling separated lawful purchase and format conversion from the acquisition of pirated material.

    Anthropic later agreed to a $1.5 billion settlement concerning pirated books, without admitting wrongdoing. The settlement did not erase the court’s earlier fair-use ruling on training and the destructive scanning of lawfully purchased copies.

    Were rare books really destroyed?

    This is where the viral version of the story often runs ahead of the evidence.

    It is confirmed that Anthropic destructively scanned millions of purchased books. It is also true that booksellers in several countries have reported strange bulk orders containing obscure, old and out-of-print titles. Some sellers suspect those orders are connected to AI training and that the books may be pulped after scanning.

    However, there is not yet public proof that AI companies are systematically targeting and destroying rare or antiquarian books across the industry.

    Anthropic told The Guardian that its acquisition programmes do not buy and destroy rare or antiquarian books. The identities and intentions of buyers behind many of the unusual bulk orders remain unclear.

    So the responsible conclusion is:

    Mass destructive scanning is documented. The broader destruction of genuinely rare books is a serious concern, but it has not been established at the same level of certainty.

    That nuance does not make the story unimportant. It makes the real story more credible.

    Why would an AI company want physical books?

    Because the open web is no longer enough.

    Modern AI models need enormous quantities of high-quality language. Books are especially valuable because they contain edited, structured, long-form thinking—something the internet does not always provide.

    Physical books also offer three advantages.

    1. They contain material that may not exist online

    Many older, specialist and out-of-print works were never turned into commercial ebooks. Their pages hold information that is effectively invisible to internet-scale data collection.

    2. Older books contain less AI-generated material

    As AI-generated text spreads across the web, training future systems on indiscriminate online data risks feeding models content produced by other models. Pre-generative-AI books are attractive because their human origin is easier to establish.

    3. Buying a physical copy can create a cleaner legal position

    The Anthropic ruling shows why acquisition method matters. Buying a copy, destroying it and keeping one internal digital replacement presented a stronger fair-use argument than downloading an unauthorised digital copy.

    In other words, this was not simply a knowledge project. It was also a data-sourcing and legal-risk strategy.

    The uncomfortable business lesson: your content is an input

    Most businesses still think about AI tools as products they consume.

    You pay for a chatbot, connect an API or add an AI assistant to your workflow. It feels like a normal software relationship: the vendor provides the tool, and you use it.

    But AI platforms are also built around inputs. They need language, images, behaviour, feedback and context. Your business may be a customer on one side of that system and a source of valuable data on the other.

    That does not mean every AI provider trains on every prompt or secretly takes every file. Policies, contracts and product settings differ. Enterprise and API offerings often include stronger data controls than free consumer tools.

    The point is simpler: never assume your content is protected merely because you created it or because it sits inside a tool you pay for. Protection comes from clear terms, technical controls and deliberate choices.

    What this means for ecommerce and marketing teams

    For an ecommerce business, “content” is not just blog posts.

    It includes product descriptions, photography, customer reviews, campaign concepts, brand voice, internal merchandising rules, conversion experiments, support tickets and pricing logic. Individually, these assets may look ordinary. Together, they describe how your company competes.

    If teams paste that material into AI tools without checking the terms, they may expose far more than a few paragraphs of copy.

    Consider four common situations:

    A marketer uploads next quarter’s campaign plan

    The document may contain unreleased offers, audience insights, budgets and positioning. The risk is not only copyright. It is confidentiality.

    A product team feeds an entire catalogue into a writing tool

    Generated descriptions may save time, but the input also reveals assortment strategy, attributes and product data. Who can retain it, and for how long?

    Customer service uses public AI tools to rewrite tickets

    Those tickets may contain names, addresses, order details or health information. Now the issue includes privacy and GDPR—not just content ownership.

    A creator builds a brand on a third-party model

    If the model, price, policy or output quality changes, the creator’s workflow can break overnight. Dependence becomes a platform risk.

    Five practical actions businesses should take now

    You do not need to stop using AI. You need to stop using it casually.

    1. Classify information before it enters an AI tool

    Create three simple categories: public, internal and restricted. Public material may be acceptable in approved tools. Internal content needs controls. Restricted data—such as personal information, credentials, contracts and unreleased financials—should not enter an unapproved system.

    2. Read the terms that matter

    Check whether the provider may retain inputs, use them to improve models, allow human review or share them with subprocessors. Confirm whether training is disabled by default, optional or unavailable for your plan.

    Do not let “enterprise-grade” function as a substitute for reading the contract.

    3. Keep an original source of truth

    Store product copy, research, images, prompts and campaign assets in systems you control. AI output should enter your workflow; your workflow should not live entirely inside one AI platform.

    4. Preserve human provenance

    Keep drafts, timestamps, licences and approval records for important creative assets. This helps demonstrate where work came from, what a human contributed and which material you had permission to use.

    5. Avoid single-model dependence

    Build processes around tasks and standards rather than one vendor’s interface. Where practical, keep prompts portable, retain exports and test a backup provider. The goal is not to switch tools every week. It is to maintain leverage.

    Legal does not automatically mean ethical—or wise

    The court’s decision addressed specific copyright questions under US law. It did not settle every ethical question raised by destroying physical books, nor did it create a universal rule for every AI model, dataset or country.

    A purchased mass-market paperback is not the same thing as a fragile edition with annotations, a distinctive binding or historical provenance. A digital text can preserve words while losing the object’s physical evidence.

    The environmental picture is complicated too. Recycling the paper is better than sending it to landfill, but buying, transporting, cutting and scanning millions of books still consumes material and energy. A company can follow a legally defensible process without proving it chose the most responsible one.

    For businesses, that distinction is essential. Compliance asks, “Are we allowed to do this?” Trust asks, “Will customers, creators and partners believe this is fair?”

    The strongest brands need an answer to both.

    The real story is not about paper

    Physical books make this issue visible because we understand what is being lost. We can picture the blade cutting through the binding. We can see the pages becoming data.

    Digital extraction is easier to ignore. A website can be scraped without an empty shelf. A creator’s style can be absorbed without a damaged cover. A customer conversation can become a data point without anyone hearing the paper shredder.

    That is why this story matters.

    AI is not magic floating above the economy. It is infrastructure built from human work: books, art, code, conversations, decisions and data. Businesses benefiting from these systems should ask where those inputs came from—and apply the same scrutiny to where their own information goes.

    Use AI. Experiment with it. Build with it.

    But do not confuse convenience with control.

    Frequently asked questions

    Are AI companies really destroying physical books?

    Yes, in at least one well-documented case. Court records confirm that Anthropic bought and destructively scanned millions of physical books, removing bindings or spines and recycling the remains after digitisation.

    Are AI companies destroying rare books?

    There are credible reports of unusual purchases involving obscure, old and out-of-print books, and booksellers suspect AI-related buyers. However, systematic destruction of genuinely rare or antiquarian books has not been conclusively established. Anthropic denies buying and destroying rare or antiquarian books through its acquisition programmes.

    Was Anthropic’s scanning ruled legal?

    In June 2025, a US federal judge held that converting lawfully purchased print books into internal digital replacements was fair use in the specific case. The same ruling did not excuse Anthropic’s acquisition and retention of pirated library copies.

    Can an AI company train on my business content?

    It depends on how the content is obtained, the provider’s terms, your product tier, applicable law and the settings or contract governing your account. Businesses should verify these conditions rather than assume all AI tools handle data in the same way.

    Should businesses stop using generative AI?

    No. Businesses should use approved tools, classify sensitive information, understand provider terms, preserve source files and avoid depending entirely on a single model or platform.

    Sources and further reading

  • Do OpenAI and Anthropic Really Drive 70% of AI Revenue? What It Means for Your Business

    Do OpenAI and Anthropic Really Drive 70% of AI Revenue? What It Means for Your Business

    One statistic is racing around the AI industry:

    More than 70% of AI revenue comes from OpenAI and Anthropic.

    It is a powerful number. It suggests that thousands of AI products, billions in infrastructure spending and the strategies of the world’s largest technology companies rest on two model providers.

    It is also easy to repeat incorrectly.

    The available evidence does not establish that OpenAI and Anthropic collect 70% of every dollar earned across the entire AI market. The figure comes from analyst estimates about a narrower—and in some ways more revealing—part of the ecosystem: the AI-related revenue earned by Amazon, Microsoft and Google from cloud compute, model access and associated commercial arrangements.

    In other words, the claim is less “70% of AI revenue flows to two companies” and more:

    Analysts estimate that OpenAI and Anthropic may directly or indirectly drive more than 70% of the AI-related revenue attributed to the three largest US cloud platforms.

    Some of that money flows from OpenAI and Anthropic to cloud providers for compute. Some comes from cloud customers buying access to their models. The exact totals are not disclosed cleanly by the companies and the estimates differ by analyst.

    That correction weakens the viral headline—but strengthens the business lesson.

    The AI economy may be broader than two companies. The commercial infrastructure underneath it is still remarkably concentrated. If your marketing workflow, ecommerce operation or software product depends on one foundation-model provider, you are not merely choosing a tool. You are inheriting that provider’s pricing, availability, policy and strategic risk.

    Fact check: what does the 70% figure actually measure?

    The source behind the current discussion is Ed Zitron’s analysis, “The AI Demand Bubble”. It combines estimates attributed to analysts at Barclays, UBS and Wells Fargo.

    The estimates cited in that analysis include:

    • Amazon Web Services: One Barclays estimate put OpenAI and Anthropic at 73% of AWS AI revenue in 2026 and 2027. A separate estimate cited in the same article put their direct compute contribution at 59% in 2026.
    • Google Cloud: UBS estimates cited in the analysis assigned 28% of total 2026 Google Cloud revenue to OpenAI and Anthropic, rising to more than 48% in 2027. The author then inferred that the pair could represent at least 70% of Google’s narrower AI-related revenue.
    • Microsoft: Wells Fargo estimates cited in the article put OpenAI and Anthropic at 70% or more of Microsoft’s AI revenue, reaching approximately 74% in the relevant forecast period.

    These are not one consistent, audited market-share dataset. They mix:

    • direct purchases of computing capacity;
    • cloud revenue associated with the two laboratories;
    • revenue from platforms that resell access to their models;
    • analyst forecasts for future periods;
    • the article author’s own classification and inference.

    The cloud companies do not publish a standard “AI revenue” line that allows outsiders to calculate a definitive global market share. OpenAI and Anthropic are also private companies, so their financial disclosure is more limited than that of a public company.

    The accurate version of the claim

    Use this formulation:

    Analyst estimates suggest OpenAI and Anthropic account for roughly 70% or more of the AI-related revenue attributed to Amazon, Microsoft and Google, although definitions and estimates vary.

    Avoid this formulation:

    OpenAI and Anthropic receive 70% of all revenue in the global AI industry.

    That broader statement would require a defined market covering chips, cloud infrastructure, enterprise software, consumer subscriptions, services, advertising, robotics, data platforms and other AI-related businesses. The cited analysis does not provide that calculation.

    Why the corrected number still matters

    The statistic is not a clean measure of the whole AI market, but it exposes three forms of concentration.

    1. Demand concentration

    Cloud providers have invested extraordinary amounts in data centres, accelerators and power capacity. If a large share of the associated revenue depends on two customers and their models, the infrastructure boom has a narrower demand base than the headline “AI adoption” numbers imply.

    This does not prove the boom will collapse. It means future growth depends heavily on OpenAI and Anthropic continuing to:

    • attract paying customers;
    • raise or generate enough cash to fund compute;
    • turn model capability into sustainable demand;
    • serve workloads efficiently enough to support margins;
    • maintain favourable relationships with cloud partners.

    2. Model concentration

    Many applications are not independent AI businesses in a technical sense. They are interfaces, workflows or specialised data layers built on a small number of foundation models.

    That can be a perfectly good business. Shopify did not need to build a payment network, and SaaS companies do not manufacture their own processors. Specialisation creates value.

    The risk begins when the application has no meaningful advantage beyond one provider’s output and cannot operate if that provider changes.

    3. Strategic concentration

    OpenAI and Anthropic influence more than model quality. Their decisions can shape:

    • token and subscription prices;
    • API limits and access tiers;
    • context-window and tool-use behaviour;
    • model retirement schedules;
    • safety policies and refused use cases;
    • data-processing terms;
    • regional availability;
    • integration standards;
    • which workflows become economically viable.

    A business built on top of one provider may experience these decisions as product changes—even when it had no voice in making them.

    What happens if one of the two stumbles?

    “Stumbles” does not have to mean bankruptcy. For a customer, smaller changes can create the same operational effect.

    Prices rise

    If inference pricing increases or a subsidised product becomes more expensive, an application with weak margins may become uneconomic overnight.

    This is especially dangerous when the company offers customers a fixed monthly price while paying the model provider per token, image, tool call or unit of compute.

    A model or feature is retired

    Prompts tuned for one model do not automatically behave the same on its replacement. Output format, tone, refusal patterns, latency and tool selection can all change.

    Without regression tests, a “simple upgrade” can quietly damage product listings, customer replies, campaign copy or structured data.

    Reliability declines

    An outage at the foundation-model layer can stop every workflow built above it. Even partial degradation—higher latency, elevated errors or inconsistent tool calls—can create queues, duplicate actions and failed customer experiences.

    Regulation or litigation changes access

    New regulatory restrictions, court decisions, government procurement rules or regional compliance requirements can affect how models are offered and which data may be processed.

    The correct response is not to predict one dramatic ban. It is to ensure that a single legal or policy change cannot disable an essential workflow without an alternative.

    Provider strategy shifts

    A model company can enter your category directly, prioritise enterprise contracts, discontinue a partner feature or bundle functionality that makes your product less differentiated.

    Platform risk is not only technical. Your supplier can become your competitor.

    What AI market concentration means for marketers

    For marketers, the immediate temptation is to treat model choice as a creative preference: Which assistant writes the strongest hooks? Which one follows brand voice best? Which one creates the most attractive images?

    Those questions matter, but operational dependence matters more.

    A marketing stack may use one provider for:

    • campaign research;
    • segmentation ideas;
    • advertisement variations;
    • product copy;
    • email personalisation;
    • image generation;
    • social scheduling;
    • performance analysis.

    If every step depends on one vendor, a policy update or service interruption can stop the entire content pipeline.

    The better approach is to distinguish between creative preference and business-critical dependency.

    • It is reasonable to prefer one model for campaign concepts.
    • It is risky if no other model can render the required data structure.
    • It is reasonable to use one assistant for drafts.
    • It is risky if brand knowledge exists only inside that provider’s proprietary workspace.
    • It is reasonable to optimise prompts for quality.
    • It is risky if no regression suite tells you when an update changes the output.

    Keep brand guidelines, approved claims, product facts, audience definitions and reusable prompt templates in systems you control. The model should consume your marketing intelligence, not become the only place where it exists.

    What it means for ecommerce brands

    Ecommerce companies have a deeper dependency problem because AI is moving from content generation into operational action.

    Models increasingly help with:

    • catalogue enrichment;
    • onsite search and recommendations;
    • customer-service responses;
    • translations;
    • merchandising analysis;
    • campaign creation;
    • pricing recommendations;
    • returns and order workflows.

    The closer AI gets to customers, orders and money, the more expensive provider concentration becomes.

    Your catalogue must remain the source of truth

    Store product attributes, claims, translations and policy rules in your own product-information or commerce systems. Do not let one model’s memory or proprietary knowledge feature become the authoritative record.

    If you are evaluating content tools, the criteria in our AI tools for ecommerce product listings benchmark remain useful: accuracy, structured output, brand consistency, channel adaptation and measurable workflow performance matter more than a flashy one-off result.

    Separate recommendations from execution

    An alternative model can replace a copywriting assistant relatively easily. Replacing an autonomous agent that can change prices, send campaigns or issue refunds is much harder.

    Our analysis of why humans missed one in three dangerous AI agent commands explains why manual approval alone is insufficient. Permissions, spending limits, audit trails and rollback must be enforced outside the model.

    Design graceful degradation

    If the preferred model is unavailable, decide what the store should do:

    • switch to a verified secondary model;
    • queue non-urgent work;
    • fall back to deterministic templates;
    • preserve human support for sensitive cases;
    • disable autonomous writes while keeping read-only analysis available.

    “Try again until it works” is not a resilience strategy for orders or customer data.

    What it means for AI builders

    For developers and founders, concentration creates risk and opportunity at the same time.

    The risk: your product becomes a thin wrapper

    If your product is only a prompt plus one API call, the provider can reproduce it, a competitor can copy it, and pricing changes can erase its margin.

    The strongest moat usually sits elsewhere:

    • proprietary workflow data;
    • domain-specific evaluation;
    • integrations that are difficult to maintain;
    • governance and approval controls;
    • customer-specific configuration;
    • reliable structured outputs;
    • auditability;
    • user experience and distribution;
    • measurable business outcomes.

    The opportunity: become the independence layer

    Concentration increases demand for products that help businesses use leading models without becoming trapped by them.

    Potential opportunities include:

    • model routing based on quality, cost and latency;
    • portable prompt and policy management;
    • cross-model evaluation suites;
    • provider-neutral agent tooling;
    • caching and cost controls;
    • observability across model vendors;
    • data-loss prevention and access governance;
    • fallbacks for regulated or regional workloads;
    • migration testing when models are retired.

    HelpingBrains’ AI Governance Platform is aimed at this control layer: AI inventory, prompt governance, access monitoring, risk management and audit-ready reporting should remain consistent even when the underlying model changes.

    The opportunity hidden inside a two-horse race

    Market concentration is not automatically bad for customers.

    Two strong providers can:

    • compete aggressively on model quality;
    • reduce prices through efficiency gains;
    • standardise tool-use patterns;
    • accelerate enterprise features;
    • make advanced capabilities accessible without infrastructure investment;
    • create a large ecosystem for specialised products.

    Competition between OpenAI and Anthropic may also prevent either from exercising complete control. Google, Meta, xAI, specialist providers and open-weight models add further pressure even if they are smaller in a particular revenue dataset.

    The opportunity for businesses is to use the leading platforms while retaining the ability to move.

    This is similar to a sound cloud strategy: “multi-cloud” should not mean duplicating everything across three providers at enormous cost. It should mean identifying critical dependencies, using portable interfaces where practical and maintaining tested alternatives for the failures that matter.

    A practical AI diversification checklist

    You do not need to abandon OpenAI or Anthropic. You need to know what would break if one disappeared from your stack tomorrow.

    1. Map every dependency

    • ☐ List every model, API, assistant, agent and AI-enabled SaaS product in use.
    • ☐ Record which business workflow each one supports.
    • ☐ Identify the provider behind tools that resell or abstract another model.
    • ☐ Mark workflows that affect customers, revenue, production data or legal obligations.
    • ☐ Assign one accountable owner to every critical AI system.

    2. Separate your assets from the provider

    • ☐ Store prompts, policies and templates in a controlled repository.
    • ☐ Keep product facts, brand rules and customer permissions in your own systems.
    • ☐ Export conversation or workflow data where contractually and technically possible.
    • ☐ Avoid provider-specific data formats unless the benefit clearly exceeds the switching cost.
    • ☐ Document how model output is transformed before it reaches customers or production.

    3. Build a model-independent boundary

    • ☐ Use a stable internal request and response schema.
    • ☐ Isolate provider-specific code behind adapters.
    • ☐ Validate structured output rather than trusting free text.
    • ☐ Enforce permissions, budgets, privacy rules and prohibited actions outside the model.
    • ☐ Log model, version, prompt, tool calls, latency, cost and outcome.

    4. Test at least one alternative

    • ☐ Maintain a representative evaluation set using real—but sanitised—business cases.
    • ☐ Compare quality, cost, latency, refusals and structured-output reliability.
    • ☐ Test a secondary hosted model or a suitable open-weight alternative.
    • ☐ Measure migration effort rather than assuming APIs are interchangeable.
    • ☐ Repeat tests after major model releases.

    5. Plan the failure mode

    • ☐ Define when to switch providers automatically and when to require human review.
    • ☐ Queue non-critical work instead of producing lower-quality customer-facing output.
    • ☐ Keep deterministic templates for essential communications.
    • ☐ Prevent retries from duplicating sends, refunds, catalogue edits or orders.
    • ☐ Run a provider-outage exercise and record the recovery time.

    This is the save-worthy part of the story: diversification is not buying two subscriptions. It is making your data, controls and workflows portable enough that a second provider can actually take over.

    Should you use both OpenAI and Anthropic?

    Not automatically.

    A small company may create more complexity than resilience by integrating multiple providers too early. Every additional model introduces another contract, privacy review, evaluation surface and operational path.

    Use a second provider when at least one of these is true:

    • the workflow is important enough that an outage creates material loss;
    • model pricing represents a significant share of your unit cost;
    • customers require regional or provider choice;
    • one provider frequently refuses or performs poorly on essential tasks;
    • a model retirement would require a rushed migration;
    • your product promises provider-independent results;
    • regulation, procurement or data residency makes one provider insufficient.

    For low-risk experimentation, one provider plus good abstraction may be enough. For revenue-critical execution, a tested fallback becomes much more valuable.

    The real lesson: concentration belongs on your risk register

    The viral 70% claim is too broad. OpenAI and Anthropic do not demonstrably receive 70% of all revenue across the global AI economy.

    What the available estimates suggest is still significant: a very large share of the AI-related revenue credited to Amazon, Microsoft and Google may depend directly or indirectly on two foundation-model companies.

    That concentration does not mean businesses should stop building. It means they should stop confusing easy access with independence.

    Use the best model available for the job. But keep control of:

    • your data;
    • your prompts and policies;
    • your customer relationships;
    • your business rules;
    • your evaluation criteria;
    • your permission boundaries;
    • your fallback plan.

    The winners will not necessarily be the companies that predict which AI laboratory wins the race. They will be the ones that create value above the model layer—and can keep operating regardless of who is leading next year.


    Frequently asked questions

    Do OpenAI and Anthropic earn 70% of all AI revenue?

    There is no public, audited dataset proving that they receive 70% of revenue across the entire global AI industry. The viral figure is based on analyst estimates of AI-related revenue at Amazon, Microsoft and Google, including compute purchased by OpenAI and Anthropic and cloud platforms reselling access to their models.

    Why is AI market concentration a risk for businesses?

    Heavy reliance on one or two providers exposes businesses to price changes, outages, model retirements, policy changes, regulatory restrictions and strategic competition. The risk is highest when core data, prompts and workflows cannot move to another provider.

    Is a multi-model strategy always better?

    No. Multiple providers add cost and complexity. The right approach is proportional: abstract critical integrations, maintain evaluation tests and create a verified fallback for workflows where downtime or forced migration would cause material harm.

    How can ecommerce companies avoid AI vendor lock-in?

    Keep catalogue data and business rules in company-controlled systems, use stable internal schemas, separate provider-specific code, enforce permissions outside the model and test the same workflow against at least one alternative model.

    What creates a defensible AI product if the models are commoditised?

    Defensibility usually comes from proprietary data, domain workflow, evaluation, integrations, governance, user experience, distribution and measurable outcomes—not exclusive access to a general-purpose model.