Tag: OpenAI

  • OpenAI’s Jalapeño Chip Beat Nvidia in Key Tests. Will It Cut Your AI Bill?

    OpenAI’s Jalapeño Chip Beat Nvidia in Key Tests. Will It Cut Your AI Bill?

    For most business owners, a new AI chip sounds like somebody else’s problem.

    Data-centre racks, memory bandwidth and tokens per watt belong in engineering presentations—not an ecommerce budget meeting.

    But OpenAI’s first custom AI chip could eventually affect something much more familiar: how much you pay every time an AI model answers a customer, writes a product description or runs another step in an automated workflow.

    The chip is called Jalapeño. OpenAI designed it with Broadcom specifically for inference—the work that happens after a model has been trained, when it responds to prompts and performs real tasks.

    OpenAI has now published its first performance results. On three large public models, Jalapeño reportedly delivered:

    • 1.5 to 1.9 times more AI work per watt at peak throughput
    • 1.7 to 3.6 times lower end-to-end latency
    • 2.1 to 4.1 times higher performance for highly interactive workloads

    The comparison systems used Nvidia’s GB200 or GB300 accelerators, depending on the model.

    That is an impressive debut. It is also the kind of benchmark that generates overheated headlines: OpenAI beat Nvidia. Nvidia’s moat is gone. AI is about to become cheap.

    The reality is more interesting—and more useful for builders and ecommerce teams.

    Fact check: is the Jalapeño story true?

    Yes, with important qualifications.

    OpenAI and Broadcom unveiled Jalapeño in June 2026. Engineering samples are running workloads in OpenAI’s labs, and OpenAI plans a limited deployment by the end of 2026 before increasing volume in 2027.

    On 25 August, OpenAI released detailed results using InferenceX, a public inference benchmark developed by semiconductor research company SemiAnalysis.

    OpenAI tested three open-weight models:

    • GPT‑OSS 120B
    • DeepSeek R1 670B
    • Kimi K2.5 1T

    It reported better combinations of throughput, power efficiency and latency than the best comparison results available on Nvidia GB200 or GB300 systems.

    For GPT‑OSS 120B, OpenAI reported approximately 1.9 times higher peak throughput per kilowatt and 1.7 times lower end-to-end latency than the GB200 comparison.

    For DeepSeek R1, it reported approximately 1.7 times higher peak throughput per kilowatt and 3.6 times lower latency than the GB300 comparison.

    For Kimi K2.5, it reported approximately 1.5 times higher peak throughput per kilowatt and 3.4 times lower latency than the GB300 comparison.

    So “Jalapeño beat Blackwell in key AI workloads” is defensible.

    But “OpenAI has built a better chip than Nvidia” is too broad.

    What the benchmark does—and does not—prove

    First, these are OpenAI-published results. The benchmark framework is public, but OpenAI selected the systems, models, configurations and operating points it presented. Independent operators still need to reproduce the results at scale.

    Second, Jalapeño is an inference accelerator, not a training chip. It can run trained models, but it is not currently positioned to replace the Nvidia systems used to train the largest frontier models.

    Third, OpenAI compared complete serving outcomes under specific workloads—not every type of AI computation. Different models, batch sizes, context lengths, software stacks and reliability requirements can change the economics.

    Fourth, engineering samples succeeding in a lab is not the same as thousands of racks operating reliably across data centres. Yield, manufacturing capacity, networking, cooling, software maturity and uptime will determine whether theoretical gains survive production.

    Finally, OpenAI has explicitly said it will continue deploying Nvidia and other partners’ accelerators for training and inference.

    This is not a clean replacement story. It is a diversification story.

    Why OpenAI built its own chip

    OpenAI consumes an extraordinary amount of computing power. Every additional ChatGPT user, API request, coding-agent step and generated token creates inference demand.

    Buying more general-purpose GPUs solves part of the problem, but it leaves OpenAI exposed to three constraints.

    Cost

    High-end AI systems are expensive to buy, power and cool. Nvidia can command strong margins because demand remains high and credible alternatives are limited.

    Supply

    Even a company willing to spend billions cannot instantly obtain unlimited GPUs, high-bandwidth memory, networking equipment and data-centre capacity.

    Control

    Nvidia must build hardware that serves many customers and workloads. OpenAI knows the models, kernels, serving patterns and product roadmap inside its own environment. It can optimise the chip, memory, networking and software around those requirements.

    That is the full-stack advantage.

    Apple designs chips around its devices and operating systems. Google designs TPUs around its AI infrastructure. Amazon builds Trainium and Inferentia for AWS workloads. OpenAI is now applying the same logic to ChatGPT, Codex, the API and future agents.

    Why inference matters more to your AI bill

    Training a frontier model may cost billions, but training is occasional. Inference happens every time somebody uses the model.

    For an ecommerce business, inference includes:

    • A chatbot answering a delivery question
    • A search assistant interpreting customer intent
    • A model translating a product page
    • An agent checking inventory and creating a support response
    • A marketing tool generating campaign variations
    • A recommendation system explaining why a product fits
    • A coding agent maintaining storefront integrations

    One request may be cheap. Millions of requests—and agents performing dozens or hundreds of sequential steps—are not.

    That is why Jalapeño’s combination of lower latency and more work per watt matters. If OpenAI can complete more useful inference with the same electricity and infrastructure, its cost per successful task can fall.

    The business question is what OpenAI does with that saving.

    Will OpenAI actually lower API prices?

    Possibly, but there is no guarantee.

    Lower infrastructure cost gives OpenAI several choices:

    1. Reduce prices to win more API volume.
    2. Keep prices stable and improve margins.
    3. Offer faster service tiers at existing or higher prices.
    4. Spend the efficiency gain on more capable models that use more computation per answer.
    5. Increase limits and availability rather than changing the headline price.

    Technology history shows that efficiency does not always create a smaller bill. Sometimes it creates more usage.

    If an AI agent becomes twice as cheap per step, companies may allow it to perform ten times as many steps. The unit price falls while the total invoice rises.

    So do not assume “better chip” automatically means “lower monthly spend.”

    The more realistic near-term benefits could be:

    • Faster responses
    • More responsive agents
    • Better availability during demand spikes
    • Higher rate limits
    • Cheaper high-speed inference tiers
    • More competition between model providers

    The important ecommerce angle: cost per outcome

    Most teams look at cost per token. That number is becoming less useful as AI workflows grow more complex.

    A cheap model that needs repeated corrections, extra tool calls and human cleanup may cost more than an expensive model that completes the task correctly once.

    For ecommerce, measure:

    • Cost per customer issue resolved
    • Cost per approved product description
    • Cost per translated and reviewed product page
    • Cost per merchandising decision
    • Cost per successful search session
    • Cost per campaign variation that passes brand review

    Jalapeño is designed around interactive inference and agents, where delays accumulate across sequential steps. If the chip genuinely lowers latency while preserving throughput, agents can finish multi-step tasks faster and infrastructure can serve more concurrent customers.

    That can improve cost per outcome even if the published API price barely changes.

    Does this threaten Nvidia’s moat?

    Yes—but only one layer of it.

    Nvidia’s moat is not merely a fast chip. It includes:

    • CUDA and a mature software ecosystem
    • Developer familiarity
    • Libraries and optimisation tools
    • High-performance networking
    • Complete rack-scale systems
    • Strong relationships with cloud providers
    • Manufacturing scale and a rapid product roadmap
    • Hardware capable of both training and inference workloads

    OpenAI has demonstrated that a major model company can outperform Nvidia systems on selected inference workloads by co-designing hardware and software.

    That weakens the idea that every valuable AI workload must run most efficiently on Nvidia.

    But Jalapeño is not being sold to developers or cloud customers. OpenAI says it needs the capacity internally and has no plan to commercialise the chip. Nvidia, meanwhile, sells a platform across the industry.

    The immediate threat is therefore not that OpenAI will steal Nvidia’s external chip customers. It is that Nvidia’s largest customers increasingly become their own suppliers for predictable, high-volume workloads.

    Google, Amazon, Microsoft, Meta and now OpenAI are all pursuing custom silicon. Each workload moved onto an internal accelerator reduces dependence on Nvidia at the margin and gives the buyer more negotiating power.

    Nvidia can remain dominant while losing its status as the only serious answer.

    Could OpenAI “beat” Nvidia without selling a single chip?

    Yes—in the area that matters to OpenAI.

    OpenAI does not need to build a better general-purpose accelerator for every customer. It needs to lower the cost and increase the speed of its own enormous inference workload.

    If Jalapeño serves a meaningful percentage of ChatGPT and API demand more efficiently, it succeeds even if Nvidia continues growing.

    This is why the “who wins the chip war?” framing can be misleading. Several companies can win different layers:

    • Nvidia can remain the leading general AI platform.
    • OpenAI can gain better economics for its own products.
    • Broadcom can profit from custom silicon design and networking.
    • TSMC and memory suppliers can benefit regardless of whose logo is on the accelerator.
    • Customers can benefit from increased competition and more available compute.

    What this means for builders

    Jalapeño will not appear as a chip option in your cloud account. OpenAI plans to use it internally.

    For API builders, its impact will appear indirectly through product behaviour:

    • Model pricing
    • Latency
    • Rate limits
    • Service reliability
    • Context-window economics
    • Batch discounts
    • Agent pricing
    • Availability of high-speed modes

    The strategic lesson is not to wait for OpenAI to reduce prices. Build your application so it can benefit when competition changes.

    Keep model routing flexible

    Avoid hard-wiring every workflow to one model. Different tasks may be cheaper or better on OpenAI, Anthropic, Google, open models or specialised providers.

    Measure the whole task

    Track retries, tool calls, latency and human review—not just input and output tokens.

    Use smaller models where they work

    Product classification, simple enrichment and structured extraction often do not require the most capable frontier model.

    Cache predictable outputs

    Do not repeatedly pay a model to generate information that changes rarely, such as stable category descriptions or standard policy explanations.

    Negotiate as volume grows

    Public pricing is not necessarily the final price for a large, predictable workload. Custom hardware could give providers more room to offer committed-use pricing.

    What ecommerce and marketing leaders should watch

    You do not need to follow chip specifications every week. Watch the signals that can reach your budget.

    1. Real API price changes

    Look for lower prices on inference-heavy models, batch processing or high-speed modes. A benchmark is not a discount until it appears in your invoice.

    2. Latency under real load

    Test your own prompts and workflows. Faster tokens are valuable only if end-to-end customer experiences improve.

    3. Agent pricing

    Agents can multiply inference usage because they plan, call tools, inspect results and retry. Watch whether providers price by token, task, tool call or successful outcome.

    4. Reliability and capacity

    Higher efficiency may first appear as fewer rate-limit errors and better availability rather than cheaper tokens.

    5. Competitive responses

    Nvidia will continue improving its hardware and software. Google, Amazon, Microsoft, AMD and specialised inference companies will not stand still. The broader price effect comes from competition, not one chip.

    6. Whether OpenAI deploys Jalapeño at meaningful scale

    Small volumes prove the concept. Gigawatt-scale production determines the economics.

    What should businesses do today?

    Do not rewrite your AI strategy because of one benchmark release.

    Instead:

    • Record your current AI cost per business outcome.
    • Identify which workflows are latency-sensitive.
    • Separate high-value reasoning from bulk repetitive inference.
    • Keep provider switching technically possible.
    • Re-test price and performance quarterly.
    • Avoid long contracts based solely on promised hardware efficiency.

    The biggest mistake would be assuming AI prices only move downward. Providers may use cheaper inference to make models perform more work, which can deliver more value while leaving your total spend unchanged—or higher.

    Your job is to ensure the additional computation produces an additional business result.

    Final thought

    Jalapeño does not dethrone Nvidia.

    It does something more strategically important: it proves that the company building the model can also redesign the infrastructure beneath it and win on selected workloads.

    That puts pressure on Nvidia, improves OpenAI’s negotiating position and creates another path toward cheaper and faster inference.

    For ecommerce brands and builders, the opportunity is real—but indirect.

    Do not watch the chip race simply to learn who has the fastest processor.

    Watch for what reaches your product:

    lower cost per successful task, faster customer experiences and enough competition to prevent one supplier from setting the price of intelligence.

    Frequently asked questions

    What is OpenAI’s Jalapeño chip?

    Jalapeño is OpenAI’s first custom AI inference accelerator, co-developed with Broadcom. It is designed to run large language models and interactive AI agents rather than train frontier models.

    Did Jalapeño beat Nvidia Blackwell?

    In OpenAI-published InferenceX results, Jalapeño delivered higher performance per watt and lower latency than comparison systems using Nvidia GB200 or GB300 accelerators across GPT‑OSS 120B, DeepSeek R1 and Kimi K2.5. This does not establish superiority across all models or workloads.

    Is Jalapeño available to developers?

    No. OpenAI plans to deploy it inside its own infrastructure and says it has no current plan to sell the chip externally.

    Will the new chip make the OpenAI API cheaper?

    It could reduce OpenAI’s inference costs, but OpenAI has not guaranteed a direct price reduction. Efficiency may appear through lower prices, faster tiers, higher limits, better availability or more computation per response.

    Does this mean Nvidia is losing its AI leadership?

    Not yet. Nvidia retains major advantages in training, software, networking, scale and broad availability. Jalapeño shows that custom chips can challenge Nvidia on specialised, high-volume inference workloads.

    Why should ecommerce brands care about inference chips?

    Ecommerce AI workloads—customer support, search, translation, content enrichment, recommendations and agents—are primarily inference. More efficient inference can improve response speed, capacity and ultimately cost per completed business task.

    Sources

  • Do OpenAI and Anthropic Really Drive 70% of AI Revenue? What It Means for Your Business

    Do OpenAI and Anthropic Really Drive 70% of AI Revenue? What It Means for Your Business

    One statistic is racing around the AI industry:

    More than 70% of AI revenue comes from OpenAI and Anthropic.

    It is a powerful number. It suggests that thousands of AI products, billions in infrastructure spending and the strategies of the world’s largest technology companies rest on two model providers.

    It is also easy to repeat incorrectly.

    The available evidence does not establish that OpenAI and Anthropic collect 70% of every dollar earned across the entire AI market. The figure comes from analyst estimates about a narrower—and in some ways more revealing—part of the ecosystem: the AI-related revenue earned by Amazon, Microsoft and Google from cloud compute, model access and associated commercial arrangements.

    In other words, the claim is less “70% of AI revenue flows to two companies” and more:

    Analysts estimate that OpenAI and Anthropic may directly or indirectly drive more than 70% of the AI-related revenue attributed to the three largest US cloud platforms.

    Some of that money flows from OpenAI and Anthropic to cloud providers for compute. Some comes from cloud customers buying access to their models. The exact totals are not disclosed cleanly by the companies and the estimates differ by analyst.

    That correction weakens the viral headline—but strengthens the business lesson.

    The AI economy may be broader than two companies. The commercial infrastructure underneath it is still remarkably concentrated. If your marketing workflow, ecommerce operation or software product depends on one foundation-model provider, you are not merely choosing a tool. You are inheriting that provider’s pricing, availability, policy and strategic risk.

    Fact check: what does the 70% figure actually measure?

    The source behind the current discussion is Ed Zitron’s analysis, “The AI Demand Bubble”. It combines estimates attributed to analysts at Barclays, UBS and Wells Fargo.

    The estimates cited in that analysis include:

    • Amazon Web Services: One Barclays estimate put OpenAI and Anthropic at 73% of AWS AI revenue in 2026 and 2027. A separate estimate cited in the same article put their direct compute contribution at 59% in 2026.
    • Google Cloud: UBS estimates cited in the analysis assigned 28% of total 2026 Google Cloud revenue to OpenAI and Anthropic, rising to more than 48% in 2027. The author then inferred that the pair could represent at least 70% of Google’s narrower AI-related revenue.
    • Microsoft: Wells Fargo estimates cited in the article put OpenAI and Anthropic at 70% or more of Microsoft’s AI revenue, reaching approximately 74% in the relevant forecast period.

    These are not one consistent, audited market-share dataset. They mix:

    • direct purchases of computing capacity;
    • cloud revenue associated with the two laboratories;
    • revenue from platforms that resell access to their models;
    • analyst forecasts for future periods;
    • the article author’s own classification and inference.

    The cloud companies do not publish a standard “AI revenue” line that allows outsiders to calculate a definitive global market share. OpenAI and Anthropic are also private companies, so their financial disclosure is more limited than that of a public company.

    The accurate version of the claim

    Use this formulation:

    Analyst estimates suggest OpenAI and Anthropic account for roughly 70% or more of the AI-related revenue attributed to Amazon, Microsoft and Google, although definitions and estimates vary.

    Avoid this formulation:

    OpenAI and Anthropic receive 70% of all revenue in the global AI industry.

    That broader statement would require a defined market covering chips, cloud infrastructure, enterprise software, consumer subscriptions, services, advertising, robotics, data platforms and other AI-related businesses. The cited analysis does not provide that calculation.

    Why the corrected number still matters

    The statistic is not a clean measure of the whole AI market, but it exposes three forms of concentration.

    1. Demand concentration

    Cloud providers have invested extraordinary amounts in data centres, accelerators and power capacity. If a large share of the associated revenue depends on two customers and their models, the infrastructure boom has a narrower demand base than the headline “AI adoption” numbers imply.

    This does not prove the boom will collapse. It means future growth depends heavily on OpenAI and Anthropic continuing to:

    • attract paying customers;
    • raise or generate enough cash to fund compute;
    • turn model capability into sustainable demand;
    • serve workloads efficiently enough to support margins;
    • maintain favourable relationships with cloud partners.

    2. Model concentration

    Many applications are not independent AI businesses in a technical sense. They are interfaces, workflows or specialised data layers built on a small number of foundation models.

    That can be a perfectly good business. Shopify did not need to build a payment network, and SaaS companies do not manufacture their own processors. Specialisation creates value.

    The risk begins when the application has no meaningful advantage beyond one provider’s output and cannot operate if that provider changes.

    3. Strategic concentration

    OpenAI and Anthropic influence more than model quality. Their decisions can shape:

    • token and subscription prices;
    • API limits and access tiers;
    • context-window and tool-use behaviour;
    • model retirement schedules;
    • safety policies and refused use cases;
    • data-processing terms;
    • regional availability;
    • integration standards;
    • which workflows become economically viable.

    A business built on top of one provider may experience these decisions as product changes—even when it had no voice in making them.

    What happens if one of the two stumbles?

    “Stumbles” does not have to mean bankruptcy. For a customer, smaller changes can create the same operational effect.

    Prices rise

    If inference pricing increases or a subsidised product becomes more expensive, an application with weak margins may become uneconomic overnight.

    This is especially dangerous when the company offers customers a fixed monthly price while paying the model provider per token, image, tool call or unit of compute.

    A model or feature is retired

    Prompts tuned for one model do not automatically behave the same on its replacement. Output format, tone, refusal patterns, latency and tool selection can all change.

    Without regression tests, a “simple upgrade” can quietly damage product listings, customer replies, campaign copy or structured data.

    Reliability declines

    An outage at the foundation-model layer can stop every workflow built above it. Even partial degradation—higher latency, elevated errors or inconsistent tool calls—can create queues, duplicate actions and failed customer experiences.

    Regulation or litigation changes access

    New regulatory restrictions, court decisions, government procurement rules or regional compliance requirements can affect how models are offered and which data may be processed.

    The correct response is not to predict one dramatic ban. It is to ensure that a single legal or policy change cannot disable an essential workflow without an alternative.

    Provider strategy shifts

    A model company can enter your category directly, prioritise enterprise contracts, discontinue a partner feature or bundle functionality that makes your product less differentiated.

    Platform risk is not only technical. Your supplier can become your competitor.

    What AI market concentration means for marketers

    For marketers, the immediate temptation is to treat model choice as a creative preference: Which assistant writes the strongest hooks? Which one follows brand voice best? Which one creates the most attractive images?

    Those questions matter, but operational dependence matters more.

    A marketing stack may use one provider for:

    • campaign research;
    • segmentation ideas;
    • advertisement variations;
    • product copy;
    • email personalisation;
    • image generation;
    • social scheduling;
    • performance analysis.

    If every step depends on one vendor, a policy update or service interruption can stop the entire content pipeline.

    The better approach is to distinguish between creative preference and business-critical dependency.

    • It is reasonable to prefer one model for campaign concepts.
    • It is risky if no other model can render the required data structure.
    • It is reasonable to use one assistant for drafts.
    • It is risky if brand knowledge exists only inside that provider’s proprietary workspace.
    • It is reasonable to optimise prompts for quality.
    • It is risky if no regression suite tells you when an update changes the output.

    Keep brand guidelines, approved claims, product facts, audience definitions and reusable prompt templates in systems you control. The model should consume your marketing intelligence, not become the only place where it exists.

    What it means for ecommerce brands

    Ecommerce companies have a deeper dependency problem because AI is moving from content generation into operational action.

    Models increasingly help with:

    • catalogue enrichment;
    • onsite search and recommendations;
    • customer-service responses;
    • translations;
    • merchandising analysis;
    • campaign creation;
    • pricing recommendations;
    • returns and order workflows.

    The closer AI gets to customers, orders and money, the more expensive provider concentration becomes.

    Your catalogue must remain the source of truth

    Store product attributes, claims, translations and policy rules in your own product-information or commerce systems. Do not let one model’s memory or proprietary knowledge feature become the authoritative record.

    If you are evaluating content tools, the criteria in our AI tools for ecommerce product listings benchmark remain useful: accuracy, structured output, brand consistency, channel adaptation and measurable workflow performance matter more than a flashy one-off result.

    Separate recommendations from execution

    An alternative model can replace a copywriting assistant relatively easily. Replacing an autonomous agent that can change prices, send campaigns or issue refunds is much harder.

    Our analysis of why humans missed one in three dangerous AI agent commands explains why manual approval alone is insufficient. Permissions, spending limits, audit trails and rollback must be enforced outside the model.

    Design graceful degradation

    If the preferred model is unavailable, decide what the store should do:

    • switch to a verified secondary model;
    • queue non-urgent work;
    • fall back to deterministic templates;
    • preserve human support for sensitive cases;
    • disable autonomous writes while keeping read-only analysis available.

    “Try again until it works” is not a resilience strategy for orders or customer data.

    What it means for AI builders

    For developers and founders, concentration creates risk and opportunity at the same time.

    The risk: your product becomes a thin wrapper

    If your product is only a prompt plus one API call, the provider can reproduce it, a competitor can copy it, and pricing changes can erase its margin.

    The strongest moat usually sits elsewhere:

    • proprietary workflow data;
    • domain-specific evaluation;
    • integrations that are difficult to maintain;
    • governance and approval controls;
    • customer-specific configuration;
    • reliable structured outputs;
    • auditability;
    • user experience and distribution;
    • measurable business outcomes.

    The opportunity: become the independence layer

    Concentration increases demand for products that help businesses use leading models without becoming trapped by them.

    Potential opportunities include:

    • model routing based on quality, cost and latency;
    • portable prompt and policy management;
    • cross-model evaluation suites;
    • provider-neutral agent tooling;
    • caching and cost controls;
    • observability across model vendors;
    • data-loss prevention and access governance;
    • fallbacks for regulated or regional workloads;
    • migration testing when models are retired.

    HelpingBrains’ AI Governance Platform is aimed at this control layer: AI inventory, prompt governance, access monitoring, risk management and audit-ready reporting should remain consistent even when the underlying model changes.

    The opportunity hidden inside a two-horse race

    Market concentration is not automatically bad for customers.

    Two strong providers can:

    • compete aggressively on model quality;
    • reduce prices through efficiency gains;
    • standardise tool-use patterns;
    • accelerate enterprise features;
    • make advanced capabilities accessible without infrastructure investment;
    • create a large ecosystem for specialised products.

    Competition between OpenAI and Anthropic may also prevent either from exercising complete control. Google, Meta, xAI, specialist providers and open-weight models add further pressure even if they are smaller in a particular revenue dataset.

    The opportunity for businesses is to use the leading platforms while retaining the ability to move.

    This is similar to a sound cloud strategy: “multi-cloud” should not mean duplicating everything across three providers at enormous cost. It should mean identifying critical dependencies, using portable interfaces where practical and maintaining tested alternatives for the failures that matter.

    A practical AI diversification checklist

    You do not need to abandon OpenAI or Anthropic. You need to know what would break if one disappeared from your stack tomorrow.

    1. Map every dependency

    • ☐ List every model, API, assistant, agent and AI-enabled SaaS product in use.
    • ☐ Record which business workflow each one supports.
    • ☐ Identify the provider behind tools that resell or abstract another model.
    • ☐ Mark workflows that affect customers, revenue, production data or legal obligations.
    • ☐ Assign one accountable owner to every critical AI system.

    2. Separate your assets from the provider

    • ☐ Store prompts, policies and templates in a controlled repository.
    • ☐ Keep product facts, brand rules and customer permissions in your own systems.
    • ☐ Export conversation or workflow data where contractually and technically possible.
    • ☐ Avoid provider-specific data formats unless the benefit clearly exceeds the switching cost.
    • ☐ Document how model output is transformed before it reaches customers or production.

    3. Build a model-independent boundary

    • ☐ Use a stable internal request and response schema.
    • ☐ Isolate provider-specific code behind adapters.
    • ☐ Validate structured output rather than trusting free text.
    • ☐ Enforce permissions, budgets, privacy rules and prohibited actions outside the model.
    • ☐ Log model, version, prompt, tool calls, latency, cost and outcome.

    4. Test at least one alternative

    • ☐ Maintain a representative evaluation set using real—but sanitised—business cases.
    • ☐ Compare quality, cost, latency, refusals and structured-output reliability.
    • ☐ Test a secondary hosted model or a suitable open-weight alternative.
    • ☐ Measure migration effort rather than assuming APIs are interchangeable.
    • ☐ Repeat tests after major model releases.

    5. Plan the failure mode

    • ☐ Define when to switch providers automatically and when to require human review.
    • ☐ Queue non-critical work instead of producing lower-quality customer-facing output.
    • ☐ Keep deterministic templates for essential communications.
    • ☐ Prevent retries from duplicating sends, refunds, catalogue edits or orders.
    • ☐ Run a provider-outage exercise and record the recovery time.

    This is the save-worthy part of the story: diversification is not buying two subscriptions. It is making your data, controls and workflows portable enough that a second provider can actually take over.

    Should you use both OpenAI and Anthropic?

    Not automatically.

    A small company may create more complexity than resilience by integrating multiple providers too early. Every additional model introduces another contract, privacy review, evaluation surface and operational path.

    Use a second provider when at least one of these is true:

    • the workflow is important enough that an outage creates material loss;
    • model pricing represents a significant share of your unit cost;
    • customers require regional or provider choice;
    • one provider frequently refuses or performs poorly on essential tasks;
    • a model retirement would require a rushed migration;
    • your product promises provider-independent results;
    • regulation, procurement or data residency makes one provider insufficient.

    For low-risk experimentation, one provider plus good abstraction may be enough. For revenue-critical execution, a tested fallback becomes much more valuable.

    The real lesson: concentration belongs on your risk register

    The viral 70% claim is too broad. OpenAI and Anthropic do not demonstrably receive 70% of all revenue across the global AI economy.

    What the available estimates suggest is still significant: a very large share of the AI-related revenue credited to Amazon, Microsoft and Google may depend directly or indirectly on two foundation-model companies.

    That concentration does not mean businesses should stop building. It means they should stop confusing easy access with independence.

    Use the best model available for the job. But keep control of:

    • your data;
    • your prompts and policies;
    • your customer relationships;
    • your business rules;
    • your evaluation criteria;
    • your permission boundaries;
    • your fallback plan.

    The winners will not necessarily be the companies that predict which AI laboratory wins the race. They will be the ones that create value above the model layer—and can keep operating regardless of who is leading next year.


    Frequently asked questions

    Do OpenAI and Anthropic earn 70% of all AI revenue?

    There is no public, audited dataset proving that they receive 70% of revenue across the entire global AI industry. The viral figure is based on analyst estimates of AI-related revenue at Amazon, Microsoft and Google, including compute purchased by OpenAI and Anthropic and cloud platforms reselling access to their models.

    Why is AI market concentration a risk for businesses?

    Heavy reliance on one or two providers exposes businesses to price changes, outages, model retirements, policy changes, regulatory restrictions and strategic competition. The risk is highest when core data, prompts and workflows cannot move to another provider.

    Is a multi-model strategy always better?

    No. Multiple providers add cost and complexity. The right approach is proportional: abstract critical integrations, maintain evaluation tests and create a verified fallback for workflows where downtime or forced migration would cause material harm.

    How can ecommerce companies avoid AI vendor lock-in?

    Keep catalogue data and business rules in company-controlled systems, use stable internal schemas, separate provider-specific code, enforce permissions outside the model and test the same workflow against at least one alternative model.

    What creates a defensible AI product if the models are commoditised?

    Defensibility usually comes from proprietary data, domain workflow, evaluation, integrations, governance, user experience, distribution and measurable outcomes—not exclusive access to a general-purpose model.