
Conceptual illustration of efficient AI computing alongside a modern data center and power infrastructure, highlighting how cheaper AI could still increase demand for the systems that support it. AI-generated image via ChatGPT (OpenAI)
AI Is Getting Cheaper Fast. So Why Could Compute Demand Keep Rising?
OpenAI and Anthropic just cut the cost of using their newest high-end AI models, the latest sign of a much bigger shift: the cost of achieving a given level of AI performance is falling remarkably fast.
For businesses, that could change the practical question from “Can we afford this level of AI?” to “What can we now do with it — and can the infrastructure support that scale?”
OpenAI positions GPT-6 Sol for difficult work, while Luna delivers substantial capability at a much lower cost. Anthropic says Opus 5.5 leads its benchmarks in agentic coding, computer use and knowledge work.
OpenAI cut API prices for its new GPT-6 Sol and Luna models by 50% compared with GPT-5.6 promotional pricing. GPT-6 Sol dropped from $4 to $2 per million input tokens and from $20 to $10 per million output tokens. GPT-6 Luna fell from $0.20 to $0.10 for input and from $1.20 to $0.50 for output. OpenAI says improvements in caching and inference — essentially, finding more efficient ways to process and reuse information — helped bring those costs down.
Anthropic made a similar move with Claude Opus 5.5. The company says the new model costs about 40% less than Opus 5 on typical workloads and requires less computing power to run. Input and output prices are 20% lower, cached information costs 60% less to reuse, and the model generates responses more than 30% faster.
At first glance, this looks like a pricing story: two major AI developers making powerful models cheaper. But the implications reach beyond OpenAI and Anthropic. Falling AI costs could affect businesses deploying AI, the companies building and operating AI infrastructure, and eventually industries using AI in everything from software and research to vehicles, factories and robots.
So why is AI intelligence getting cheaper so quickly? The short answer is that several improvements are happening at once: better algorithms, better chips, more efficient models, smarter inference software and systems that avoid repeating work unnecessarily.
But that creates a second question. If AI is becoming dramatically more efficient, shouldn't we eventually need less compute?
Not necessarily.
AI can use fewer resources to perform a fixed amount of work while total demand still rises if cheaper, more capable AI leads us to use much more of it.
And that is where this story becomes much bigger than two model price cuts.
Key Takeaways: Why AI Is Getting Cheaper — and Why Compute Demand May Still Rise
The cost of achieving a given level of AI performance is falling rapidly as algorithms, hardware, models and software become more efficient. But cheaper AI can also make far more AI use economical, allowing total compute and infrastructure demand to keep growing.
The cost of achieving a given level of AI performance has fallen about 47% per quarter since 2023, equivalent to roughly 13-fold per year, according to Epoch AI.
OpenAI and Anthropic's latest price cuts are part of a longer trend in which AI inference costs have fallen dramatically across general knowledge, coding, science and chatbot tasks.
AI is getting cheaper because better algorithms, more efficient hardware, improved model designs, inference software and caching are reducing the resources needed to produce useful AI work.
More efficient AI does not necessarily reduce total compute demand because people can use more AI, give it more complicated tasks and build workflows that require many steps, tools and model calls.
As cheaper AI makes more applications economical, the limiting factor may increasingly become the physical infrastructure needed to run it, including chips, data centers, electricity, cooling and networking.
It remains unclear whether AI costs can keep falling at their recent pace because the modern generative-AI market has only a few years of comparable history and future efficiency gains will depend on continued investment and innovation.
The Cost of AI Performance Is Falling at an Extraordinary Rate
The OpenAI and Anthropic price cuts are striking on their own. But they are part of a much larger trend.
Epoch AI estimates that since 2023, the cost of achieving a given level of AI performance has fallen by about 47% per quarter — the equivalent of roughly 13 times per year.
That pace is unusual even compared with technologies known for becoming dramatically cheaper over time. According to Epoch, AI's recent rate of price decline has been about four times faster than DNA sequencing, six times faster than computing, 18 times faster than lithium-ion batteries and 54 times faster than electricity over the historical periods it compared.
One example makes the scale easier to understand.
In January 2025, Epoch estimates that OpenAI's o3 could score about 75% on GPQA Diamond, a difficult science benchmark, at a cost of roughly 30 cents per question. Less than 18 months later, GPT-5.6 Luna could reach approximately the same score for about $0.0004 per question — four-hundredths of a cent.
That's a 725-fold decline in the estimated cost of reaching roughly the same level of performance.
There are important limits to what those numbers tell us. AI benchmarks are not the same thing as real-world work, and Epoch's dataset covers only about three years. Different models, tasks and ways of measuring performance also produce different rates of decline, so Epoch cautions that its estimates should be treated as reasonable approximations rather than exact measurements.
Still, the direction is hard to miss: getting a given level of AI performance has become dramatically cheaper.
And this didn't begin with this week's OpenAI and Anthropic announcements.
AI Costs Were Already Falling Before 2026
The rapid drop in AI costs didn't begin with this week's OpenAI and Anthropic launches.
Stanford's 2025 AI Index found that the cost of running a model at roughly GPT-3.5-level performance on a common AI benchmark fell from about $20 per million tokens in November 2022 to just $0.07 by October 2024 — a decline of more than 280-fold in less than two years.
The decline wasn't limited to one type of AI task, although the pace varied considerably. Costs fell across general knowledge, coding, science and chatbot performance, but not at the same speed.
For example, the price of reaching GPT-4-level performance on a coding benchmark fell from about $37.50 per million tokens in March 2023 to $0.10 by July 2024. For GPT-4-level performance on a chatbot benchmark, Epoch's data show the price falling from $15 per million tokens in January 2024 to $0.12 by December 2024.
Across all the tasks and performance levels Epoch studied, annual price declines ranged from about 9-fold to 900-fold, with a median of roughly 50-fold. That enormous range reflects differences in both the task and the level of performance being measured.
AI models were also becoming much smaller while still reaching similar levels of performance.
In 2022, Google's PaLM needed about 540 billion parameters to score above 60% on MMLU, a benchmark covering a wide range of academic subjects. By 2024, Microsoft's Phi-3-mini reached that same threshold with just 3.8 billion parameters.
That's about a 142-fold reduction in model size for roughly comparable performance on that benchmark.
Why does that matter? Because comparable AI performance no longer necessarily requires the enormous model sizes it once did. Smaller, more capable models are part of the same broader shift making advanced AI more efficient, affordable and accessible.
So the latest OpenAI and Anthropic price cuts aren't isolated discounts. They're the newest examples of a cost-and-efficiency trend that has already been moving remarkably quickly for several years.
Why Is AI Getting Cheaper So Quickly?
There isn't one reason. AI is getting cheaper because several parts of the technology are becoming more efficient at the same time.
One piece is better algorithms — essentially, better ways of designing and training AI models so they can get more useful performance from the same amount of computing power.
In a historical analysis of language models, Epoch AI estimated that the amount of training compute needed to reach a given level of performance was cut roughly in half every eight months, although its estimate ranged from about five to 14 months.
But that doesn't mean better algorithms have eliminated the need for more computing power. Quite the opposite.
Epoch estimated that about 60% to 95% of historical performance gains came from increases in computing power and training data, while about 5% to 40% came from new algorithms. In other words, AI has been getting more efficient at the same time that developers have been using vastly more compute.
The hardware itself is also improving.
Stanford's 2025 AI Index found that machine-learning hardware performance increased about 43% per year, while price-performance improved about 30% per year and energy efficiency improved around 40% per year. Put simply, newer hardware has been able to do more AI work for the money and electricity being spent.
Then there is the software used to actually run AI models after they have been trained.
Nvidia says improvements to its inference software — the software that helps models generate answers and perform tasks — reduced the cost per token for DeepSeek V4 on its Blackwell hardware by as much as fivefold in about one month. That's a company claim rather than an independent industry-wide measurement, but it shows how software improvements can squeeze substantially more work out of the same underlying hardware.
And sometimes the cheapest computation is the computation you don't have to repeat.
AI systems often reuse the same instructions, documents or other information across many requests. Caching allows them to reuse work they have already processed instead of starting from scratch every time. OpenAI says cached GPT-6 input tokens receive a 90% discount, and GitHub reports that recent caching improvements reduced the share of prompt tokens needing fresh processing by more than 50% across billions of requests to OpenAI models.
Put all of this together, and the falling price of AI isn't being driven by one breakthrough.
Better algorithms, better hardware, more efficient model designs, smarter inference software, caching and other system improvements are all pushing costs downward — and some of those gains compound on top of one another.
But there is an important catch: AI has been getting more efficient while the industry has continued building and using more computing power.
Which brings us back to the central question: If AI needs fewer resources to perform a given amount of work, shouldn't we eventually need less compute?
Why Cheaper AI May Still Require More Compute
At first, this seems contradictory.
If AI models can do the same work with less computing power, shouldn't all these efficiency gains eventually mean we need less compute?
For a fixed amount of AI work, yes. If a model becomes more efficient, it can use fewer resources to perform the same task.
But that's only half of the equation.
The amount of AI work being done is also growing rapidly. Epoch AI reported that inference revenue — money companies earn from customers actually using AI models — at major developers including OpenAI and Anthropic had been growing threefold per year or more, even as their models became smaller and cheaper compared with 2023.
And what counts as an AI request is changing.
A few years ago, a user might ask a chatbot a question and get an answer. Increasingly, AI models are being asked to reason through complicated problems, use tools, search through information and work through multiple steps before producing a result. Epoch notes that models are now reasoning for longer periods and being used inside increasingly elaborate AI-agent systems.
OpenAI makes a similar distinction. A quick answer may require very little compute, while a coding, research or financial task can involve deeper reasoning, tool use and many separate actions — requiring considerably more computing power to complete the work.
At the same time, greater efficiency makes existing AI tasks cheaper, while greater capability makes entirely new kinds of work possible.
That helps explain the apparent contradiction.
AI can use less compute to perform a fixed amount of work while total compute demand still rises because we are asking AI to do more work, more complicated work and entirely new kinds of work.
That doesn't mean cheaper AI has been proven to cause all of the growth in compute demand. But falling costs and rapidly growing AI use are already happening at the same time.
And as the price of using AI continues to fall, more tasks that were once too expensive to justify may become economical.
One AI Request Can Now Trigger Far More Computing Work
The way people use AI is also changing what a single request can require.
With traditional chatbots, a request might be relatively simple: ask a question, the model processes it and returns an answer. AI agents can work very differently. They may break a larger goal into multiple steps, call outside tools, retrieve information, use memory, run different models and create additional AI agents to handle parts of the job.
Nvidia says some agent systems can turn a single request into hundreds of specialized subagents and thousands of individual tasks, involving multiple AI models and different types of computing and storage. That's Nvidia's description of emerging agent infrastructure, not a measure of how every AI agent operates. But it illustrates an important change: one request no longer necessarily means one model producing one answer.
So even as each individual piece of AI work becomes cheaper and more efficient, the amount of computing happening behind a single user request can grow substantially.
Why Falling AI Costs Can Drive More Compute Investment
Falling AI costs don't necessarily reduce the incentive to build more infrastructure. They may do the opposite.
OpenAI argues that AI development works as a compounding cycle: better infrastructure helps accelerate research; that research produces more capable and efficient models; better models improve products; better products drive more adoption, learning and revenue; and that growth can support further investment in research, computing power and deployment. (openai.com)
In simple terms, the cycle can look like this:
More compute → better models → cheaper and more useful AI → more adoption → more reason to invest in compute.
That helps explain why rapidly falling AI costs don't necessarily signal the end of infrastructure expansion. If cheaper, more capable AI makes more applications worth using, lower prices can actually support greater demand for the systems needed to run them.
As AI Gets Cheaper, the Bottleneck May Shift to Infrastructure
This creates a strange economic picture: the output of AI systems is getting dramatically cheaper while some of the physical resources needed to produce that output are becoming more expensive.
Epoch AI points to this paradox directly. The AI boom is creating enough demand to push up the price of inputs such as chips and electricity even as the cost of producing a given level of AI performance continues to fall.
That suggests the constraint may be shifting.
If useful AI becomes cheap enough to deploy much more widely, the limiting factor may increasingly be the physical infrastructure required to generate and deliver enormous amounts of it: computing power, chips, data centers, electricity, cooling, networking and the capital needed to build it all.
In other words, making AI itself cheaper doesn't necessarily make the entire system behind it cheaper — especially if lower costs lead to much more AI use.
Perhaps making AI more efficient doesn’t remove the bottleneck. It moves it — from the cost of producing intelligence toward the infrastructure required to generate intelligence at enormous scale.
Can AI Costs Keep Falling This Fast?
The honest answer is: we don't know.
The recent decline in AI costs has been extraordinary, but there simply isn't much history to tell us how long a trend like this can continue.
Epoch AI's estimates cover only about three years — not because there is a long history being ignored, but because today's market for widely used generative AI models is still very young. That makes it harder to know whether the rapid declines we've seen so far represent a lasting trend or an unusually fast early period.
There are also limits to the measurements themselves. AI benchmarks are useful for comparing performance over time, but they don't perfectly capture the value or difficulty of real-world work. Different tasks and different performance thresholds can also produce very different rates of cost decline.
And there is no guarantee that future efficiency gains will arrive at the same pace as past ones. In separate research on algorithmic progress, Epoch said continued improvements will depend on factors including investment, computing power and further innovation.
So we know the historical direction: achieving a given level of AI performance has become dramatically cheaper in a remarkably short period of time.
What we don't know is how long that pace can continue — or whether AI use will grow so quickly that total demand keeps rising despite those efficiency gains.
What This Means: Cheaper AI Could Expand Use — and Infrastructure Demand
As AI becomes cheaper, applications that once cost too much to justify may suddenly become practical.
That affects more than AI companies. Businesses may be able to automate or expand work that previously wasn't economical. Cloud and data-center providers may face more demand, while utilities may have to support greater electricity demand and infrastructure companies may have to build more capacity. And as AI moves into vehicles, factories, robots and other physical systems, those demands could spread far beyond traditional data centers.
For businesses, the practical question may increasingly shift from:
“Can we afford this level of AI?”
to:
“What can we now do with it — and can the infrastructure support that scale?”
The important trade-off is that greater efficiency at the task level does not necessarily mean lower resource use overall. If AI becomes cheap enough that people use far more of it — and use it for more complicated work — total demand for computing power and energy can still rise.
We also don't know whether AI costs will keep falling at anything close to their recent pace. The historical decline has been extraordinary, but this is still a young market, and future gains will depend on continued improvements in algorithms, hardware, software and infrastructure.
So perhaps making AI more efficient doesn't remove the bottleneck. It moves it — from the cost of producing intelligence toward the infrastructure required to generate intelligence at enormous scale.
And that is where this story connects directly to what happens next.
In a separate AiNews analysis, contributing writer Jim Harris looks at what happens when increasingly inexpensive AI moves beyond the data center and into vehicles, factories, robots and other physical systems — and what that could mean for electricity demand.
If AI keeps getting cheaper, we’re likely to use more of it in more places — and the infrastructure behind it will have to keep up.
Q&A: Why AI Costs Are Falling, Why Compute Demand May Rise, and What It Means for Businesses
Q: What changed with OpenAI and Anthropic's new AI models?
A: OpenAI cut API prices for GPT-6 Sol and Luna by 50% compared with GPT-5.6 promotional pricing, while Anthropic says Claude Opus 5.5 costs about 40% less than Opus 5 on typical workloads. The cuts are the latest examples of a broader decline in the cost of achieving a given level of AI performance.
Q: How fast are AI costs falling?
A: Epoch AI estimates that since 2023, the cost of achieving a given level of AI performance has fallen about 47% per quarter, equivalent to roughly 13-fold per year. The exact rate varies by task and performance level.
Q: Why is AI getting cheaper so quickly?
A: AI is getting cheaper because several efficiency gains are happening at the same time. Better algorithms, more capable hardware, improved model designs, smarter inference software and caching are all reducing the resources needed to produce useful AI work.
Q: If AI is becoming more efficient, why would we still need more compute?
A: AI can use less compute to perform a fixed task while total compute demand rises because people are using more AI, asking it to handle more complicated work and creating workflows that involve many steps, tools and model calls.
Q: What does cheaper AI mean for businesses?
A: Cheaper AI can make applications practical that previously cost too much to justify. That could allow businesses to use AI for more tasks and more complicated work, changing the practical question from “Can we afford this level of AI?” to “What can we now do with it — and can the infrastructure support that scale?”
Q: Could infrastructure become the next AI bottleneck?
A: It could. As AI becomes cheaper and more widely used, demand may increasingly fall on the physical systems required to run it, including computing power, chips, data centers, electricity, cooling and networking. Epoch AI has already pointed to rising prices for inputs such as chips and electricity even as the cost of AI performance continues to fall.
Q: Will AI costs keep falling this quickly?
A: No one knows yet. Modern generative AI has only a few years of comparable history, benchmarks do not perfectly represent real-world work, and future efficiency gains will depend on continued advances in algorithms, hardware, software, computing power and investment.
Sources:
OpenAI: Introducing GPT-6 Sol and Luna
https://openai.com/index/introducing-gpt-6-sol-and-luna/Anthropic: Claude Opus 5.5
https://www.anthropic.com/claude-opus-5-5Epoch AI: The plunging price of thought
https://epoch.ai/publications/the-plunging-price-of-thoughtStanford HAI: AI Index 2025: State of AI in 10 Charts
https://hai.stanford.edu/news/ai-index-2025-state-of-ai-in-10-chartsStanford HAI: The 2025 AI Index Report
https://hai.stanford.edu/ai-index/2025-ai-index-reportNVIDIA Blog: How NVIDIA’s Inference Software Stack Powers the Lowest Token Cost
https://blogs.nvidia.com/blog/inference-software-lowest-token-cost/Epoch AI: Inference economics of language models
https://epoch.ai/publications/inference-economics-of-language-modelsEpoch AI: Algorithmic progress in language models
https://epoch.ai/publications/algorithmic-progress-in-language-modelsOpenAI: A scorecard for the AI age
https://openai.com/index/a-scorecard-for-the-ai-age/
Editor’s Note: This article was created by Alicia Shapiro, CMO of AiNews.com, with writing support, AEO/GEO/SEO optimization, image concept development, and editorial structuring support from ChatGPT, an AI assistant. All final editorial decisions, perspectives, and publishing choices were made by Alicia Shapiro.
