
New cost data from Artificial Analysis shows that DeepSeek's V4-Flash can be pushed through a full AI benchmark suite for roughly three cents, making it the cheapest widely recognized model to operate. The figure is far below every comparable frontier model and underscores how quickly the economics of AI inference are changing.
The research firm's Intelligence Index is designed to measure the practical capabilities of large language models by running a standard battery of tasks. It also tracks the cost of completing that battery on each model. DeepSeek's V4-Flash came in at about three cents. Moonshot's Kimi K3 cost 86 cents. OpenAI's GPT-5.6 Sol cost $1.86. Anthropic's Claude Fable 5 cost $3.15. None of the well-known competitors came close to DeepSeek's number.
What DeepSeek charges
On published pricing, DeepSeek charges $0.14 per million input tokens and $0.28 per million output tokens for V4-Flash. For developers handling millions of requests, those rates matter more than broad marketing talk about model quality. The model is the lighter sibling in DeepSeek's latest lineup, which the Chinese lab introduced with V4-Pro and V4-Flash. The Pro variant is aimed at harder reasoning tasks, while Flash is designed for speed, volume, and lower operating cost.
Capability trade-offs
The low price does come with a performance ceiling. V4-Flash scored 50 out of 100 on the Intelligence Index, level with Google's Gemini 3.6 Flash and just behind Meta's Muse Spark 1.1 and Z.ai's GLM-5.2, both on 51. The frontier still sits clearly ahead. Kimi K3 scored 57, while Claude Opus 5, Claude Fable 5, and GPT-5.6 landed roughly nine points higher again. Those differences matter for complex reasoning, long-horizon planning, and high-stakes professional work, but they are not the whole story for many everyday tasks.
How DeepSeek reached this point
DeepSeek has become one of the most closely watched AI labs by competing on efficiency rather than trying to outspend everyone in centralized training runs. The company showed with earlier models that strong performance could be achieved with leaner infrastructure and careful architectural choices. Its latest release pairs a powerful reasoning model with a cheaper, faster option, a pattern that is becoming common across the industry. The launch also gave DeepSeek a renewed presence in the market after a period of speculation about what it would do next.
Part of the reason V4-Flash can be priced so aggressively is that it is designed for high-volume use cases. Chatbots, coding assistants, and back-office automation systems generate millions of requests, and every fraction of a cent per request compounds into a real bill. Models like Flash are built to absorb that load without forcing developers onto more expensive frontier systems. That makes them especially attractive to startups and enterprises that want AI features without betting the budget on a premium model.
A price war DeepSeek started
The release lands in the middle of an AI price war that DeepSeek has done more than anyone to start. The company made a 75% discount permanent earlier this year, and rivals have been cutting in response. OpenAI trimmed GPT-5.6 pricing sharply, and the general drift of the market has been down, and fast, on a curve that looks less like software margins and more like a commodity. Every major lab is now aware that staying expensive is risky when a comparable service exists at a fraction of the cost.
DeepSeek has the balance sheet to keep pushing. The company recently closed its first outside funding, a round of more than $7bn, which buys room to subsidize aggressive pricing while it takes share. The funding signals that investors see the strategy as sustainable, at least for now. It also gives DeepSeek resources for the next generation of models, which could widen the gap between what Chinese labs charge and what Western rivals need to charge to cover their own research costs.
Engineering advantages
DeepSeek's edge is as much engineering as pricing. The company has leaned on efficient training and inference methods to hold costs down. This is what lets it charge so little without, it says, simply setting money on fire. Efficient inference means fewer compute cycles per answer, which lowers the marginal cost of every API call. The result is a business model that does not depend on maintaining high prices across the board, and that has forced competitors to rethink the assumptions behind their own pricing.
Those engineering choices also have an indirect effect on the market. When a low-cost model can handle a large share of real-world prompts, there is less reason to send every request to the most capable system. Applications can route simple requests to cheap models and save frontier models for hard problems. That kind of routing is becoming a standard practice, and it puts additional pressure on labs whose revenue depends on charging premium prices for every interaction.
Pressure on Western AI valuations
Analysts have argued that relentless discounting from Chinese labs puts the eventual OpenAI and Anthropic IPOs under pressure. Premium pricing is hard to defend when a rival is tens of times cheaper per task. If investors measure AI value by revenue potential, then falling prices directly threaten the projections at the center of private market valuations. The prospect of an IPO makes those numbers even more sensitive, because public investors will expect clear evidence that the business can stay profitable in a competitive market.
Not everyone is convinced the quality gap still matters. Zack Kass, OpenAI's former head of go-to-market, has framed the moment as one of 'diminishing model returns'. His argument is that once models are close enough, the next one barely moves the needle and price does the deciding. If that view is right, then DeepSeek's strategy of selling solid performance at very low cost could become even more influential. It also suggests that chasing a few extra benchmark points may not guarantee commercial success if the cheaper option is good enough for most users.
Chinese labs have been setting that pace. Moonshot's Kimi K3 spooked markets on release, and the broader worry is that a wave of cheap, open-weight models erodes the economics that Western AI valuations quietly assume. Open-weight creates another pressure: developers can self-host, bypassing API fees entirely, or use a low-cost hosted model as a baseline for their own fine-tuning. The combination of open availability and aggressive pricing is challenging the notion that AI infrastructure will always operate with software-like margins.
What benchmarks can and cannot say
Benchmarks are an imperfect proxy, and cost per test turns on how efficiently a model spends tokens as much as on its sticker price. A model can be cheap per token but wasteful in practice, while another can charge more per token and finish tasks with fewer calls. Artificial Analysis tries to capture the full picture by measuring the cost to complete a fixed test battery, but no single benchmark can capture every real-world need. Different workloads place different demands on context length, latency, reliability, and safety.
Even so, the direction is not in doubt, and Artificial Analysis has put hard figures on what developers have felt for months. The cost of high-quality AI has fallen dramatically, and the gap between price and capability is now visible in a way that is hard to ignore. For buyers, the sum is getting simpler. If a model that costs three cents to run can do most of the job, the burden shifts onto the expensive models to prove what those extra nine points on a benchmark are really worth.
