Why Today May Have Changed AI Forever
31 Jul
Today may prove to be one of the most important days in the economics of artificial intelligence.
For the past several weeks, the evidence has been accumulating.
Kimi K3.
More efficient routing.
Distillation.
Continuous learning.
Better memory.
Improved post-training.
Open-weight frontier models.
Each release chipped away at a long-held assumption.
Today, DeepSeek V4 Flash may have turned that trend into something much harder to ignore.
This wasn’t simply another benchmark improvement.
It was another demonstration that frontier-level capability can increasingly be achieved with dramatically less computation, dramatically lower cost and hardware that, until recently, many believed would never be sufficient.
This isn’t simply another model launch.
It challenges one of the largest investment assumptions in modern technology.
The Assumption
For much of the AI race, progress appeared straightforward.
Build larger models.
Train on more data.
Buy more GPUs.
Build more data centres.
Raise more capital.
Every improvement seemed to require a corresponding increase in infrastructure.
That assumption justified hundreds of billions of dollars of investment.
Hyperscalers committed enormous capital expenditure.
Cloud providers expanded at unprecedented speed.
Power generation accelerated.
Supply chains stretched.
Balance sheets absorbed long-term infrastructure commitments under a single assumption.
More intelligence would always require proportionally more compute.
That assumption is beginning to break.
The Last Few Weeks Changed The Conversation
Look at what has happened.
Kimi K3.
More efficient routing.
Distillation.
Continuous learning.
Better memory.
Improved post-training.
Open-weight frontier models.
Now DeepSeek V4 Flash.
Individually these look like product releases.
Collectively they tell a completely different story.
Every release pushes intelligence higher…
while pushing computation lower.
The trend is no longer subtle.
It is becoming impossible to ignore.
Efficiency Is Becoming The New Scaling Law
For years the industry measured progress through scale.
More parameters.
More FLOPs.
More GPUs.
Larger clusters.
Increasingly, progress comes from something different.
Better architectures.
Sparse activation.
Improved routing.
Higher-quality post-training.
Distillation.
Speculative decoding.
Better memory.
Continuous learning.
The result is remarkable.
DeepSeek V4 Flash contains roughly 304 billion parameters, but only a small fraction are active during inference.
That is the point.
Modern intelligence is no longer brute force.
It is increasingly about activating exactly the computation you need…
and avoiding everything you don’t.
Models with far fewer active parameters are approaching—and on some workloads matching—the capability of much larger systems.
Not because intelligence has stopped improving.
Because waste is disappearing.
Intelligence Is Decoupling From Infrastructure
Today’s announcement captured the shift perfectly.
DeepSeek released an open-weight model that is competitive with leading proprietary systems while being dramatically cheaper to run.
Even more remarkably…
It can increasingly run locally.
Not only inside hyperscale data centres.
On powerful workstations.
On AI PCs.
On Apple Silicon Macs.
And increasingly…
On laptops.
That doesn’t eliminate cloud infrastructure.
Frontier training will continue requiring enormous compute.
Large enterprise deployments will continue using hyperscale infrastructure.
But it changes something fundamental.
The assumption that every increase in AI capability requires a proportional increase in centralised inference becomes much harder to defend.
Inference is beginning to move towards the edge.
CapEx Was Built Around A Different Future
Infrastructure investment has always involved forecasting future demand.
The current AI build-out assumes that inference demand scales alongside intelligence.
That assumption underpins hundreds of billions of dollars of committed capital expenditure.
Multi-year GPU purchases.
Purpose-built AI data centres.
Long-term power agreements.
Expanded manufacturing capacity.
Balance sheets built around the expectation that ever-larger infrastructure would always be required.
But if every new generation becomes dramatically more efficient…
If comparable capability requires fewer GPUs…
If more inference runs locally…
If open-weight models continue compressing costs…
Then some of today’s capital assumptions deserve revisiting.
The question is no longer simply:
“How much AI will the world need?”
It is increasingly:
“How much infrastructure will that AI actually require?”
Those are very different questions.
And for companies that have already committed enormous capital programmes…
The distinction matters.
History Repeats
Technology repeatedly follows the same pattern.
Computers became smaller.
Storage became cheaper.
Bandwidth became abundant.
Solar panels collapsed in cost.
Semiconductors became exponentially more efficient.
Artificial intelligence now appears to be entering exactly the same phase.
Capability continues rising.
Cost continues falling.
The Scarcity Moves
As intelligence becomes abundant…
Value migrates elsewhere.
Not to producing more intelligence.
To deploying it better.
To organising it better.
To remembering.
To trust.
To resolving uncertainty.
To becoming the system an intelligent agent can confidently defend.
Intelligence increasingly becomes infrastructure.
Trusted resolution becomes the scarce resource.
The Resolution Economy
The Model Economy rewarded whoever built the largest model.
The Resolution Economy rewards whoever resolves uncertainty with the greatest confidence while performing the least unnecessary computation.
Every major release over the past few weeks points in exactly the same direction.
Capability converges.
Efficiency compounds.
Costs compress.
Inference moves closer to the edge.
Capital assumptions begin to change.
The economics begin to change.
This isn’t simply another model launch.
It may prove to be the beginning of one of the largest reallocations of value the computing industry has experienced in decades.
Because the biggest breakthrough isn’t that models are becoming more intelligent.
It’s that they’re becoming dramatically cheaper to be intelligent.
For years the industry believed intelligence scaled by spending more capital.
Increasingly, it looks like intelligence scales by wasting less computation.
History suggests those are two very different futures.
And 31 July 2026 may be remembered as the day that became impossible to ignore.