The Information reports that Google is building a new server chip, internally dubbed “Frozen v2,” meant to run its Gemini models far more efficiently. The idea, since picked up by everyone from Reuters to CNBC, is easy to say and hard to do- instead of running Gemini on general-purpose hardware, you etch parts of Gemini’s architecture directly into the silicon.
The weights can still change. Engineers can load new numbers into the model. But the shape- the structure of the network itself- stays fixed; frozen. Hence the name.
Now the number everyone is quoting: Google’s engineers reportedly project six to ten times more tokens per unit of power than the company’s newest TPUs. For context, a normal generational leap in chips buys you two to three times better performance per watt.
Deployment is targeted around 2028. Google didn’t confirm it. It didn’t deny it either.
Why would anyone freeze a model into a chip?
Because flexibility is expensive.
A general-purpose chip has to be ready for anything- any model, any architecture, any workload you throw at it tomorrow. That readiness costs energy. Data gets shuttled back and forth, instructions get decoded, the hardware keeps its options open. Freezing the model removes the options. You stop paying for what you don’t use.
This is a very old move wearing new clothes. Google isn’t asking “how do we build a faster chip?” It’s asking “what business is this chip actually in?” And the answer it seems to have landed on is that it isn’t in the general-compute business at all- it’s in the run-Gemini business. Everything else is overhead. Frozen would sit as a specialized branch of Google’s chip portfolio, not a replacement for the TPUs.
There’s a real reason for the urgency, too. The project is reportedly aimed at easing internal compute shortages that have limited Google Cloud’s ability to serve some enterprise customers. Read that again. The bottleneck isn’t demand. It’s supply. They have people who want to buy and not enough silicon to sell them.
The catch
Here’s the part the stock-pop headlines skip.
The thing that makes Frozen fast is the same thing that makes it fragile. You get the efficiency because the hardware and the model are welded together- but weld two things together and you can no longer move one without the other. If Gemini’s architecture shifts in a big way, the chip built for the old shape becomes an expensive paperweight.
So Frozen is a bet on stability. It only pays off if Google believes the fundamental shape of a transformer isn’t going to be reinvented before 2028. That’s a confident thing to believe in a field that redesigns itself every six months. Freezing cuts both ways- it’s efficient precisely because it refuses to change, and it’s risky for exactly the same reason.
What it actually tells you
Forget the chip for a second.
The story underneath the story is that the AI industry has quietly stopped competing on who has the smartest model and started competing on who can run it cheapest. The frontier is moving from intelligence to economics- from “can it think?” to “can you afford to let it think at scale?”
That’s why Google is willing to hardwire its crown jewel into metal. That’s why it’s reportedly hiring to help businesses actually use this stuff. The model was never the moat. The cost per token is.
And if you’re building anything on top of these systems, that’s the shift to watch. The next advantage won’t come from access to a better model- everyone will have that. It’ll come from whoever figured out how to serve it for a tenth of the power. Google is freezing a chip to win that fight.
The rest of us should be asking the same question it is: what are we paying for that we don’t actually use?


