Google Developing "Frozen v2" AI Chip to Supercharge Gemini Efficiency
Alphabet designs custom server silicon targeting up to tenfold improvements in token generation and power consumption.
On July 20, 2026, details emerged regarding a high-stakes internal semiconductor project inside Google. Alphabet, Google’s parent company, is designing a new custom server chip engineered specifically to help its in-house Gemini models operate with unprecedented efficiency. This initiative directly addresses global shortages in AI computing capacity and aims to wean Alphabet off its heavy reliance on chipmaker Nvidia, whose industry dominance has long dictated the financial and infrastructural margins of foundation model laboratories.
Key Details
The custom silicon project, internally dubbed “Frozen v2,” represents a direct pivot toward vertically integrated hardware-software co-design. While Google has not publicly confirmed the development, reports indicate that the chip could represent a massive leap over the company's current processing units.
- Dramatic Efficiency Gains: According to internal source projections, the "Frozen v2" chip is targeted to be between six and ten times more efficient than Google’s existing custom TPU processors. This is measured by the number of tokens generated per unit of electric power.
- Projected Release Timeline: Alphabet aims to bring the new server chip into production and deployment sometime around 2028, representing a medium-term hedge against escalating infrastructure bills.
- Hardware-Software Co-Design: Rather than building general-purpose accelerators, Google’s hardware engineers are building Frozen v2 from the ground up to align with the specific mechanisms underlying the Gemini model family.
- A Full-Stack Stance: In response to inquiries, Google emphasized its commitment to co-designing hardware and software. The company noted that this rigorous exploration ensures systems are highly integrated and optimized for real-world workloads.
- Capital Expenditure Justification: Alphabet’s planned AI infrastructure capital expenditures for the year are expected to range between $180 billion and $190 billion. An in-house efficiency breakthrough is vital to demonstrating long-term return on investment to shareholders.
What This Means
As concerns over the massive capital expenditure required to sustain the AI boom intensify, tech giants are entering a critical second phase. The initial race for raw model scale is shifting into a high-stakes competition over operational margins. It is no longer enough to build the smartest model; companies must build the most cost-effective intelligence pipelines.
By reducing the power cost per token, Google could lower the price of Gemini API calls for developers and enterprise clients. Furthermore, achieving a tenfold efficiency improvement would allow Google to run complex agentic loops and multi-step reasoning processes that are currently cost-prohibitive under existing GPU clusters.
Technical Breakdown
To achieve a six-to-tenfold efficiency leap, Alphabet’s hardware team is prioritizing highly specialized neural processing:
- Mathematical Specialization: Frozen v2 is designed to execute low-precision matrix multiplication and specialized attention algorithms natively in silicon, eliminating instruction overhead.
- On-Chip Memory Optimization: The chip optimizes data locality, keeping model weights and context memory physically closer to the processing cores to minimize energy-intensive data transfer bottlenecks.
- Dynamic Voltage Scaling: The processor dynamically scales its power draw based on prompt complexity, ensuring that simpler conversational queries do not waste precious energy.
- Native Mixture-of-Experts (MoE) Routing: The hardware architecture is optimized to support MoE models like Gemini, routing activation pathways dynamically through specialized sub-networks directly at the physical chip layer.
Industry Impact
Google's internal push for custom silicon signals a broader industry trend toward hardware independence. Major AI laboratories are realizing that relying solely on off-the-shelf accelerators from Nvidia is an unsustainable long-term strategy, both due to premium pricing and supply chain bottlenecks.
By designing Frozen v2, Google is joining a crowded field of proprietary silicon builders:
- OpenAI: The startup announced its first custom inference processor, "Jalapeño," built in partnership with Broadcom, in June 2026.
- Anthropic: Earlier this month, Anthropic entered talks to partner with Samsung for custom chip design to optimize its Claude model family.
- Nvidia: The dominant GPU manufacturer faces a delicate balancing act, as its largest cloud customers increasingly transform into direct hardware competitors.
Looking Ahead
While the projected 2028 release timeline means Google must continue relying on its existing TPUs and Nvidia’s current-generation architectures for the immediate future, Frozen v2 sets a clear trajectory for the company's research. As AI systems migrate from simple chat interfaces to persistent, long-running autonomous agents, the demand for continuous compute will swell exponentially. Achieving a tenfold power reduction is not just a financial victory; it is a prerequisite for the next stage of agentic intelligence. If Alphabet successfully brings Frozen v2 to production, it could secure a permanent infrastructural moat, ensuring that Gemini remains competitive even as the global energy grid faces unprecedented strain.
Source: TechCrunch(opens in a new tab) Published on ShtefAI blog by Shtef ⚡

