Google launches TPU 8t for training and TPU 8i for inference

For ten years, AI accelerators have been general-purpose silicon. Google's eighth-generation TPU launch ended that. The TPU 8t is for training, the TPU 8i is for inference, and the rest of the industry now has to pick a side.

Google launches TPU 8t for training and TPU 8i for inference

At Google Cloud Next '26, Google introduced two new chips and one big idea. The chips are the TPU 8t — for training — and the TPU 8i — for inference. The idea is that the era of one chip doing both jobs is over. Google's official launch post is here; CNBC frames the Nvidia angle.

The Two Chips

The TPU 8t is the training chip. A single 8t superpod scales to 9,600 chips with two petabytes of shared HBM, hitting 121 FP4 exaFLOPS per pod — roughly 2.7x the per-dollar training performance of Ironwood, last year's TPU. It is co-designed with Broadcom and built on TSMC's 2nm process node.

The TPU 8i is the inference chip. It connects 1,152 chips per pod, triples on-chip SRAM versus Ironwood, and delivers up to 80% better performance-per-dollar at low-latency targets — particularly for the kinds of large mixture-of-experts models that increasingly underpin modern frontier systems. It is co-designed with MediaTek. Both chips share Google's Axion ARM-based custom CPUs and ship on JAX, PyTorch, MaxText, SGLang, and vLLM. Generally available later this year.

Google calls it a launch for "the agentic era," and the framing is intentional. Their argument: training and inference are no longer the same workload, and pretending they are leaves performance on the table. Training needs maximum throughput across enormous synchronous batches. Inference for AI agents needs sustained low-latency parallelism across hundreds of concurrent reasoning loops. Different problem, different chip.

Why This Is the Real Story

For most of the last decade, "AI chip" has meant "Nvidia GPU." H100, H200, B100, B200 — all general-purpose tensor accelerators that do training and inference on the same silicon. That generality is what built Nvidia's 80%-plus share of the $400 billion AI accelerator market. It is also what created the vendor concentration risk that every hyperscaler has spent the last two years trying to escape.

Google splitting its eighth generation into two SKUs is the most credible counterargument anyone has put on the table. It says: if you specialize, you can beat Nvidia on price-performance for whichever workload matters most to you. If you are a frontier lab spending nine figures a year on training, the 8t. If you are an enterprise running thousands of agent invocations per minute, the 8i. The Nvidia roadmap does not currently have a clean answer to either of those bets.

The numbers around the launch underscore why this matters. Anthropic's annual recurring revenue passed $30 billion in March, having grown 30x in fifteen months. Google itself just committed up to $40 billion to Anthropic, much of it in compute credit. TechCrunch broke that down last week. Anthropic has separately committed up to $100 billion in spend with Amazon for around 5 gigawatts of compute. The infrastructure layer of the AI economy is now operating at a scale where 2-3% efficiency gains are real money — not bragging rights.

The Bigger Pattern

Splitting silicon by workload is not a Google original. Apple has built its M-series the same way for years (separate cores for performance vs. efficiency vs. neural). Amazon's Inferentia and Trainium pair has long been Trainium-and-Inferentia. What is new is Google taking the bifurcation public with the most aggressive performance numbers anyone has put on a TPU launch — and packaging it as the canonical answer to what an "agentic era" of AI workloads requires.

Expect the response to come from two directions. Nvidia will likely lean harder on its own software moat — CUDA, NeMo, the developer ecosystem — to argue that workload-specific silicon does not matter when your tooling spans both. Hyperscalers without their own chip programs (most acutely Meta and Microsoft) will face renewed pressure to either accelerate internal silicon or lock in deeper TPU contracts. Meta has already been a Google TPU tenant since this winter; the 8t/8i bifurcation gives Google a sharper pitch for that business.

The TPU 8t and 8i ship later this year. The pricing has not been published. The full specs will get a deeper read at Cloud Next's later sessions and in the inevitable independent benchmarks. But the strategic move is already legible: the era of one-size-fits-all AI chips ended Tuesday, and the entire industry now has to decide whether to follow Google's split or argue against it. That same specialize-or-consolidate logic is now reshaping the merchant chip market, where onsemi agreed to acquire Synaptics for $7 billion to build out purpose-built edge-AI silicon.

Related on Uristocrat: Qualcomm's AI-wearable chip deals with OpenAI and Meta.

Comments

Get tomorrow's roundup. Free.

One email each morning. Sneakers, sports, culture, tech.

Link copied