ZKSF logo, a neon quantum brainZKSF
← All articles

Google Cloud TPU vs GPU for Quantum Computing, Measured

Last updated · 9 min read · ZKSF team

The short version

  • It tied the GPU everywhere they met. On every problem both chips ran, the Google Cloud TPU matched the NVIDIA GPU within ordinary sampling noise
  • It won the one exact rematch. Given the identical neural network request, the TPU finished three times closer to the true ground state
  • Best on the table four times. On satellite tasking, vehicle routing, job shop scheduling and the neural wavefunction, the TPU run returned the best answer on its table, ahead of three superconducting quantum processors on three of them
  • From your phone too. Both TPU engines sit in our Android app beside the CPUs, GPUs and quantum processors

Everything here is runnable on your own circuit. Try it in the console

GPUs own the conversation about AI compute. NVIDIA sells the cards, the cards train the models, and most benchmarks stop right there, which leaves anyone curious about the alternatives with very little to go on.

We wanted to see what happens when a Google Cloud TPU is handed the same homework. Twelve of the sixteen quantum computing applications we publish now carry a TPU row beside CPU, GPU and real quantum hardware rows, every one of them a billed run with its result attached.

We expected the NVIDIA GPU to take the lead. It didn't. On a small sample that deserves a far deeper follow up, the Tensor Processing Unit tied the GPU wherever they met and came out ahead on the one problem where the two received exactly the same request. We were pleasantly surprised, and we will admit to a little smugness.

What is a Tensor Processing Unit (TPU)?

A Tensor Processing Unit is an AI chip Google designed in house for the large matrix multiplications at the heart of neural networks. Inside sits a systolic array, a grid of multiply and accumulate units that hands numbers from cell to cell so they rarely make the slow round trip to memory. That one specialisation is where its speed and its efficiency come from.

You can't buy a TPU. Google keeps them in its own data centres and rents them out as Google Cloud TPU, the same AI compute Google uses to train its own models. For how the four kinds of compute differ in general, our guide to CPU vs GPU vs TPU vs QPU walks through each one. This post is the TPU's turn in the spotlight.

TPU v4, TPU v5e and Cloud TPU v6e (Trillium)

Google has released several TPU generations, and three names come up most when people compare them.

GenerationKnown forHow we use it
TPU v4Large pods for training big models, PaLM among themNot used here
TPU v5eEfficiency, built to keep the cost per chip low for training and inferenceFirst choice for both TPU engines
Cloud TPU v6e, TrilliumThe generation after v5e, with more compute and memory per chipSecond choice for both TPU engines

Every job here runs on a single chip, which covers everything we ask of it. One v5e chip already holds a 29 qubit statevector and trains our neural networks comfortably, so v5e answers first and v6e is the second choice.

TPU vs GPU vs CPU vs QPU on twelve quantum workloads

Here is every application where the TPU ran, set against the NVIDIA GPU and the quantum processors that tried the same problem. Smaller gaps and higher accuracy are better.

ApplicationGoogle TPUNVIDIA GPUQuantum processorOutcome
Satellite tasking, 14 qubitsgap 1.30gap 7.36best, Rigetti, gap 5.49TPU best on the table
Vehicle routing, 16 qubitsthe exact optimumgap 2.28best, IQM Garnet, gap 4.05TPU best, the only exact hit
Job shop scheduling, 16 qubitsgap 2.81gap 3.28best, IQM Emerald, gap 4.30TPU best on the table
Neural quantum states, 6 spins0.130 from exact0.402 from exactnone ranTPU ahead of the GPU
Portfolio optimisation, 12 assetsthe provable optimumbeats 96.4%best, IQM Garnet, beats 99.2%TPU ahead of the GPU, tied with a CPU tensor network
Reservoir computing, 4 qubits0.783 accuracy0.767 accuracynone ranTPU ahead by one test point
Born machine, 2 qubitsTVD 0.049TVD 0.041Rigetti TVD 0.102, Garnet 0.104Tie within shot noise
Reinforcement learning, 2 qubits0.147 and 0.6320.158 and 0.624Rigetti 0.199 and 0.650Tie
Quantum kernel classifier, 2 qubits0.6750.675none ranTie
Shor's algorithm, 7 bit keykey recoveredkey recoverednone ranTie
Traffic routing, 29 to 30 qubitsno valid routeno valid routeno valid routeNobody solved it
H2 molecule, neural engines0.0203 Ha from exactno matching runa different measurementTie with the CPU

Count it up and the TPU finished ahead of the GPU six times and tied it five times. The GPU did not finish clearly ahead once.

A word on fairness, because these numbers deserve it. On the four optimisation problems each chip sampled its own build of the instance's circuit at 500 shots, and the best of 500 samples always carries some luck of the draw. One run per chip is a small sample too. So read the pattern across all twelve rows, and watch this space for the repeats.

Where the TPU beat the GPU head to head

The cleanest comparison in the whole set is a neural network quantum state, a neural network trained to represent the ground state of a ring of six interacting spins. Both chips received the same Hamiltonian and the same settings, twenty optimisation steps, 1,024 samples and one restart, so the only real difference was the silicon underneath.

The GPU stopped at an energy of -6.647. The TPU reached -6.919. The exact answer is -7.049, which puts the TPU three times closer to the truth.

Run on our engines

A spin Hamiltonian solved on the TPU tier, and the H2 molecule solved on both neural engines for comparison. On 25 September the identical spin request ran on neural.gpu, an NVIDIA GPU, reaching energy -6.647038 with a ceiling of -6.630361. Submitted to each kind of compute we offer, on 16, 18 and 25 September 2026. Every figure below is a real job on the service, priced as any customer would be priced.

DeviceEngineKindQubitsResultCost
Google Cloud TPUneural.tpuTPU6energy -6.918861, ceiling -6.886729, against an exact ground state of -7.048804$0.0395
NVIDIAneural.gpuGPU6energy -6.647038, ceiling -6.630361, against an exact ground state of -7.048804 certificate$0.0253
CPUneural.cpuCPU2H2 at -1.116981 Ha, 0.0203 Ha above exact certificate$0.0001
Google Cloud TPUneural.tpuTPU2the same H2 molecule, -1.116981 Ha certificate$0.0740

The same problem is yours to run: every instance here is seeded, so it rebuilds exactly. Open the console and a cost estimate is free before anything executes.

Why would a TPU do better here? A neural quantum state is a neural network, and its inner loop is dense matrix multiplication over batches of sampled spin configurations, which is precisely the workload Google built the chip for. Our neural.tpu engine runs it in JAX, the framework Google itself uses on TPUs. One run on each chip is a hint rather than a verdict, so the repeats are already on our list.

TPU and GPU tie, and that is the surprise

A tie sounds dull until you remember the usual assumption, that a chip built for low precision neural network maths has no business doing quantum simulation. Our exact.tpu engine holds the quantum state in single precision complex numbers, where GPU engines usually work in double. For shot based results that difference sits far below ordinary sampling noise, and the table bears it out.

On the Born machine, the reinforcement learning policy, the kernel classifier and Shor's algorithm, the TPU landed on the same answer as the GPU. On the Born machine it put none of its 1,000 shots on the two outcomes the target forbids, exactly what an ideal simulator promises.

That makes a TPU more than one more fast chip. It is an independent second opinion. Different silicon, a different compiler and a different arithmetic path arrived at the same answers, so if you want to check a GPU result before you publish it or bet money on it, a Google Cloud TPU is now a credible referee.

Where the TPU beat everything, quantum processors included

Satellite observation tasking, vehicle routing and job shop scheduling are three optimisation problems we ran on everything we have. On all three, the TPU run returned the best answer on the table, ahead of three superconducting quantum processors from Rigetti and IQM. On vehicle routing it was the only run of eight to land exactly on the optimum, -18.7897.

Run on our engines

A random 16-variable QUBO, seed 20260902, whose exact optimum is -18.7897: the size a machine can hold, with no routing structure in it, generated by small_qubo.py in the public benchmarks folder. It is deliberately small, because the 100-stop round this benchmark is about needs 10,000 qubits and fits nothing. Submitted to each kind of compute we offer, on 16 and 25 September 2026 at 500 shots. Every figure below is a real job on the service, priced as any customer would be priced.

DeviceEngineKindQubitsResultCost
CPUmps.quimb.cpuCPU16-16.3521, gap 2.44 certificate$0.0001
CPUexact.cpuCPU16-14.0904, gap 4.70 certificate$0.0001
NVIDIAexact.gpuGPU16-16.5132, gap 2.28 certificate$0.0001
Rigettiqpu.rigettiQPU16-216.06, gap 76.30 on the routing instance * certificate$0.5125
IQMqpu.iqm.garnetQPU16-14.7420, gap 4.05 certificate$1.025
IQMqpu.iqm.emeraldQPU16-10.5934, gap 8.20 certificate$1.100
Google Cloud TPUexact.tpuTPU16-18.7897, the optimum certificatebest outcome$0.0776
Google Cloud TPUneural.tpuTPU—a QUBO is diagonal, which is not the shape a neural ansatz is for—

* The Rigetti row is not scored against the optimum named above. The run used the 4-vehicle routing instance that sector_instances.py generates instead. Its exhaustive minimum is -292.3588, and the 76.30 gap is measured against that rather than against -18.7897. The IQM rows are our own internal testing. The exact.tpu row, added 25 September, is the only run on this instance to reach the exact optimum of -18.7897. The steps are in the docs.

A note on the hardware certificates: they state Hellinger fidelity against the exact distribution. For an optimisation circuit that distribution is spread across many outcomes rather than concentrated on one, so the figure is low by construction and is not a measure of whether the device found a good answer. The result column above is.

The same problem is yours to run: every instance here is seeded, so it rebuilds exactly. Open the console and a cost estimate is free before anything executes.

All three are QAOA circuits, the workhorse algorithm for optimisation on quantum hardware, and the full story of how each instance was built and scored lives on its application page.

Train machine learning models on a TPU, quantum ones included

Most people meet TPUs when they train machine learning models in JAX, TensorFlow or PyTorch through XLA. Quantum machine learning fits the same mould. The Born machine, the reinforcement learning policy, the kernel classifier and the reservoir computer on our AI applications page all ran their circuits on the TPU tier, and the neural quantum states engine trains its network on the chip itself.

If you already know how to train a model on a TPU, you already understand most of what neural.tpu does. The difference is the loss. Ours is the energy of a quantum system, and the answer comes back with a certified upper bound on the true ground state.

Google Cloud AI compute without the setup

Renting a Cloud TPU directly means a Google Cloud project, quotas, a machine image and usually an afternoon of configuration. Here it is a choice in a dropdown. Pick exact.tpu for gate circuits up to 29 qubits or neural.tpu for Hamiltonians, then submit from the web console, the Python SDK or the Android app, and pay for the seconds it runs.

import qsim_sdk

client = qsim_sdk.Client(token="YOUR_TOKEN")
job = client.run(qc, engine="exact.tpu", shots=1000)
print(job["result"]["counts"])
The ZKSF console with its full engine roster open, CPU, GPU and every quantum processor on the platform in one place, above a job history showing what each tier actually cost
The ZKSF console with its full engine roster open, CPU, GPU and every quantum processor on the platform in one place, above a job history showing what each tier actually cost. Try it yourself in the console

Use a Google TPU from your phone

Here is the part we enjoy most. Both TPU engines are in our quantum computing mobile app for Android, alongside the CPUs, the NVIDIA GPUs and eight real quantum processors, 24 engines in all. Open a template, choose exact.tpu, and a Tensor Processing Unit in a Google data centre runs your circuit while you finish your coffee.

Try the TPU rows yourself

The docs carry seeds for the two newest TPU rows, the trained Born machine weights and the policy weights, so you can send the very same circuits to exact.tpu and check our numbers against yours. Every other row links to its application page, where the circuits, shot counts and costs are published in full.

TPU questions, answered

Can a TPU run quantum circuits?

Yes. Our exact.tpu engine runs gate circuits of up to 29 qubits as an exact statevector simulation on a single TPU chip, and neural.tpu solves Hamiltonians with a neural network.

Is a TPU better than a GPU?

On these quantum workloads the TPU matched the GPU everywhere both ran and finished ahead on the neural network problem. For software written for CUDA, or for graphics work, a GPU stays the natural choice.

Which is better, TPU v5e or TPU v6e?

For single chip jobs like ours, v5e gives the same answers at a lower price per chip hour. v6e brings more compute and memory per chip, which pays off on larger models spread across many chips.

Can I use a Google TPU from my phone?

Yes. The ZKSF Android app offers both TPU engines, so a TPU job is a template and a tap away.

The verdict, for now

We set out expecting to write the usual story, the one where the GPU wins and everyone nods along. Instead the TPU matched it at every turn and edged ahead where the workload looked most like the job Google built it for. Twelve applications and one run per chip is a starting point, and we will keep adding repeats and larger instances to the same tables as they finish. If you have only ever reached for a GPU, give the TPU a run. It might surprise you too.

Run your own 100-qubit circuit, with an error bar.

Share this articleLink copied