Quantization on Blackwell: what FP4 actually costs you
On a GB10, quantization isn't just how you make a model fit. It's the main dial you have on how fast the model answers — and the first place quality quietly goes missing. Here's the arithmetic, and the bill.
Most quantization advice is written for people trying to squeeze a model onto a card that's too small for it. The GB10 has the opposite problem. With 128 GB of unified LPDDR5X, nearly everything fits. So the interesting question stops being "will it load?" and becomes "what am I actually buying, and what am I paying for it?"
The answer has two halves. Quantization buys you room — that part is obvious. It also buys you speed, for a reason that has nothing to do with arithmetic precision and everything to do with memory bandwidth. And it costs you something real, just not evenly across tasks.
The memory math
Start with weights. A 70-billion-parameter model stores one number per parameter, so the precision you pick sets the floor directly:
- FP16 — about 140 GB. Two bytes a parameter. Doesn't fit in 128 GB, full stop.
- FP8 — about 70 GB. Fits, and leaves a bit over half the machine free.
- 4-bit — about 40 GB once you account for the scales, embeddings and layers that usually stay at higher precision. Uses under a third of the pool.
Then there's the tenant nobody budgets for: the KV cache. Every token you keep in context parks its attention state in memory for the life of the session. For a 70B model with grouped-query attention — 80 layers, 8 key/value heads, 128 values per head, keys and values both, two bytes each — that's about 320 KB per token. Run the context out to 128K tokens and the cache alone wants roughly 40 GB, the same order as the quantized weights it sits next to.
Put those together and the shape of the machine appears. A 4-bit 70B with a long context is an 80 GB workload on a 128 GB box — comfortable. The same model at FP8 is 70 GB of weights plus that cache, and you're suddenly rationing context against precision. That trade is the whole game.
Why quantizing makes it faster, not just smaller
Here's the part that surprises people coming from data-center GPUs. To produce a single token, the model reads all of its active weights out of memory. Not some. All of them, every token. So single-stream decoding isn't limited by how fast the chip can multiply — it's limited by how fast weights can be dragged across the memory bus.
The DGX Spark's unified memory runs at 273 GB/s. Divide that by the size of your model and you get a hard ceiling on tokens per second:
- 70B at 4-bit (~40 GB): 273 ÷ 40 ≈ 7 tokens/second, best case.
- 70B at FP8 (~70 GB): 273 ÷ 70 ≈ 4 tokens/second, best case.
- A 14B at 8-bit (~14 GB): 273 ÷ 14 ≈ 19 tokens/second, best case.
Halving the bytes per parameter roughly doubles the ceiling. That's the real argument for FP4 on this hardware: not that the 70B wouldn't fit at FP8, but that at FP8 it answers at roughly half the speed.
Every number above is arithmetic — memory bandwidth divided by bytes read per token — not something we measured. Real throughput lands below it, because attention reads the KV cache too, sampling and detokenization cost something, and nothing achieves 100% of theoretical bandwidth. Treat the ceiling as the thing you cannot exceed and the shape of the trade-off as the thing worth trusting.
NVFP4 isn't the INT4 you remember
Four-bit has a bad reputation, and it earned it honestly. Early INT4 schemes mapped a whole group of weights onto sixteen evenly spaced integers with one shared scale. Neural network weights are not evenly distributed — a handful of outliers carry disproportionate weight — so a single scale stretched across a large group either clipped the outliers or wasted most of its range on them.
Blackwell's native format, NVFP4, fixes this from two directions. Each value is a 4-bit float (one sign bit, two exponent bits, one mantissa bit) rather than an integer, so its precision is finest near zero where most weights actually live. And the scaling is hierarchical: every block of 16 values gets its own FP8 scale factor, with a single FP32 scale applied across the tensor on top. Small blocks mean one outlier contaminates fifteen neighbours instead of hundreds, and an FP8 block scale can land on the value it needs rather than rounding to the nearest power of two the way MXFP4's coarser 32-value blocks do.
The practical upshot is that the second-generation Transformer Engine does this in hardware, on the chip, as part of the matmul — so the format isn't a compression trick you pay to unpack. It's how the numbers are stored and multiplied.
What it actually costs you
Quality loss from 4-bit isn't a flat tax. It concentrates, and knowing where lets you spend it deliberately:
- Nearly free: conversational chat, summarization, classification, extraction, and anything grounded in retrieved text. When the answer is mostly in the prompt, a slightly fuzzier model finds it just fine.
- Where it shows first: long multi-step reasoning, where small per-step errors compound over a chain; code generation, where one wrong token is a syntax error rather than a slightly worse word; and arithmetic or rare proper nouns, which live in exactly the low-probability tail that quantization blurs.
- A separate dial entirely: KV-cache quantization. Storing the cache at 8-bit halves that 40 GB and buys back context, but it degrades differently — it erodes what the model remembers from early in a long context, which is precisely why you wanted the long context. Change one dial at a time.
The useful instinct: if your workload is retrieval-grounded and conversational, take the 4-bit speed. If it's agentic, long-horizon, or writing code you intend to run, test both precisions on your own prompts before deciding the faster one is good enough.
Picking a dial on 128 GB
Rules of thumb for a single GB10, assuming interactive single-user use:
- 70B class — 4-bit is the default. FP8 is the better model on paper and roughly half the speed in practice, while leaving much less room for context. Reach for it only when you've confirmed 4-bit is failing your task.
- 27B–32B class — 8-bit is comfortable. Around 30 GB of weights leaves plenty of pool for a long context, and the bandwidth ceiling is high enough that the extra precision is close to free.
- 8B–14B class — 8-bit or FP16. Down here bandwidth stops being the binding constraint. Spend the memory on precision and context; you won't miss the speed.
None of which replaces trying it. The dial that matters is the one that holds up on your prompts, with your context, in your tooling — and that's a half-hour experiment, not a research project.
Test your own quantization.
Rent a full Grace Blackwell by the hour, load the precision you're arguing about, and run your own prompts against it.
Spin up a GB10 session