← The Log
Deep Dive·Sep 15, 2026·7 min read

Quantization on Blackwell: what FP4 actually costs you

On a GB10, quantization isn't just how you make a model fit. It's the main dial you have on how fast the model answers — and the first place quality quietly goes missing. Here's the arithmetic, and the bill.

Most quantization advice is written for people trying to squeeze a model onto a card that's too small for it. The GB10 has the opposite problem. With 128 GB of unified LPDDR5X, nearly everything fits. So the interesting question stops being "will it load?" and becomes "what am I actually buying, and what am I paying for it?"

The answer has two halves. Quantization buys you room — that part is obvious. It also buys you speed, for a reason that has nothing to do with arithmetic precision and everything to do with memory bandwidth. And it costs you something real, just not evenly across tasks.

The memory math

Start with weights. A 70-billion-parameter model stores one number per parameter, so the precision you pick sets the floor directly:

Then there's the tenant nobody budgets for: the KV cache. Every token you keep in context parks its attention state in memory for the life of the session. For a 70B model with grouped-query attention — 80 layers, 8 key/value heads, 128 values per head, keys and values both, two bytes each — that's about 320 KB per token. Run the context out to 128K tokens and the cache alone wants roughly 40 GB, the same order as the quantized weights it sits next to.

~40 GB70B weights at 4-bit
~320 KBKV cache per token
~40 GBThat cache at 128K context

Put those together and the shape of the machine appears. A 4-bit 70B with a long context is an 80 GB workload on a 128 GB box — comfortable. The same model at FP8 is 70 GB of weights plus that cache, and you're suddenly rationing context against precision. That trade is the whole game.

Why quantizing makes it faster, not just smaller

Here's the part that surprises people coming from data-center GPUs. To produce a single token, the model reads all of its active weights out of memory. Not some. All of them, every token. So single-stream decoding isn't limited by how fast the chip can multiply — it's limited by how fast weights can be dragged across the memory bus.

The DGX Spark's unified memory runs at 273 GB/s. Divide that by the size of your model and you get a hard ceiling on tokens per second:

Halving the bytes per parameter roughly doubles the ceiling. That's the real argument for FP4 on this hardware: not that the 70B wouldn't fit at FP8, but that at FP8 it answers at roughly half the speed.

These are ceilings, not benchmarks

Every number above is arithmetic — memory bandwidth divided by bytes read per token — not something we measured. Real throughput lands below it, because attention reads the KV cache too, sampling and detokenization cost something, and nothing achieves 100% of theoretical bandwidth. Treat the ceiling as the thing you cannot exceed and the shape of the trade-off as the thing worth trusting.

NVFP4 isn't the INT4 you remember

Four-bit has a bad reputation, and it earned it honestly. Early INT4 schemes mapped a whole group of weights onto sixteen evenly spaced integers with one shared scale. Neural network weights are not evenly distributed — a handful of outliers carry disproportionate weight — so a single scale stretched across a large group either clipped the outliers or wasted most of its range on them.

Blackwell's native format, NVFP4, fixes this from two directions. Each value is a 4-bit float (one sign bit, two exponent bits, one mantissa bit) rather than an integer, so its precision is finest near zero where most weights actually live. And the scaling is hierarchical: every block of 16 values gets its own FP8 scale factor, with a single FP32 scale applied across the tensor on top. Small blocks mean one outlier contaminates fifteen neighbours instead of hundreds, and an FP8 block scale can land on the value it needs rather than rounding to the nearest power of two the way MXFP4's coarser 32-value blocks do.

The practical upshot is that the second-generation Transformer Engine does this in hardware, on the chip, as part of the matmul — so the format isn't a compression trick you pay to unpack. It's how the numbers are stored and multiplied.

What it actually costs you

Quality loss from 4-bit isn't a flat tax. It concentrates, and knowing where lets you spend it deliberately:

The useful instinct: if your workload is retrieval-grounded and conversational, take the 4-bit speed. If it's agentic, long-horizon, or writing code you intend to run, test both precisions on your own prompts before deciding the faster one is good enough.

Picking a dial on 128 GB

Rules of thumb for a single GB10, assuming interactive single-user use:

None of which replaces trying it. The dial that matters is the one that holds up on your prompts, with your context, in your tooling — and that's a half-hour experiment, not a research project.

Test your own quantization.

Rent a full Grace Blackwell by the hour, load the precision you're arguing about, and run your own prompts against it.

Spin up a GB10 session