What we're seeing as the NVIDIA Grace Blackwell GB10 and DGX Spark move from keynote stages to desks — and what it means for everyone who wants to own their inference instead of renting it from a hyperscaler.
On a bandwidth-bound machine, quantization is the main dial on how fast a model answers — not just whether it fits. The memory math for a 70B, what NVFP4 does differently, and where quality actually goes missing.
Read the deep dive →A petaflop of FP4 and 128GB of coherent memory on a desk you can carry. The Spark isn't a workstation — it's a category reset.
03Per-token cloud pricing looks cheap until you run the numbers at volume. What an owned GB10 actually costs per hour — and where the break-even sits.
04NVLink-C2C, a coherent CPU–GPU address space, and why "unified memory" is the spec that actually changes which models you can run.
05Project DIGITS became the DGX Spark, partners shipped their own GB10 systems, and the software stack caught up fast. A timeline.
06What it's actually like to serve Llama 3.3 70B from a single Grace Blackwell — context limits, tokens per second, and where it shines.
Multi-GB10 clustering and KV-cache tuning for long contexts. Subscribe from your dashboard to get them first.