JSPARK3:: GLM-5.3 Flash recipe for 3x NVIDIA DGX Sparks

4 views
Three DGX Sparks serving one GLM-5.3 Flash endpoint. The recipes I kept seeing were for two Sparks or four, and the usual line was that three did not make sense. Here is the recipe and why it works.

The recipes I kept seeing were for two Sparks or four. GLM-5.3 Flash runs well on two DGX Sparks, and the community recipes prove it. Four is the next size people talk about. The usual line on three was that it did not make sense. I had three, and I wanted three to make sense: each Spark a full peer, not two workers and a spare. JSpark3 is what came out of that.

JSpark3 is a serving recipe, not a model. It runs one GLM-5.3 Flash endpoint across three Sparks, tensor parallel 3 and expert parallel 3. Every input is pinned: the checkpoint, the draft, the container image, the transforms. If a byte drifts, it refuses to start. The evidence is published with the misses left in, including two internal gates the release did not pass.

The numbers I trust most came from other people's benchmarks. I ran FlyCockpit's own three-Spark benchmark script against JSpark3, changing only the endpoint: about 1.2x their published decode rate on all three prompts, with slightly higher draft acceptance. On Mia's bench_decode, JSpark3 on three Sparks ran about 1.35x Mia's published two-Spark numbers, with acceptance within a point of theirs. Our own frozen screen agrees. Each number sits next to its conditions in the full write-up, including the two-stream row that lands below Mia's figure.

The full write-up. Architecture, evidence, reproducibility, provenance, and licensing are at jakejh.com/jspark3. The code and results are on GitHub. The Hugging Face repo has the model card, the provenance chain, and a shard-by-shard manifest of the exact target weights, with the upstream license and attribution intact.

Credit. JSpark3 builds on FlyCockpit's three-Spark loader, MiaAI-Lab's two-Spark recipe, vcruz305's fixes, Brandon M. Music's EXL3/TR3 quantization as published by Mia-AiLab, Inco AI's DFlash2 draft, and Z.AI's GLM-5.3 Flash.

If you have three Sparks, try it out. Run the preflight, it’ll tell you what is wrong with your fleet before anything starts. It’s a feature.