Recipe comparison methods
Historical September 2026 measurements. Tempo is recorded as L5-P in the original experiments. These charts compare specific recipes and workloads, not the latest releases or overall model quality.
Short speed screen
Tempo and Mia were measured on our three DGX Sparks. The short-screen chart uses aggregate end-to-end throughput: total completion tokens across the simultaneous wave divided by the time from the first HTTP request start to the last HTTP completion. It includes initial wait. It is not per-stream decode speed.
There is one wave per category and concurrency, with 150–256-token output caps. The eight-category mean is the arithmetic mean of the category wave rates, excluding the separate counting ceiling. Code and prose views show individual categories. These are speed fixtures, not answer-quality grades.
Mia's active request cap was 4, versus Tempo's 8. The short-screen comparison stops at four offered streams, within Mia's active-request cap. Runtime, model storage, quantization, speculative decoding, and configuration differ together; these are recipe-level comparisons, not isolated kernel or quantization experiments.
Tony's eight-category series is the author-published TP3 reference captured in our result records. It comes from his three Sparks, not a short-screen rerun on ours. His aggregate definition also includes the batch wall time and initial wait, but hardware state and recipe settings differ. Published figures do not establish a controlled performance ranking. The chart preserves Tony's higher reported C3–C4 means.
The chart opens at Code / C1. Each line chart fits its Y-axis to all plotted C1–C4 values, rounded outward to tens, and labels the nonzero range. The range stays fixed when selecting a stream count. Legend keys share the lines' colors, dash patterns, and circle markers. Bar charts keep a zero baseline.
Three-minute agent tasks
This is a separate local comparison: Tempo, a locally adapted Tony recipe, and Mia's counted rerun were all measured on our three Sparks. The rate is accepted backend output tokens divided by the fixed 180-second task window, including unfinished output, health probes, and small launch/drain overhead. It is not GPU-only decode speed.
The Work comparison shows three concurrent agents, with first and repeat rounds. Initial task payloads match, but later tool histories diverge. Repeat rounds use fresh Pi sessions with caches intact. One cohort does not establish a causal cache-speed effect, statistical significance, or a productivity gain. Output quality and task completion are not scored.
Three concurrent agents fit within every recipe's active-request cap. The Tony local run includes required fabric, storage, loader, and allocator adaptations and a shared local EXL3 kernel clamp fallback. It is not an exact reproduction of his published fleet binaries. Mia's counted rerun supplies generation counters that were absent from the preliminary Work run; the preliminary unavailable values are not used.
Recipes and credit
Tempo builds on Tony and Kai's DeepSeek Spark serving recipe and bot-lab-21's model files, made smaller with WestWaters' Pollard method (EXL3). Mia's three-Spark recipe provides a comparison. Its ideas for combining model calculations and sharing memory helped inform Tempo; Tempo does not reuse its code. DeepSeek created the model. This comparison does not imply endorsement by another recipe author.
- Mia's recipe · Tested software version: de57fe2b.
- Tony's recipe · Reference software version: c2c1bb7. The cited instructions are for three Sparks (TP3); the repository also describes a four-Spark setup.
Model downloads
These recipes use different versions of DeepSeek. The software versions are linked above. This table lists the model files each setup downloads.
| Recipe | Model files it downloads |
|---|---|
| Mia | deepseek-ai/DeepSeek-V4.1-Flash, fb2764a. The setup instructions select this exact model version. |
| Tempo | bot-lab-21/DeepSeek-V4.1-Flash-EXL3-3.5bpw-Pollard, b60193e. Tempo selects this exact version for all its model downloads, including the smaller helper model. |
| Tony and Kai | The cited three-Spark download script fetches 40 changed model files from bot-lab-21 and reuses eight unchanged files from an existing official DeepSeek download. It does not select a fixed version. Our saved download records identify b60193e, but the script itself does not guarantee that version. |
Tempo does not require a separate official DeepSeek download. That release is named to record where the model comes from. Download sources and file checks. This clarification does not change the historical measurements.
Evidence
Exact chart values and source SHA-256 hashes are a selected, sanitized extraction of the saved result artifacts, with C1–C4 short-screen comparisons and C3 Work comparisons. No live model requests were made for this page.
- Short screen: Mia
M-pre-tony-c4to6/TONY-RESULTS.jsonand Tempol5-qualification/TONY-RESULTS.json. Mia's eight-category series uses the shared-window recomputation in the Tempo record. Tony's series is itsTony_publishedfield. - Work: Tempo
l4-followup-20260912/L5-WORK-RESULTS.json; local Tonywork-benchmark/WORK-RESULTS.json; Miawork-benchmark/Mia-counted/{round}/RESULT.json. Mia's generation-token counter deltas are summed and divided by 180 seconds.
Full historical report · Tempo release measurements
The page also retains Tempo's latency measurements and known limitations. Faster output on one workload does not imply faster uncached prefill or better answers.