JSPARK V2 / SEPTEMBER12, 2026
L5-P, with the measurements left in.
This public review companion reproduces the recorded result tables. L5-P remains an experimental serving build: functional code2/4 and short-behind-long TTFT40.458s remain unresolved. Measurement completion is not a blanket quality qualification. Raw local prompts, process logs and machine paths are intentionally not part of this public artifact.
Project page · Headline measurement data
Boot history at a glance
| Boot | Change | Headline | Outcome |
|---|---|---|---|
| K4 | Tony-derived EXL3, width 4 | 64K repeat 47.960 s; strong generation speed | Superseded for cache latency |
| L1 | Retention 512 | Stopped during loading to correct owner diagnostic hook | No model benchmark |
| L1b | Retention 512, prefill 4096 | 64K repeat 0.756 s; changed suffix 0.733 s | Cache fix selected |
| L2 | Prefill 1024 | Native image allowance requires at least 1025 tokens | Initialization refused; pivoted |
| L3 | Prefill 2048 | Smaller incumbent pauses; failed arriving-request TTFT improvement gate | Not selected |
| L4 | Fresh L1b configuration; broader checks | Prose 32.40–35.16, code 62.46–68.50 tok/s | Historical control |
| L5-S | Solo 4096; shared mixed prefill 2048 | Cold 64K 48.334 s; incumbent pauses 1.14–1.26 s | Scheduling retained |
| L5-B8192 | Solo 8192; mixed 2048 | Cold 64K 47.652 s; only 3.8% lower than control | Rejected below 5% gain threshold |
| L5-P | L5-S plus exact packed Engram buffered reads | Cold 64K 42.780 s; incumbent pauses 1.13–1.15 s | Live; full matrix and Work complete |
Long L1b and L5 cold/repeat figures are warmed-kernel three-trial medians. K4 is a historical control with different sample counts. Prose/code rates use full HTTP duration. S-r2 was a reconstruction for diagnostics, not another scored configuration. The canceled redundant L4 control is not a boot failure.
Measured single-stream output rates by boot
| Boot | Prose tok/s | Code tok/s | Sampling / state |
|---|---|---|---|
| Mia | 27.27–29.15 | 59.19–65.78 | One full response per task |
| E0 / initial EXL3 | 29.06–32.59 | 62.30–70.69 | One full response per task |
| K3 | 34.50–35.35 | 57.60–60.82 | Three-response task medians |
| K4 | 32.98–33.87 | 61.31–68.99 | Three-response task medians |
| L1b | 31.87–35.97 | 62.53–67.62 | One full response per task |
| L4 | 32.40–35.16 | 62.46–68.50 | One full response per task |
| L5-S / L5-B | Not measured | Not measured | Rapid latency screens only; full speed suite reserved for finalist |
| L5-P | 32.21–35.34 | 63.54–68.22 | One full response per task; code validation 2/4 |
Completion tokens / full HTTP duration, including TTFT. Ranges span three prose tasks or four code tasks. Output lengths and quality outcomes vary; these are not matched pure-decode rates. L1/L2 did not reach generation testing; L3 only received latency/overlap screening.
Experiment progression
| Candidate | Solo prefill | Mixed prefill budget | State |
|---|---|---|---|
| L4 historical control | 4096 | Shared original batch budget | Previously measured |
| L5-S | 4096 | 2048 total across prefills | Scheduling change retained; parent stopped for planned subboots |
| L5-B | 8192, with 6144 pivot if needed | Selected mixed cap | Rejected: 3.8% cold gain; no service failure |
| L5-P | 4096 | 2048 total across prefills | Live; full measurement complete, quality limitations recorded |
The new budget protects streams already decoding. It does not itself promise fast admission when a long prefill starts first. Failed or rejected candidates will remain visible.
First solo/cache screen · seconds to first output
| Measurement | L4 archived | L5-S first screen | L5-B first screen | L5-P first screen |
|---|---|---|---|---|
| 4K uncached | 2.273 s | 2.247 s | 2.156 s | 2.253 s |
| 4K repeat | 0.239 s | 0.233 s | 0.238 s | 0.222 s |
| 16K uncached | 9.106 s | 8.491 s | 15.503 s | 8.420 s |
| 16K repeat | 0.259 s | 0.254 s | 0.268 s | 0.241 s |
| 64K uncached / first-use | 54.332 s | 57.985 s | 56.043 s | 47.741 s |
| 64K repeat | 0.764 s | 0.765 s | 0.748 s | 0.699 s |
| 64K changed suffix | 0.729 s | 0.705 s | 0.714 s | 0.674 s |
One screen per boot; these are not medians. First-use64K includes cold runtime effects and is separate from the three salted warmed-kernel trials. L5-S changes mixed-traffic scheduling; solo prefill remains 4096. L5-B solo prefill is 8192; its first 16K request was slower. Repeated 64K trials determine budget selection. L5-P keeps solo 4096 / mixed 2048 and changes only the exact Engram storage reader/layout.
Warmed mixed traffic · medians of repeats 2 and3
| Case | L4 pause | L5-S pause | L4 new-request TTFT | L5-S new-request TTFT | L5-B pause | L5-B new-request TTFT | L5-P pause | L5-P new-request TTFT |
|---|---|---|---|---|---|---|---|---|
| 4K arrives / C2 | 2.146 s | 1.216 s | 2.439 s | 2.439 s | 1.148 s | 2.362 s | 1.152 s | 2.353 s |
| 4K arrives / C3 | 2.104 s | 1.137 s | 2.459 s | 2.419 s | 1.127 s | 2.389 s | 1.128 s | 2.386 s |
| 16K arrives / C2 | 2.104 s | 1.264 s | 8.707 s | 9.050 s | 1.139 s | 8.848 s | 1.126 s | 8.771 s |
| 16K arrives / C3 | 2.125 s | 1.177 s | 8.884 s | 8.983 s | 1.144 s | 8.905 s | 1.137 s | 8.942 s |
Pause is each incumbent stream’s maximum gap between output events, summarized by the worst incumbent per cell and then median across the two warmed repeats. L4/L1b share semantic configuration. Repeat 1 stays in the raw evidence; no outlier was removed from repeats 2/3.
L5-S completed screen
| Check | Result |
|---|---|
| 64K uncached / warmed-kernel median | 48.334 s vs 49.528 s historical; three salted trials |
| 64K exact repeat / changed suffix | 0.725 / 0.720 s |
| Native ID screen | 12/12 match L4; 5/5 repeat pairs exact |
| Shared scheduler cap | Native trace: 2048 maximum across two distinct prefills |
| Memory minima / Sparks 1, 2, 3 | 8.049 / 8.308 / 8.523 GiB during this screen |
| Swap /OOM /restart /preemption | None observed; preemption counter 0 |
This historical screen supported the next experiment. The selected L5-P finalist has now completed the full cache, quality, Pi, concurrency and Work matrices below. Auto-profiled logical KV capacity is 1,165,659 tokens for this boot; historical L4 had 1,444,887, so memory comparisons must account for allocation drift.
L5-B8192 decision
| Check | Result |
|---|---|
| 64K uncached / warmed-kernel median | 47.652 s vs 49.528 s historical |
| 64K exact repeat / changed suffix | 0.720 / 0.724 s |
| Mixed-traffic and latency safety gates | Passed; larger solo budget not promoted |
| Memory minima / Sparks 1, 2, 3 | 6.275 / 7.662 / 6.833 GiB |
| Decision | Keep solo 4096 / shared mixed 2048; test packed storage next |
Three salted trials completed. No resource failure required a 6144 pivot. Native ID promotion screen was not run after the performance rejection; the healthy boot is not described as fully qualified.
Engram diagnostic on selected L5-S settings
| Spark | Cold large-chunk row-read time | Cold large-chunk CPU gather time |
|---|---|---|
| 1 | 8.909 s | 9.662 s |
| 2 | 13.001 s | 14.476 s |
| 3 | 16.549 s | 17.974 s |
Unscored 64K cold TTFT: 56.871 s. These are sums over 16 full cold prefill chunks on each rank; warm/suffix records are excluded. Parallel rank times must not be added together or treated as guaranteed removable latency. The small cold tail and decode steps were below the trace threshold.
Packed Engram CPU read replay
| Spark | Original cold | Packed buffered cold | Original repeat | Packed buffered repeat | Packed direct repeat |
|---|---|---|---|---|---|
| 1 | 6.163 s | 3.157 s | 3.096 s | 1.595 s | 3.810 s |
| 2 | 9.158 s | 4.817 s | 4.346 s | 2.216 s | 5.326 s |
| 3 | 11.470 s | 6.234 s | 5.041 s | 2.587 s | 6.330 s |
Replay of the first eight recorded cold-prefill batches; one file-advised-cold pass and three repeats per mode. Same read functions, 32 threads / chunk 16, prestarted pool, 2 CPU quota and 4 GiB cgroup. Physical read bytes confirm cold I/O; exact output bytes agree. No API work occurred. Buffered passes; direct is rejected because warm reads regress 23–26%. These are CPU read times, not model TTFT.
L5-P repeated latency screen
| Check | Result |
|---|---|
| Cold 64K / warmed-kernel median | 42.780 s; trials 42.928, 42.780, 42.729 s |
| Change vs archived L4/L1b | 13.6% lower TTFT; control median 49.528 s |
| Change vs L5-S scheduling parent | 11.5% lower TTFT; parent median 48.334 s |
| 64K exact repeat / changed suffix | 0.694 / 0.684 s |
| Memory minima / Sparks 1, 2, 3 | 7.791 / 8.380 / 8.595 GiB |
| Latency and sampled resource gates | Rapid gates passed; full results recorded below |
These are three salted cold-prefix trials with warmed runtime kernels, separate from the first-use screen. Logical KV is 1,291,987 tokens on this boot versus 1,444,887 on historical L4; raw memory differences cannot all be attributed to packed storage. This table retains the rapid latency screen; full quality, concurrency and Work results are recorded below.
L5-P full qualification · repeated long latency
| Prompt | Cold TTFT | Exact repeat TTFT | Changed suffix TTFT |
|---|---|---|---|
| 64K | 42.623 s | 0.708 s | 0.691 s |
| 76K | 49.659 s | 0.471 s | 0.475 s |
Three independently salted trials per length, all 18responses passed. Separate from rapid screen; all native long and stable-tool companions now complete.
Long tool reuse and reverse arrival
| Check | L5-P result | Interpretation |
|---|---|---|
| Stable-tool 64K continuation | 0.595 s; 65,536 reused tokens | Native IDs and committed-prefix bound passed |
| Stable-tool 76K continuation | 0.550 s; 76,160 reused tokens | Native IDs and committed-prefix bound passed |
| Short prompt arrives two seconds into 64K prefill | 40.458 s short-request TTFT | Remaining admission weakness; one supplemental trial |
All pairs passed transport and task checks. Tool schemas are stable from the first request. Inverse arrival is a distinct ordering from the incumbent-decoder protection screen; no historical matched inverse control is claimed.
L5-P streamed full-answer speed · separate cohort
| Task | Output tokens | TTFT | Whole-request tok/s | After-first-output tok/s estimate |
|---|---|---|---|---|
| prose-explanation | 426 | 0.216 s | 33.78 | 34.28 |
| prose-memo | 478 | 0.237 s | 33.51 | 34.01 |
| prose-story | 546 | 0.220 s | 32.83 | 33.21 |
| ttl_lru_cache | 430 | 0.425 s | 70.53 | 75.64 |
| interval_rooms | 432 | 0.426 s | 64.74 | 68.99 |
| structured_log_summary | 562 | 0.411 s | 68.14 | 71.59 |
| dependency_waves | 541 | 0.396 s | 64.26 | 67.31 |
Same application payload as the original nonstream cohort, separately salted and streamed. All seven outputs finished with stop; texts differ across the two passes and are not averaged together. After-first-output estimate uses (tokens − 1)/(HTTP time − TTFT), including transport/finalization. Speculative chunks may hold multiple tokens; this is not GPU-only decode timing. Functional quality scores come from the original nonstream cohort.
Pi validation · L4 vs L5-P
| Measurement | L4 historical | L5-P |
|---|---|---|
| Three requests / two reads wall time | 16.753 s | 15.436 s |
| Aggregate engine prefix hits | 14,848 | 14,848 |
| 64K history off wall time | 52.654 s | 47.241 s |
| 64K history max wall time | 53.517 s | 48.158 s |
| Actual output cap | 65,536 tokens | 65,536 tokens |
One run per cell. Installed Pi expands synthetic long inputs to 71,030/71,058 tokens; these are supplied histories, not measured warm interactive turns. Pi usage still omits cache details; isolated engine counter deltas establish aggregate reuse. Both successful reads and exact final marker were checked.
L5-P concurrency · completed C1–C6
| Offered streams | L4 aggregate tok/s | Mia aggregate tok/s | Tony published tok/s | L5-P aggregate tok/s | L5-P mean TTFT |
|---|---|---|---|---|---|
| C1 | 48.26 | 42.54 | 46.00 | 49.05 | 0.241 s |
| C2 | 73.06 | 66.48 | 73.37 | 73.55 | 0.372 s |
| C3 | 92.88 | 90.17 | 100.14 | 97.21 | 0.450 s |
| C4 | 112.92 | 106.69 | 118.41 | 114.16 | 0.517 s |
| C5 | 132.29 | 80.60 | 134.44 | 131.46 | 0.564 s |
| C6 | 143.43 | 87.34 | 152.89 | 146.12 | 0.611 s |
All 189 primary streams passed; three warmups and six separate coding-repeat streams also passed. Eight-category arithmetic mean of aggregate completion tokens / shared HTTP window, counting task excluded. C6 coding first/repeat: 213.70 / 228.10 tok/s, max TTFT 0.659 / 0.555 s; original matrix average remains 146.12. No swap/OOM/restart/preemption, sampled waiting 0 and running max 6. Tony reference is published; Mia cap 4 affects offered C5/C6.
L5-P Work · all four rounds
| Round | L4 backend tok/s | L5-P backend tok/s | L4 initial TTFT | L5-P initial TTFT | L5-P tools |
|---|---|---|---|---|---|
| c3-first | 80.70 | 80.12 | 2.119 s | 2.122 s | 26 |
| c3-repeat | 79.54 | 81.98 | 0.676 s | 0.599 s | 25 |
| c6-first | 100.72 | 102.79 | 2.347 s | 2.283 s | 45 |
| c6-repeat | 103.00 | 110.49 | 0.878 s | 0.864 s | 42 |
All 18 initial payloads match L4 and all 9 repeat pairs match. All 18 sessions were deliberately cut at 180 seconds; no model/tool API errors, 3 nonzero shell exits, no pending first output at cutoff. Backend counters include unfinished accepted output. Later agent trajectories differ; this is one cohort, not an output-quality score or universal speed claim. Work waiting max 3/1 at C6 first/repeat; preemptions 0.
Work: four recipes, first and repeat
Accepted backend output tokens / fixed180-second window. Deliberately unfinished tasks; no output-quality score. Mia is the counted rerun, with active cap4; local EXL3 lanes allow8. Later tool trajectories differ. First/repeat is one cohort, not a causal warming experiment.
| Recipe | Round | Backend tok/s | Initial TTFT, s | Later median TTFT, s |
|---|---|---|---|---|
| L4 | c3-first | 80.70 | 2.119 | 1.233 |
| L4 | c3-repeat | 79.54 | 0.676 | 1.246 |
| L4 | c6-first | 100.72 | 2.347 | 1.328 |
| L4 | c6-repeat | 103.00 | 0.878 | 1.708 |
| Tony | c3-first | 76.36 | 2.485 | 1.752 |
| Tony | c3-repeat | 73.04 | 2.125 | 1.435 |
| Tony | c6-first | 101.69 | 3.561 | 1.416 |
| Tony | c6-repeat | 92.82 | 3.528 | 1.539 |
| Mia | c3-first | 68.15 | 2.667 | 1.227 |
| Mia | c3-repeat | 71.83 | 2.587 | 1.239 |
| Mia | c6-first | 82.71 | 4.309 | 4.135 |
| Mia | c6-repeat | 87.17 | 1.510 | 3.301 |
| L5 | c3-first | 80.12 | 2.122 | 1.274 |
| L5 | c3-repeat | 81.98 | 0.599 | 1.355 |
| L5 | c6-first | 102.79 | 2.283 | 1.655 |
| L5 | c6-repeat | 110.49 | 0.864 | 1.518 |