JSPARK V2 / SEPTEMBER12, 2026

L5-P, with the measurements left in.

This public review companion reproduces the recorded result tables. L5-P remains an experimental serving build: functional code2/4 and short-behind-long TTFT40.458s remain unresolved. Measurement completion is not a blanket quality qualification. Raw local prompts, process logs and machine paths are intentionally not part of this public artifact.

Project page · Headline measurement data

Boot history at a glance

BootChangeHeadlineOutcome
K4Tony-derived EXL3, width 464K repeat 47.960 s; strong generation speedSuperseded for cache latency
L1Retention 512Stopped during loading to correct owner diagnostic hookNo model benchmark
L1bRetention 512, prefill 409664K repeat 0.756 s; changed suffix 0.733 sCache fix selected
L2Prefill 1024Native image allowance requires at least 1025 tokensInitialization refused; pivoted
L3Prefill 2048Smaller incumbent pauses; failed arriving-request TTFT improvement gateNot selected
L4Fresh L1b configuration; broader checksProse 32.40–35.16, code 62.46–68.50 tok/sHistorical control
L5-SSolo 4096; shared mixed prefill 2048Cold 64K 48.334 s; incumbent pauses 1.14–1.26 sScheduling retained
L5-B8192Solo 8192; mixed 2048Cold 64K 47.652 s; only 3.8% lower than controlRejected below 5% gain threshold
L5-PL5-S plus exact packed Engram buffered readsCold 64K 42.780 s; incumbent pauses 1.13–1.15 sLive; full matrix and Work complete

Long L1b and L5 cold/repeat figures are warmed-kernel three-trial medians. K4 is a historical control with different sample counts. Prose/code rates use full HTTP duration. S-r2 was a reconstruction for diagnostics, not another scored configuration. The canceled redundant L4 control is not a boot failure.

Measured single-stream output rates by boot

BootProse tok/sCode tok/sSampling / state
Mia27.27–29.1559.19–65.78One full response per task
E0 / initial EXL329.06–32.5962.30–70.69One full response per task
K334.50–35.3557.60–60.82Three-response task medians
K432.98–33.8761.31–68.99Three-response task medians
L1b31.87–35.9762.53–67.62One full response per task
L432.40–35.1662.46–68.50One full response per task
L5-S / L5-BNot measuredNot measuredRapid latency screens only; full speed suite reserved for finalist
L5-P32.21–35.3463.54–68.22One full response per task; code validation 2/4

Completion tokens / full HTTP duration, including TTFT. Ranges span three prose tasks or four code tasks. Output lengths and quality outcomes vary; these are not matched pure-decode rates. L1/L2 did not reach generation testing; L3 only received latency/overlap screening.

Experiment progression

CandidateSolo prefillMixed prefill budgetState
L4 historical control4096Shared original batch budgetPreviously measured
L5-S40962048 total across prefillsScheduling change retained; parent stopped for planned subboots
L5-B8192, with 6144 pivot if neededSelected mixed capRejected: 3.8% cold gain; no service failure
L5-P40962048 total across prefillsLive; full measurement complete, quality limitations recorded

The new budget protects streams already decoding. It does not itself promise fast admission when a long prefill starts first. Failed or rejected candidates will remain visible.

First solo/cache screen · seconds to first output

MeasurementL4 archivedL5-S first screenL5-B first screenL5-P first screen
4K uncached2.273 s2.247 s2.156 s2.253 s
4K repeat0.239 s0.233 s0.238 s0.222 s
16K uncached9.106 s8.491 s15.503 s8.420 s
16K repeat0.259 s0.254 s0.268 s0.241 s
64K uncached / first-use54.332 s57.985 s56.043 s47.741 s
64K repeat0.764 s0.765 s0.748 s0.699 s
64K changed suffix0.729 s0.705 s0.714 s0.674 s

One screen per boot; these are not medians. First-use64K includes cold runtime effects and is separate from the three salted warmed-kernel trials. L5-S changes mixed-traffic scheduling; solo prefill remains 4096. L5-B solo prefill is 8192; its first 16K request was slower. Repeated 64K trials determine budget selection. L5-P keeps solo 4096 / mixed 2048 and changes only the exact Engram storage reader/layout.

Warmed mixed traffic · medians of repeats 2 and3

CaseL4 pauseL5-S pauseL4 new-request TTFTL5-S new-request TTFTL5-B pauseL5-B new-request TTFTL5-P pauseL5-P new-request TTFT
4K arrives / C22.146 s1.216 s2.439 s2.439 s1.148 s2.362 s1.152 s2.353 s
4K arrives / C32.104 s1.137 s2.459 s2.419 s1.127 s2.389 s1.128 s2.386 s
16K arrives / C22.104 s1.264 s8.707 s9.050 s1.139 s8.848 s1.126 s8.771 s
16K arrives / C32.125 s1.177 s8.884 s8.983 s1.144 s8.905 s1.137 s8.942 s

Pause is each incumbent stream’s maximum gap between output events, summarized by the worst incumbent per cell and then median across the two warmed repeats. L4/L1b share semantic configuration. Repeat 1 stays in the raw evidence; no outlier was removed from repeats 2/3.

L5-S completed screen

CheckResult
64K uncached / warmed-kernel median48.334 s vs 49.528 s historical; three salted trials
64K exact repeat / changed suffix0.725 / 0.720 s
Native ID screen12/12 match L4; 5/5 repeat pairs exact
Shared scheduler capNative trace: 2048 maximum across two distinct prefills
Memory minima / Sparks 1, 2, 38.049 / 8.308 / 8.523 GiB during this screen
Swap /OOM /restart /preemptionNone observed; preemption counter 0

This historical screen supported the next experiment. The selected L5-P finalist has now completed the full cache, quality, Pi, concurrency and Work matrices below. Auto-profiled logical KV capacity is 1,165,659 tokens for this boot; historical L4 had 1,444,887, so memory comparisons must account for allocation drift.

L5-B8192 decision

CheckResult
64K uncached / warmed-kernel median47.652 s vs 49.528 s historical
64K exact repeat / changed suffix0.720 / 0.724 s
Mixed-traffic and latency safety gatesPassed; larger solo budget not promoted
Memory minima / Sparks 1, 2, 36.275 / 7.662 / 6.833 GiB
DecisionKeep solo 4096 / shared mixed 2048; test packed storage next

Three salted trials completed. No resource failure required a 6144 pivot. Native ID promotion screen was not run after the performance rejection; the healthy boot is not described as fully qualified.

Engram diagnostic on selected L5-S settings

SparkCold large-chunk row-read timeCold large-chunk CPU gather time
18.909 s9.662 s
213.001 s14.476 s
316.549 s17.974 s

Unscored 64K cold TTFT: 56.871 s. These are sums over 16 full cold prefill chunks on each rank; warm/suffix records are excluded. Parallel rank times must not be added together or treated as guaranteed removable latency. The small cold tail and decode steps were below the trace threshold.

Packed Engram CPU read replay

SparkOriginal coldPacked buffered coldOriginal repeatPacked buffered repeatPacked direct repeat
16.163 s3.157 s3.096 s1.595 s3.810 s
29.158 s4.817 s4.346 s2.216 s5.326 s
311.470 s6.234 s5.041 s2.587 s6.330 s

Replay of the first eight recorded cold-prefill batches; one file-advised-cold pass and three repeats per mode. Same read functions, 32 threads / chunk 16, prestarted pool, 2 CPU quota and 4 GiB cgroup. Physical read bytes confirm cold I/O; exact output bytes agree. No API work occurred. Buffered passes; direct is rejected because warm reads regress 23–26%. These are CPU read times, not model TTFT.

L5-P repeated latency screen

CheckResult
Cold 64K / warmed-kernel median42.780 s; trials 42.928, 42.780, 42.729 s
Change vs archived L4/L1b13.6% lower TTFT; control median 49.528 s
Change vs L5-S scheduling parent11.5% lower TTFT; parent median 48.334 s
64K exact repeat / changed suffix0.694 / 0.684 s
Memory minima / Sparks 1, 2, 37.791 / 8.380 / 8.595 GiB
Latency and sampled resource gatesRapid gates passed; full results recorded below

These are three salted cold-prefix trials with warmed runtime kernels, separate from the first-use screen. Logical KV is 1,291,987 tokens on this boot versus 1,444,887 on historical L4; raw memory differences cannot all be attributed to packed storage. This table retains the rapid latency screen; full quality, concurrency and Work results are recorded below.

L5-P full qualification · repeated long latency

PromptCold TTFTExact repeat TTFTChanged suffix TTFT
64K42.623 s0.708 s0.691 s
76K49.659 s0.471 s0.475 s

Three independently salted trials per length, all 18responses passed. Separate from rapid screen; all native long and stable-tool companions now complete.

Long tool reuse and reverse arrival

CheckL5-P resultInterpretation
Stable-tool 64K continuation0.595 s; 65,536 reused tokensNative IDs and committed-prefix bound passed
Stable-tool 76K continuation0.550 s; 76,160 reused tokensNative IDs and committed-prefix bound passed
Short prompt arrives two seconds into 64K prefill40.458 s short-request TTFTRemaining admission weakness; one supplemental trial

All pairs passed transport and task checks. Tool schemas are stable from the first request. Inverse arrival is a distinct ordering from the incumbent-decoder protection screen; no historical matched inverse control is claimed.

L5-P streamed full-answer speed · separate cohort

TaskOutput tokensTTFTWhole-request tok/sAfter-first-output tok/s estimate
prose-explanation4260.216 s33.7834.28
prose-memo4780.237 s33.5134.01
prose-story5460.220 s32.8333.21
ttl_lru_cache4300.425 s70.5375.64
interval_rooms4320.426 s64.7468.99
structured_log_summary5620.411 s68.1471.59
dependency_waves5410.396 s64.2667.31

Same application payload as the original nonstream cohort, separately salted and streamed. All seven outputs finished with stop; texts differ across the two passes and are not averaged together. After-first-output estimate uses (tokens − 1)/(HTTP time − TTFT), including transport/finalization. Speculative chunks may hold multiple tokens; this is not GPU-only decode timing. Functional quality scores come from the original nonstream cohort.

Pi validation · L4 vs L5-P

MeasurementL4 historicalL5-P
Three requests / two reads wall time16.753 s15.436 s
Aggregate engine prefix hits14,84814,848
64K history off wall time52.654 s47.241 s
64K history max wall time53.517 s48.158 s
Actual output cap65,536 tokens65,536 tokens

One run per cell. Installed Pi expands synthetic long inputs to 71,030/71,058 tokens; these are supplied histories, not measured warm interactive turns. Pi usage still omits cache details; isolated engine counter deltas establish aggregate reuse. Both successful reads and exact final marker were checked.

L5-P concurrency · completed C1–C6

Offered streamsL4 aggregate tok/sMia aggregate tok/sTony published tok/sL5-P aggregate tok/sL5-P mean TTFT
C148.2642.5446.0049.050.241 s
C273.0666.4873.3773.550.372 s
C392.8890.17100.1497.210.450 s
C4112.92106.69118.41114.160.517 s
C5132.2980.60134.44131.460.564 s
C6143.4387.34152.89146.120.611 s

All 189 primary streams passed; three warmups and six separate coding-repeat streams also passed. Eight-category arithmetic mean of aggregate completion tokens / shared HTTP window, counting task excluded. C6 coding first/repeat: 213.70 / 228.10 tok/s, max TTFT 0.659 / 0.555 s; original matrix average remains 146.12. No swap/OOM/restart/preemption, sampled waiting 0 and running max 6. Tony reference is published; Mia cap 4 affects offered C5/C6.

L5-P Work · all four rounds

RoundL4 backend tok/sL5-P backend tok/sL4 initial TTFTL5-P initial TTFTL5-P tools
c3-first80.7080.122.119 s2.122 s26
c3-repeat79.5481.980.676 s0.599 s25
c6-first100.72102.792.347 s2.283 s45
c6-repeat103.00110.490.878 s0.864 s42

All 18 initial payloads match L4 and all 9 repeat pairs match. All 18 sessions were deliberately cut at 180 seconds; no model/tool API errors, 3 nonzero shell exits, no pending first output at cutoff. Backend counters include unfinished accepted output. Later agent trajectories differ; this is one cohort, not an output-quality score or universal speed claim. Work waiting max 3/1 at C6 first/repeat; preemptions 0.

Work: four recipes, first and repeat

Accepted backend output tokens / fixed180-second window. Deliberately unfinished tasks; no output-quality score. Mia is the counted rerun, with active cap4; local EXL3 lanes allow8. Later tool trajectories differ. First/repeat is one cohort, not a causal warming experiment.

RecipeRoundBackend tok/sInitial TTFT, sLater median TTFT, s
L4c3-first80.702.1191.233
L4c3-repeat79.540.6761.246
L4c6-first100.722.3471.328
L4c6-repeat103.000.8781.708
Tonyc3-first76.362.4851.752
Tonyc3-repeat73.042.1251.435
Tonyc6-first101.693.5611.416
Tonyc6-repeat92.823.5281.539
Miac3-first68.152.6671.227
Miac3-repeat71.832.5871.239
Miac6-first82.714.3094.135
Miac6-repeat87.171.5103.301
L5c3-first80.122.1221.274
L5c3-repeat81.980.5991.355
L5c6-first102.792.2831.655
L5c6-repeat110.490.8641.518