← Back to the playground · this is the detailed study behind it (calibration, held-out thresholds, confidence intervals).
System 1 / System 2: confidence-gated routing, measured
A fast, cheap System 1 classifies first and reports a confidence. A strong System 2 LLM is
called only when that confidence is below a threshold. This page replays recorded results from public data;
no key, no live call. Code and reproduction: github.com/Iskandeur/system1-system2.
Loading results…
The router
Input text e.g. a voice-assistant request, a ticket, an issue
→
System 1 typed-decision model, ~300 ms label + confidence
→
confidence ≥ t ? yes → answer no → System 2 (LLM)
Every item is scored by both systems once, so every threshold, and the two extremes (System 1 only at t = 0, System 2 only at t = 1), come from the same predictions.
hybrid, one dot per thresholdSystem 1 onlySystem 2 only
Held-out operating point
Strategy
Accuracy (95% CI)
Cost / 1k items
Mean latency
Escalation
Is the confidence worth routing on?
A reliability diagram puts items in ten confidence bins and compares each bin's mean confidence with its accuracy. Bars on the diagonal mean "when it says 80 %, it is right 80 % of the time". ECE is the count-weighted gap; AUROC says how well confidence ranks right answers above wrong ones (0.5 = useless).
Prompt injection through the router
System
Defense
Targeted success
Flipped vs clean
Control flip
Confidence clean → attacked
Successful & above t
Through the router
Pair
Defense
t
Targeted success
Escalated under attack
Passed the router
System 1 · defense
t
Successful attacks kept (conf ≥ t)
Escalation clean → attacked
Hybrid targeted success
Passed the router
The six templates
Template
Appended payload
Same items, in French
System
EN accuracy
FR accuracy
Gap
EN ECE
FR ECE
Open-weight System 1, run locally on CPU
System
Runs on
EN accuracy
FR accuracy
FR − EN
Median latency
ECE EN / FR
Cost / 1k
Reproduce
git clone https://github.com/Iskandeur/system1-system2 && cd system1-system2
cp .env.example .env # OPENROUTER_API_KEY is enough for the default pair
node scripts/run-eval.mjs --limit 20
node scripts/analyze.mjs --dataset massive-en # recompute every number on this page from data/predictions/
All predictions, the datasets and the analysis code are in the repository. Methods, limitations and sources are in the README.