← Back to the playground · this is the detailed study behind it (calibration, held-out thresholds, confidence intervals).

System 1 / System 2: confidence-gated routing, measured

A fast, cheap System 1 classifies first and reports a confidence. A strong System 2 LLM is called only when that confidence is below a threshold. This page replays recorded results from public data; no key, no live call. Code and reproduction: github.com/Iskandeur/system1-system2.

Loading results…

The router

Input text
e.g. a voice-assistant request, a ticket, an issue
→
System 1
typed-decision model, ~300 ms
label + confidence
→
confidence ≥ t ?
yes → answer
no → System 2 (LLM)

Every item is scored by both systems once, so every threshold, and the two extremes (System 1 only at t = 0, System 2 only at t = 1), come from the same predictions.

Accuracy and cost as the threshold moves

0.50
hybrid accuracySystem 2 onlySystem 1 onlyescalation rate
hybrid, one dot per thresholdSystem 1 onlySystem 2 only

Held-out operating point

StrategyAccuracy (95% CI)Cost / 1k itemsMean latencyEscalation

Is the confidence worth routing on?

A reliability diagram puts items in ten confidence bins and compares each bin's mean confidence with its accuracy. Bars on the diagonal mean "when it says 80 %, it is right 80 % of the time". ECE is the count-weighted gap; AUROC says how well confidence ranks right answers above wrong ones (0.5 = useless).

Reproduce

git clone https://github.com/Iskandeur/system1-system2 && cd system1-system2
cp .env.example .env            # OPENROUTER_API_KEY is enough for the default pair
node scripts/run-eval.mjs --limit 20
node scripts/analyze.mjs --dataset massive-en      # recompute every number on this page from data/predictions/

All predictions, the datasets and the analysis code are in the repository. Methods, limitations and sources are in the README.