ci_leaderboard recipe: judge-free, 30 rows/axis, small-n directional scores.
| # | model | Elo (BT-MLE) | MS (public) | MS (experimental ⚠) | knowledge | nlu | cultural | axes | run date |
|---|---|---|---|---|---|---|---|---|---|
| 1 | qwen2.5-1.5b-instruct | 1009.0 | 0.3667 [0.23,0.48] | 0.2759 ⚠ [0.28,0.28] | 0.533 n=30 | 0.200 n=30 | 0.276 ⚠ n=29 | 3 | 2026-07-25 |
| 2 | smollm2-1.7b-instruct | 1001.8 | 0.3333 [0.22,0.43] | 0.3448 ⚠ [0.34,0.34] | 0.500 n=30 | 0.167 n=30 | 0.345 ⚠ n=29 | 3 | 2026-07-25 |
| 3 | qwen2.5-0.5b-instruct | 989.2 | 0.3500 [0.23,0.45] | 0.2414 ⚠ [0.24,0.24] | 0.533 n=30 | 0.167 n=30 | 0.241 ⚠ n=29 | 3 | 2026-07-25 |
⚠ Experimental = self-built axes below the measurement-grade item floor: authored by the benchmark team, not yet independently human-validated, small-n (wide CI). DIRECTIONAL ONLY — a rank-order hint, never a citable score. Kept visible (not hidden) so progress is trackable while the native-annotation program brings them to headline grade (n≥200, κ≥0.7). Experimental axes here: cultural.
118 verifiable pairwise matches fed this board. Elo ranking is BT-MLE (order-invariant); a model is credibly above another only when their bootstrap CIs don't overlap — see glossobench elo for full CIs. This page is generated by CI from a small, judge-free, tiny-row-cap recipe (directional / demo only, not a real GlossoBench-5/7 claim) — see glossobench for the real harness.