GlossoBench — ms leaderboard v1.0

ci_leaderboard recipe: judge-free, 30 rows/axis, small-n directional scores.

#modelElo (BT-MLE)MS (public)MS (experimental ⚠)knowledgenluculturalaxesrun date
1qwen2.5-1.5b-instruct1009.00.3667 [0.23,0.48]0.2759 ⚠ [0.28,0.28]0.533 n=300.200 n=300.276 ⚠ n=2932026-07-25
2smollm2-1.7b-instruct1001.80.3333 [0.22,0.43]0.3448 ⚠ [0.34,0.34]0.500 n=300.167 n=300.345 ⚠ n=2932026-07-25
3qwen2.5-0.5b-instruct989.20.3500 [0.23,0.45]0.2414 ⚠ [0.24,0.24]0.533 n=300.167 n=300.241 ⚠ n=2932026-07-25

Experimental = self-built axes below the measurement-grade item floor: authored by the benchmark team, not yet independently human-validated, small-n (wide CI). DIRECTIONAL ONLY — a rank-order hint, never a citable score. Kept visible (not hidden) so progress is trackable while the native-annotation program brings them to headline grade (n≥200, κ≥0.7). Experimental axes here: cultural.

118 verifiable pairwise matches fed this board. Elo ranking is BT-MLE (order-invariant); a model is credibly above another only when their bootstrap CIs don't overlap — see glossobench elo for full CIs. This page is generated by CI from a small, judge-free, tiny-row-cap recipe (directional / demo only, not a real GlossoBench-5/7 claim) — see glossobench for the real harness.