Rubricon.
Measurement infrastructure for rubric-based evaluation: reliability first, claims second. Every claim passes a signal gate that can refuse it.
4 tracks 168 items 504 responses 1,512 annotations 51.1% of claims blocked agentic: invest · grounding: iterate · reasoning: iterate · refusal: stop
SIM Simulated annotators — these numbers are not empirical findings
SIMULATED ANNOTATORS. No human labels were collected. Agreement statistics characterise the generative annotator model in rubricon.annotation.pool, not real annotator behaviour. The measurement machinery is real; the findings are demonstrations, not empirical claims about language models.

Portfolio overview

Four rubric-based evaluation tracks, each with its own depth contract. A track's depth contract fixes what may be claimed on it; the signal gate then decides, claim by claim, whether the measurement actually supports the claim.

tracks
4
evaluation programmes
items
168
distinct task items
responses
504
system outputs scored
annotations
1,512
rubric judgements
dimensions
21
rubric dimensions
failure codes
50
taxonomy entries

Tracks

Select a card to open the full track detail.

Side-by-side

One row per track. Reliability band is derived from mean alpha against the 0.50 / 0.667 / 0.80 policy thresholds.
trackdepthmean alphareliability bandgold acc.replicationrecommendationdetection blind spots
Agentic tool-use failure evaluation
agentic
production0.584warn79.3%3.00investnone
Grounding and citation integrity
grounding
pilot0.479block86.7%3.00iterateGF-12, GF-02
Reasoning process quality
reasoning
pilot0.559warn76.9%3.00iteratenone
Refusal calibration and over-refusal
refusal
exploratory0.355block78.4%3.00stopnone