BENCHMARKS · QUALITY
Canonical quality
Dataset
Five hundred positives cover noun, pronoun, numeral, verb, adjective, determiner, and adverb target spans. Five hundred negatives include identical surface forms that serve other morphological functions.
Only manually reviewed standard sentences are included; noisy sentences remain in Robust.
Metrics
Each backend reports a raw confusion matrix plus precision, recall, and F1. No contract review applies, so the adjusted matrices equal raw and the reviewed-case count is zero.
The D3 chart renders both series rather than hiding their equality.
Interpretation limits
Canonical F1 measures lemma search on standard sentences. It includes neither file-scan throughput nor noisy-sentence robustness.
POS strata have different denominators, so their percentages are not averaged directly into the overall result.