VERITAS · interactive research note

A fairness gap you can measure with no unfairness present

When the reference standard is imperfect and disease prevalence differs between two groups, a benchmark reports a subgroup performance gap for a model that is exactly equally good in both. Move the panel below: the number on the right is the gap you would measure for an identical model — the artifact floor. A reported gap beneath it has measured the labels, not the model.

0.90
P(reader marks positive | truly positive), per reader.
0.90
The weak axis in real reads: reference false positives drive the artifact.
30%
10%
3 (2+1)
Two readers plus a third on disagreement is algebraically a 3-reader majority — the field standard.
artifact floor · apparent sensitivity gap
0.114
For a model identical in both groups. The true gap is zero.
constructed reference: Seref 0.972 Spref 0.972
0.05
below its floor

Artifact floor vs. panel size at the current reader accuracy and prevalences. Readers are assumed conditionally independent — the best case, since readers who share an image share their mistakes. The dot marks your current panel.

Measured on four real radiologists

Reader accuracy above is a dial. It doesn't have to be. LIDC-IDRI had four thoracic radiologists read 1,016 CT scans independently and releases the annotations without the images — the one ungated source of real per-reader labels. Binary task: is a nodule ≥3 mm present?

28.5%of scans on which the four radiologists disagree about whether a ≥3 mm nodule is present at all. Unanimity on 726 of 1,016.
0.239artifact floor at their measured accuracy, 3-reader reference, for a 30%-vs-10% prevalence contrast — worse than the parametric guess. (Press the LIDC button above.)
Latent-class reader estimates, LIDC-IDRI — identical across 12 random starts.
readersensitivityspecificity
10.9600.801
20.9720.800
30.9360.739
40.9230.806
mean0.9480.787

The conditional-independence assumption the whole construction rests on is rejected on this same data: G² = 44.1 on 6 df, p ≈ 7e-8 (parametric bootstrap p ≤ 0.0005). LIDC's phase-2 reads were made with the other radiologists' marks visible, so dependence is expected by design — the point is that the canonical public multi-reader dataset violates the assumption every consensus label built from it silently relies on.

What is known, what is new, what is open

The prevalence contrasts are illustrative: LIDC's own prevalence is 0.72, so the 0.239 figure pairs measured accuracy with hypothetical subgroup prevalences — stated, not hidden. No published finding is claimed here to be wrong; whether a specific reported disparity is an artifact is an empirical question needing that study's reader data.