A fairness gap you can measure with no unfairness present
When the reference standard is imperfect and disease prevalence differs
between two groups, a benchmark reports a subgroup performance gap for a model that is
exactly equally good in both. Move the panel below: the number on the right is the
gap you would measure for an identical model — the artifact floor. A reported
gap beneath it has measured the labels, not the model.
0.90
P(reader marks positive | truly positive), per reader.
0.90
The weak axis in real reads: reference false positives drive the artifact.
30%
10%
3 (2+1)
Two readers plus a third on disagreement is algebraically a 3-reader majority — the field standard.
artifact floor · apparent sensitivity gap
0.114
For a model identical in both groups. The true gap is zero.
constructed reference:Seref0.972Spref0.972
0.05
below its floor
Artifact floor vs. panel size at the current reader accuracy and prevalences.
Readers are assumed conditionally independent — the best case, since readers who
share an image share their mistakes. The dot marks your current panel.
Measured on four real radiologists
Reader accuracy above is a dial. It doesn't have to be. LIDC-IDRI had four
thoracic radiologists read 1,016 CT scans independently and releases the annotations without the
images — the one ungated source of real per-reader labels. Binary task: is a nodule ≥3 mm
present?
28.5%of scans on which the four radiologists disagree about whether a ≥3 mm nodule is present at all. Unanimity on 726 of 1,016.
0.239artifact floor at their measured accuracy, 3-reader reference, for a 30%-vs-10% prevalence contrast — worse than the parametric guess. (Press the LIDC button above.)
Latent-class reader estimates, LIDC-IDRI — identical across 12 random starts.
reader
sensitivity
specificity
1
0.960
0.801
2
0.972
0.800
3
0.936
0.739
4
0.923
0.806
mean
0.948
0.787
The conditional-independence assumption the whole construction rests on is
rejected on this same data: G² = 44.1 on 6 df, p ≈ 7e-8
(parametric bootstrap p ≤ 0.0005). LIDC's phase-2 reads were made with the other
radiologists' marks visible, so dependence is expected by design — the point is that the
canonical public multi-reader dataset violates the assumption every consensus label built from
it silently relies on.
What is known, what is new, what is open
knownThe mechanism. Biesheuvel, Irwig & Bossuyt (Clin Chem 2007): an imperfect reference interacting with differing prevalence spuriously shifts apparent sensitivity and specificity. Bossuyt authored STARD. This is not a discovery.
newQuantification & a design rule. The closed-form floor (validated by a 400,000-sample Monte Carlo), the reader-count needed to get beneath a target gap, the floor at measured radiologist accuracy, and the observation that the AI-fairness literature computing these gaps has never cited the mechanism.
openCorrelated reader error. The floor assumes independent readers. Real readers share image-driven error, which only makes the reference worse — so every number here is a lower bound. Quantifying that from real data is the honest frontier, not a solved problem.
The prevalence contrasts are illustrative: LIDC's own prevalence is 0.72, so the
0.239 figure pairs measured accuracy with hypothetical subgroup prevalences —
stated, not hidden. No published finding is claimed here to be wrong; whether a specific reported
disparity is an artifact is an empirical question needing that study's reader data.