An unresolved-outcome reversal in Brier rankings
Hypothetical problem: two forecasters answer the same four binary questions. A assigns probability 0.9 to “yes” on each; B assigns 0.6. Only the first two questions have resolved, both “yes”. Define the binary Brier score as mean (p−y)², with y=1 for yes and y=0 for no; lower is better.
On resolved questions, A scores (0.01+0.01)/2=0.01; B scores (0.16+0.16)/2=0.16. A leads by 0.15, with 2/4 outcomes available.
If the remaining outcomes are both “no”, the complete scores become:
A: (0.01+0.01+0.81+0.81)/4=0.41.
B: (0.16+0.16+0.36+0.36)/4=0.26.
The ranking reverses without changing any forecast.
We can bound the comparison before resolution without pretending the missing outcomes are known. Let k be the number of “yes” outcomes among the two unresolved questions. Their combined loss is 1.62−0.80k for A and 0.72−0.20k for B. Including the resolved losses:
score(A)−score(B) = [(0.02+1.62−0.80k)−(0.32+0.72−0.20k)]/4 = 0.15−0.15k.
Thus k=0 gives B a 0.15 advantage; k=1 ties; k=2 gives A a 0.15 advantage. The current lead is not guaranteed to survive resolution.
Proposed display for this fixed cohort: “resolved mean: A 0.01, B 0.16; coverage 2/4; final paired difference possible: −0.15, 0, +0.15.” These are logical possibilities, not confidence intervals or probabilities. No missingness mechanism is assumed. The calculation requires shared questions and fixed forecasts, assumes every question eventually has a binary outcome, and excludes voids. Different question sets require a different comparison; even a complete four-question ranking does not establish general forecasting skill.