Which option really performs better?
An aggregate comparison can tell a different story from every subgroup comparison when the groups have different weights. Edward Simpson analyzed how combining contingency tables changes interpretation.
Inside this model
In each difficulty group A succeeds 10 percentage points more often: easy 90% versus 80%; hard 30% versus 20%. Overall A = 0.30+0.60×easyA and B = 0.20+0.60×easyB, with shares expressed as fractions. These are illustrative expected rates for equal-sized programs.
Out in the world
A practical use
Before judging two hospitals or schools by one headline rate, compare the difficulty of the cases each accepts.
A useful lens. Not a universal law.
- This toy comparison is not a causal estimate. Whether to pool or split real data depends on how groups arise, what was measured, and the question being asked.
Associated thinkers
Further reading
Explore the original research or the teaching reference behind this experiment.
Edward Simpson — The Interpretation of Interaction in Contingency Tables ↗