Ask which population you are counting.
A detection rate starts with actual spam and asks how much gets flagged. The chance that a flagged message is spam starts with all flagged messages. Swapping those denominators is an easy mistake.
In this experiment, only 1% of messages are spam initially. A 90% detection rate produces about 9 correct flags per 1,000 messages. A 5% false alarm rate among the other 990 messages produces about 49.5 false flags on average. The probability that a flag indicates spam is therefore about 15.4%, not 90%.
Inside this model
P(spam | flag) = p × d / [p × d + (1 − p) × f], where p is the base rate, d is the detection rate, and f is the false alarm rate.
The dot grid samples 1,000 independent messages. Sample counts fluctuate; the model probability is calculated exactly from the controls. A sample with no flags has no observed proportion.
Why it matters
Evidence should update a prior belief, not erase it. Ask both how often a signal finds its target and how often it appears without that target.
A useful lens. Not a universal law.
- Fixed error rates and independent messages simplify a real classifier. Rates can differ across populations and change over time.
- This model illustrates reasoning about evidence; it does not measure how people actually make judgments.
Associated thinkers
Associations marked provisional are awaiting source review.
Further reading
Explore conditional probability and Bayes’ rule in Brown University’s Seeing Theory.
Seeing Theory — Bayesian inference ↗