THE IDEA

Which machine would you play next?

A bandit problem asks how to balance immediate reward and learning from experiments.

Inside this model

Three seeded machines have fixed hidden win chances. The comparison agent uses epsilon-greedy choice over sample means, beginning with one trial of each. The oracle always chooses the best hidden chance. The plot shows cumulative reward over 60 rounds.

Out in the world

A practical use

Trying a new page design while still showing the strongest current design.

WHERE IT BREAKS

A useful lens. Not a universal law.

  • Independent stationary rewards and a single epsilon rule omit changing audiences, costs and delayed outcomes.

Associated thinkers

Further reading

Explore the original research or the teaching reference behind this experiment.

William R. Thompson — On the Likelihood that One Unknown Probability Exceeds Another ↗