Which machine would you play next?
A bandit problem asks how to balance immediate reward and learning from experiments.
Inside this model
Three seeded machines have fixed hidden win chances. The comparison agent uses epsilon-greedy choice over sample means, beginning with one trial of each. The oracle always chooses the best hidden chance. The plot shows cumulative reward over 60 rounds.
Out in the world
A practical use
Trying a new page design while still showing the strongest current design.
A useful lens. Not a universal law.
- Independent stationary rewards and a single epsilon rule omit changing audiences, costs and delayed outcomes.
Associated thinkers
Further reading
Explore the original research or the teaching reference behind this experiment.
William R. Thompson — On the Likelihood that One Unknown Probability Exceeds Another ↗