Learning by trying

Most machine learning starts with examples. Someone gathers a pile of cases where the answer is already known, the system studies them, and it learns to produce similar answers on cases it has never seen. Reinforcement learning starts somewhere else entirely. There is no pile of answers. There is a situation, a set of possible moves and a score.

The system tries something. The result is scored. A better score makes that choice a little more likely next time, a worse score makes it a little less likely, and the whole loop repeats until the choices stop drifting. That is the entire idea. Everything else is detail about how the trying is organised and how the scoring is worked out.

The familiar version happens outside computers all the time. A cook who tastes and adjusts, a child learning to stay upright on a bicycle, an apprentice told that the seam is crooked and sent back to try again: none of them was handed the answer, and all of them ended up with it.

The score is the hard part

If the idea is simple, the difficulty is easy to locate. Almost everything depends on what gets rewarded, because a system learning from a score will pursue that score faithfully, including along routes nobody intended.

That is why the interesting work sits rarely in the trying and almost always in the scoring. A few things tend to separate a score worth learning from a score that will mislead.

  • It rewards the outcome that is actually wanted rather than a convenient stand in for it.
  • It arrives often enough to be useful, because a score that comes only at the very end teaches slowly.
  • It cannot be raised by a shortcut that leaves the real task undone.
  • It survives someone else reading it and agreeing that a higher number really does mean a better result.

What it does not do

Plain language also means being plain about the limits. Learning this way takes a great many attempts, which is fine where attempts are cheap and a problem where they are not. It suits situations that can be repeated and scored, and it has little to offer where neither is true.

It carries no opinion of its own either. A system trained this way has found a way to raise a number inside the setting it was given. Whether that number deserved to be raised is a human question, asked before the training starts and worth asking again afterwards.

Held at that size, the idea becomes useful rather than mysterious. Something tries, something scores, and the trying leans towards whatever the scoring rewards. Most of the care belongs in deciding what to reward.

Frequently asked questions

What is reinforcement learning in one sentence?

Something tries an action, the result is scored, and a better score makes that choice more likely next time, repeated until the choices settle.

How is it different from learning from examples?

Learning from examples begins with a collection of cases where the answer is already known. Reinforcement learning begins with no answers at all, only a situation, some possible moves and a score.

Why does the reward matter so much?

Because the score is pursued faithfully. If it can be raised by a shortcut that leaves the real task undone, the shortcut is what gets learned.