Deep Blue never learned anything: the real difference between AI and machine learning

In May 1997 a computer beat the world chess champion. In March 2016 a computer beat one of the world’s best Go players. Both made headlines as triumphs of artificial intelligence, and both were. But the two machines had almost nothing in common. Deep Blue did not learn to play chess. AlphaGo did almost nothing but learn. The difference between those two machines is the difference between artificial intelligence and machine learning, and once you see it you will notice it everywhere.

Two words with two histories

Artificial intelligence was named in 1955, in a proposal for a summer workshop at Dartmouth College written by John McCarthy and three colleagues. Their ambition was the whole thing: machines that use language, form concepts, solve problems reserved for humans, and improve themselves. AI is the goal. It says nothing about the method.

Machine learning was named four years later by Arthur Samuel, an IBM engineer who had written a checkers program that got better by playing against itself. He defined it as giving computers the ability to learn without being explicitly programmed. Machine learning is one method. It is a way of getting to AI, or to a lot of things that nobody would call AI, by letting a program adjust itself from examples instead of following rules written by a person.

So the relationship is simple. Machine learning is a subset of artificial intelligence. Everything that learns from data to do something intelligent-looking is AI. But a great deal of AI, historically most of it, never learned anything.

flowchart TB

AI[Artificial intelligence: the goal] --> R[Rule-based systems]

AI --> ML[Machine learning: learn from examples]

R --> R1[Deep Blue, 1997]

R --> R2[Expert systems, 1970s to 80s]

R --> R3[Route planning, chess openings, tax software]

ML --> S[Supervised: labelled examples]

ML --> U[Unsupervised: find structure]

ML --> RL[Reinforcement: reward and punishment]

S --> DL[Deep learning: learned features]

DL --> D1[AlphaGo, 2016]

DL --> D2[Speech, vision, language models]
AI is the ambition. Machine learning is one family of methods for reaching it, and deep learning is one branch of that family.

The rules era

For its first thirty years, AI mostly meant writing rules. The Logic Theorist of 1956 proved theorems by searching through logical steps. Chess programs searched millions of positions per second and scored them with evaluation functions that grandmasters helped design. Expert systems of the 1970s and 80s, such as MYCIN for diagnosing blood infections and XCON for configuring computer orders at Digital Equipment Corporation, encoded the judgement of human specialists as thousands of if-then rules. XCON reportedly saved the company tens of millions of dollars a year.

This was genuinely intelligent behaviour by any reasonable standard of the time, and none of it was learned. Deep Blue is the peak of this era. It won because it could search two hundred million positions a second and because its evaluation function had been tuned by hand over years. Show it a new game and it would have been helpless.

The rules era ran into a wall that its own practitioners named: the knowledge acquisition bottleneck. Every rule had to be extracted from a human expert and written down. Experts disagreed, could not articulate what they knew, and the rules interacted in ways nobody could predict. The systems were brittle, expensive to build, and impossible to maintain. By the late 1980s funding collapsed in what is now called the second AI winter.

The learning era

The alternative had been there since the beginning. Frank Rosenblatt’s perceptron of 1958 learned to classify simple images by adjusting weights from examples, and Samuel’s checkers program learned from self-play. But learning needed two things the rules approach did not: large amounts of data and large amounts of computation. Neither was available.

Three things changed. The backpropagation algorithm, popularised in 1986, made it practical to train networks with many layers. The internet produced data at a scale nobody had imagined. And graphics cards built for video games turned out to be ideal for the arithmetic that neural networks need. The moment usually cited as the turning point is 2012, when a deep network called AlexNet cut the error rate on the ImageNet image recognition benchmark by a margin that made the entire field switch approaches within about two years.

AlphaGo is the peak of this era so far. Go has too many positions to search the way Deep Blue searched chess, and no human could write an evaluation function for it. So AlphaGo learned one, first from thirty million positions from expert games, then from millions of games against itself. Its successor, AlphaZero, dropped the human games entirely and learned chess, Go and shogi from nothing but the rules, and beat the best rule-based chess engine in the world in a matter of hours.

Spot the difference in daily life

The distinction is not academic. It tells you how a system will fail, what it needs to work, and who is responsible when it is wrong.

TaskRule-based approachLearned approach
Filtering spamBlock messages containing certain words. Beaten within weeks by misspellings.Since about 2002, classifiers trained on millions of labelled messages. Adapts as spam changes.
NavigationShortest-path search, as in every GPS. No learning needed; the road graph is known.Arrival-time prediction, which is learned from traffic history.
Filing taxesThe tax code is rules. Tax software encodes them exactly. A learned system here would be a liability.Fraud detection on the returns, learned from past cases.
Reading handwritingAttempts to describe letter shapes as rules failed for decades.Networks trained on handwritten digits have read postal codes and bank cheques since the early 1990s.
Detecting spikes in a neural recordingA voltage threshold, chosen by the experimenter. Still the first step in most pipelines.Sorting the detected spikes into neurons, which is clustering, a form of unsupervised learning.

Notice that the rule-based column is not the losing side. Your GPS, your tax software and your compiler are all rules, and they should be. The learned column wins only where the rules cannot be written down, which is the subject of the next article.

Where neuroscience sits in this

The brain was the inspiration for the learning side from the start. Rosenblatt was a psychologist, and the perceptron was explicitly a model of how neurons might learn. The modern neural network keeps the name and the rough idea, a network of simple units whose connections strengthen with experience, and discards nearly everything else about biology. That gap is the subject of the Neuroscience meets AI track on this site: what the brain does that our learners do not, and which of those differences matter.

It also cuts the other way. Neuroscience today uses both columns of the table above. A threshold detector is a rule. A spike sorter is a learner. A decoder that turns motor cortex activity into cursor movement is a learner, but the safety limits around it are rules. Knowing which is which is part of being able to build and trust these systems.

Further reading: Nils Nilsson, The Quest for Artificial Intelligence, is the standard history and is free online. Arthur Samuel’s 1959 paper on checkers is short and still readable.