KT for iPhone is in beta. Tap to install with TestFlight →

AlphaGo Beats Lee Sedol With Move 37

Katy & TheoEpisode 1 of The Machines Are Learning6 min

A machine broke four hundred years of Go doctrine in Seoul. Two games later, one man broke the machine.

Transcript

Katy

Seoul. March, twenty sixteen. Lee Sedol, eighteen world titles, the best player of his generation, is thirty-six moves into game two when a man reaches across the board and sets down a black stone in a place no professional would ever put one. The man is just the hands. The move came from a computer. And by that computer's own reckoning, it's a stone a human plays about one time in ten thousand. In the commentary booth, Fan Hui, the European champion who had trained against this thing, stares at the board. It's not a human move. Two hundred million people are watching. Remember that number. One in ten thousand. Because in two games' time, a human is going to fire it straight back. So what was so heretical about move thirty-seven? In Go there's a play called a shoulder hit. You lean your stone diagonally against your opponent's. And every child who learns the game learns the same commandment: shoulder hits go on the third line or the fourth. Third for territory, fourth for influence. Centuries of doctrine, taught the same way in every Go school on earth. AlphaGo played it on the FIFTH line. To a professional, that looks roughly like a grandmaster walking his king out into the middle of the chessboard for the fun of it. David Silver of DeepMind says players simply assumed the human operator had made a mistake. Now, that one-in-ten-thousand figure. It wasn't a poll of grandmasters. It came from inside the machine. AlphaGo's policy network had been trained on around thirty million positions from expert human games, and its whole job was to guess what a person would play next. Then the program played millions of games against copies of itself. A second network estimated who was winning. A search explored what came after. So move thirty-seven is really the machine saying two things in the same breath. My model of humanity says nobody plays this. My search says play it anyway. It worked. The stone grew into a beautiful shape in the centre, and the game was gone. Afterwards, Lee said something remarkable for a man who'd just been beaten in front of the planet. He'd assumed AlphaGo was only probability calculation. Merely a machine. Then he saw that move, and he changed his mind. Surely, he said, AlphaGo is creative. Then AlphaGo won game three. Three-nil. The match was over as a contest, and Lee walked into game four playing for one thing only. Not the trophy. One win. For the species. Four Seasons Hotel, Seoul. Game four opens with the exact same first eleven moves as game two. Lee tries amashi, letting the machine take the grand, influential points while he quietly pockets profit underneath. It doesn't work. AlphaGo's win probability climbs toward seventy percent. And then Lee stops. He thinks for a long, long time. And he plays move seventy-eight. A wedge. A white stone dropped into the tiny gap between two black stones, right in the belly of the board. By the machine's own maths... a move a human plays about one time in ten thousand. And AlphaGo breaks. Not poetically. Actually. Its search survives by pruning, throwing away the lines it judges irrelevant so it can look deeper down the ones that matter. Lee's wedge was so unlikely it had essentially never been considered. There was no plan waiting behind it. Move seventy-nine is a misplay. That seventy percent falls through fifty. The machine starts putting down stranger and stranger stones, and after move one hundred and eighty, it resigns. Back at DeepMind they went hunting for the bug. Engineer Ioannis Antonoglou's verdict: the bug was that Lee Sedol came up with an ingenious move. To this day, that is the only recorded game AlphaGo ever lost to a human being. Three years later, Lee Sedol retired from professional Go. His reasoning was brutally clear. Even if he became the number one player alive, there would still be, in his words, an entity that cannot be defeated. Think about what that costs a competitor. The entire logic of his life was that effort closes gaps. Move thirty-seven told him the gap had a floor. Move seventy-eight told him he could still reach through it. Once. That's the part that outlived the match. Researchers stopped asking whether machines can be creative and started asking the harder question: creative according to whom? Move thirty-seven only stunned us because it sat outside everything humans had already tried. It wasn't in the corpus. It came from a machine playing itself in the dark, with no teacher. And within months, professionals were testing those fifth-line plays themselves. Human Go got better because a machine ignored us. Every AI you touch now is running some version of that same trick: searching a space we never walked. Sometimes it comes back with nonsense. Sometimes it comes back with move thirty-seven. The whole game, still, is knowing which stone you're holding.

Sources

Katy and Theo researched this episode from these sources.

  1. AlphaGo — Google DeepMind
  2. Where does AlphaGo go now? — IT Brew
  3. AI Finds A Way (David Silver and Marc Lanctot firsthand accounts)
  4. Language Has Two Parameters (Move 37 case study)
  5. Lee Sedol and AlphaGo: The Legacy of a Historic Fight
  6. The God Move, AlphaGo, and the Rise of Agentic AI