How an AI Beat the World's Best Poker Players
Four elite pros, a hundred and twenty thousand hands, and a machine that sealed its leaks by morning.
Transcript
Somewhere in the long grind of an afternoon, the pros think they've found it. A bet size the machine handles badly. A crack they can pry open, hand after hand, for twenty days. They go to bed happy. They come back the next morning... and the crack is gone. Sealed. Rivers Casino, Pittsburgh, January twenty seventeen. Two of the four best heads-up poker players alive are sitting in a windowless back room the organizers nicknamed the Dungeon. No phones. No contact with the outside world. Across the felt from them... nobody. Just a screen, and a machine called Libratus, playing all four of them at once. So who closed the crack? Rewind twenty months. Same casino, spring of twenty fifteen. Carnegie Mellon rolls out an earlier program called Claudico against four of the world's top heads-up specialists, Doug Polk and Dong Kim among them. Eighty thousand hands. The humans finish ahead by about seven hundred thirty-three thousand in chips. And Polk's verdict is brutal and almost affectionate. Betting nineteen thousand dollars to win a seven hundred dollar pot, he said, just isn't something a person would do. Weird. Alien. Beatable. IEEE Spectrum called it a statistical draw. Tuomas Sandholm and his student Noam Brown went back to work. Twenty seventeen, the rematch. Twenty days. One hundred twenty thousand hands. Jason Les, Dong Kim, Jimmy Chou, Daniel McAulay, playing for shares of a two hundred thousand dollar purse. And behind that screen, Libratus is reaching across the city into the Pittsburgh Supercomputing Center's Bridges machine. Before the match it had built what the researchers call a blueprint — a simplified map of the game, solved by playing billions of hands against itself. No human hand histories. No poker books. No coaching. It learned the shape of the game from the rules alone. But heads-up no-limit hold'em has more possible situations than you could store in any computer on Earth. So the map goes coarse in the late streets. Here's the trick. Whenever a hand reached deeper water, Libratus put the map down and solved that one branch in real time, from scratch. And if a pro fired off a bet size the blueprint had never even considered? It built a brand-new sub-game containing that exact move, and solved THAT. Nested subgame solving, with a provable safety guarantee. Which is a polite way of saying: surprising it doesn't help. Which brings us back to the Dungeon. Module three. The researchers called it the self-improver. Every night, while the pros slept, Libratus took the moves the humans had actually made that day and treated them as a to-do list. Not to profile Jason Les. Not to learn his tells. To find the gaps in its OWN blueprint that the humans had been poking at... and fill them with freshly computed strategy. The machine wasn't studying the players. It was studying the holes the players had found in it. By breakfast, those holes were closed. And that kills the thing most of us believe about poker. Libratus never saw a face. Never heard a sigh. Its bluffs weren't psychology, they were arithmetic. If you only bet big with monsters, I fold. If you only bluff, I call. So the unexploitable answer is to do both, in precise, randomized proportion, forever. Unpredictability as a formula. The pros weren't being read. They were being balanced into a corner by something that genuinely did not care who they were. Day sixteen, Libratus crosses a million in chips. January thirtieth, last hand lands. Final margin: one million, seven hundred sixty-six thousand, two hundred fifty. A win rate of fourteen point seven big blinds per hundred hands, which in this game is a slaughter. Dong Kim did best, down only eighty-five thousand. Jason Les took the worst of it, down eight hundred eighty thousand. Sandholm afterward: "The best AI's ability to do strategic reasoning with imperfect information has now surpassed that of the best humans." Two years later, the harder version. Six players, the format most people actually play. And here's the catch — past two players, there's no guaranteed optimal solution to aim at. The math that made Libratus safe just doesn't apply. Brown and Sandholm built Pluribus anyway, and it beat elite pros including Chris Ferguson and Darren Elias across ten thousand hands. The blueprint took eight days on a sixty-four core server. No GPUs. About a hundred and forty-four dollars of cloud compute. It kept making donk bets, a move pros consider weak. Elias's take? Its strength was mixing strategies. So what got proven in that back room? Not that machines understand people. The opposite. Libratus and Pluribus are the best evidence we have that hidden-information problems — where you don't know what the other side holds, and they're actively hiding it — can be attacked with pure computation. Sandholm's team has pointed that work at negotiation, cybersecurity, military strategy, medical treatment planning, and licensed it to two startups. Careful with the leap, though. Poker has fixed rules and a scoreboard. Your boss, your landlord, a foreign government — they don't. But next time something across the table seems unreadable... consider that it might not be hiding anything at all. It might just be rolling dice you can't see.
Sources
Katy and Theo researched this episode from these sources.
- Poker Play Begins in "Brains Vs. AI: Upping the Ante" — Carnegie Mellon University
- AI Beats Poker Pros in 'Brains vs. AI' Event — PokerNews
- Libratus — Wikipedia
- Superhuman AI for heads-up no-limit poker: Libratus beats top professionals — Science
- Libratus: The Superhuman AI for No-Limit Poker (IJCAI 2017 demonstration paper)
- Team reveals inner workings of victorious AI: Libratus — TechXplore
- Poker Pros Battle Artificial Intelligence to a Statistical Draw — IEEE Spectrum
- Man Proves Greater Than Machine: Players Win $732,713 Against Bot "Claudico" — PokerNews
- Claudico — Wikipedia
- Superhuman AI for multiplayer poker — Science
- AI program beats pros in six-player poker—a first — Phys.org
- Let's Read: Superhuman AI for multiplayer poker — LessWrong