How Palantir Foundry Handles Streaming Data
Append-only logs, tumbling windows and live Ontology objects — what Palantir's own developer docs say about moving data in under a second.
Transcript
A truck is idling outside a distribution centre. A sensor in the trailer fires a reading. Temperature, two degrees too high. And the question Palantir's engineers had to answer is brutally simple. How many seconds pass before a human being somewhere sees that number... and can still do something about it?
And the honest answer, for most enterprise software, is what — minutes? Hours?
Historically, hours. Because the standard way data moves inside a big company is batch. You collect everything, run a job on a schedule, write out a new version of the table. But Palantir's entire pitch is real-time operational decisions. And you cannot make a real-time decision on a table that refreshes at two in the morning.
So this episode is... the plumbing.
This episode is the plumbing. And one caveat before we start. Everything I'm describing comes from Palantir's own public developer documentation for Foundry. It's what they SAY the system does. It isn't independent benchmarking, and I'm not going to pretend it is.
Noted. Okay. Smallest possible unit. What is a stream in Foundry?
An append-only log of records. Picture a notebook where you can only ever write the next line. You never go back and erase line four hundred. Records arrive, they get appended, and everything downstream reads forward.
Versus a normal dataset, which is...
Versioned by transactions. You commit a chunk, that's a new version, downstream jobs wake up and chew through the whole thing. A batch dataset is a photograph. A stream is the video feed. Genuinely different shapes, which is why Foundry treats them as separate dataset types instead of pretending one is the other.
Hang on. Append-only, never stops... doesn't it just grow forever until somebody's storage bill explodes?
Good instinct, and no. Streams have a retention window. You configure how long records stay, and old ones age out. The stream is a moving window on recent reality. Not an archive.
Then where does the history go? Because somebody always wants the six-month chart.
Foundry can archive the stream into batch-readable storage alongside it. So one source feeds the live alert AND the six-month trend. That's the quietly clever bit — most companies run two entirely separate stacks for those two jobs.
What happens when someone adds a field to the sensor and everything downstream catches fire?
Streams are schema'd up front. Declared structure, validated on arrival. Which sounds bureaucratic right up until you're the person at three in the morning asking why the temperature column is suddenly a string.
Been that person. Right — so how do you actually transform the thing while it's moving?
Pipeline Builder. Foundry's visual pipeline tool, with a streaming mode. Same interface you'd use for batch, except the logic runs continuously. Record by record, as data lands.
So nothing waits for a batch to fill up.
That's the core difference. In batch, your latency is capped by your schedule — hourly job, worst case an hour stale. In streaming, the transform fires on the record. The documented target is sub-second class, rather than minutes.
But some things you just can't do one record at a time. An average needs more than one number.
And that's the beautiful constraint. Average trailer temperature — average over WHAT? The stream never ends. So streaming answers with windows. You slice infinity into finite buckets. A tumbling window: every five minutes, back to back, no overlap. Or a sliding one that keeps moving and overlapping.
So the aggregation isn't over a table. It's over a slice of time.
Over a slice of time, grouped by a key. Per vehicle. Per sensor. Per warehouse. The system holds state for each key while the window is open, then emits a result when it closes.
And the truck goes through a tunnel.
And the truck goes through a tunnel. Loses signal, comes out the other side, dumps ten minutes of readings at once — into a window that already closed. Every streaming system on earth wrestles with that one, and Foundry is no exception. There are documented ways to handle late data, but there's no magic. It's a trade-off between waiting longer and being right.
Okay. Fast numbers, windowed, grouped. What do I actually DO with them?
Here's the part that's specifically Palantir. The stream syncs into the Ontology.
Which, for anyone who skipped the earlier episodes...
Is Foundry modelling your business as objects instead of tables. Not row four hundred in a database — a TRUCK. A shipment. A turbine. With properties, and links to other objects. When a stream feeds that, the properties update live. The truck's temperature isn't last night's figure. It's now.
So the dashboard, the alert and the analyst's model are all staring at the same live object.
That's the architectural bet. One representation, many consumers. And you hang automations on top — rules that watch for a condition and fire an action. Temperature over threshold, flag the shipment, notify the depot.
Who's actually running this?
The documented patterns are industrial. Sensor telemetry off machinery. Fleet and supply chain tracking. Manufacturing lines, energy, aviation — places where the whole value is noticing something in the next thirty seconds instead of the next quarter.
I notice nobody's claiming it makes the decision for you.
No. And that's the criticism worth sitting with. Faster data is not better judgement. Critics have argued for years that slick real-time interfaces manufacture a feeling of certainty the underlying data hasn't earned. A number arriving in four hundred milliseconds is still only as good as the sensor that sent it.
And the bull case is that the speed IS the moat.
That's the company's case. Anyone can build a dashboard. Very few can wire live telemetry into a shared model of the business that an operator, an analyst and an automated rule all act on in the same moment. Whether that's worth what customers pay is a completely separate argument.
So — the truck. How many seconds?
That's the number the whole architecture exists to shrink. The gap between something happening in the physical world and somebody being able to act on it. The append-only log, the retention window, the tumbling window, the Ontology sync — all of it is engineering in service of that one gap.
The distance between the world and the screen.
Nobody writes headlines about retention policies. But when a company tells you it runs the world in real time... that claim is only ever as true as the pipe underneath it.
Sources
Katy and Theo researched this episode from these sources.