KT for iPhone is in beta. Tap to install with TestFlight →

How Palantir Foundry Handles Streaming Data

Katy & TheoEpisode 4 of Palantir8 min

Append-only logs, tumbling windows and live Ontology objects — what Palantir's own developer docs say about moving data in under a second.

Transcript

Katy

A truck is idling outside a distribution centre. A sensor in the trailer fires a reading. Temperature, two degrees too high. And the question Palantir's engineers had to answer is brutally simple. How many seconds pass before a human being somewhere sees that number... and can still do something about it?

Theo

And the honest answer, for most enterprise software, is what — minutes? Hours?

Katy

Historically, hours. Because the standard way data moves inside a big company is batch. You collect everything, run a job on a schedule, write out a new version of the table. But Palantir's entire pitch is real-time operational decisions. And you cannot make a real-time decision on a table that refreshes at two in the morning.

Theo

So this episode is... the plumbing.

Katy

This episode is the plumbing. And one caveat before we start. Everything I'm describing comes from Palantir's own public developer documentation for Foundry. It's what they SAY the system does. It isn't independent benchmarking, and I'm not going to pretend it is.

Theo

Noted. Okay. Smallest possible unit. What is a stream in Foundry?

Katy

An append-only log of records. Picture a notebook where you can only ever write the next line. You never go back and erase line four hundred. Records arrive, they get appended, and everything downstream reads forward.

Theo

Versus a normal dataset, which is...

Katy

Versioned by transactions. You commit a chunk, that's a new version, downstream jobs wake up and chew through the whole thing. A batch dataset is a photograph. A stream is the video feed. Genuinely different shapes, which is why Foundry treats them as separate dataset types instead of pretending one is the other.

Theo

Hang on. Append-only, never stops... doesn't it just grow forever until somebody's storage bill explodes?

Katy

Good instinct, and no. Streams have a retention window. You configure how long records stay, and old ones age out. The stream is a moving window on recent reality. Not an archive.

Theo

Then where does the history go? Because somebody always wants the six-month chart.

Katy

Foundry can archive the stream into batch-readable storage alongside it. So one source feeds the live alert AND the six-month trend. That's the quietly clever bit — most companies run two entirely separate stacks for those two jobs.

Theo

What happens when someone adds a field to the sensor and everything downstream catches fire?

Katy

Streams are schema'd up front. Declared structure, validated on arrival. Which sounds bureaucratic right up until you're the person at three in the morning asking why the temperature column is suddenly a string.

Theo

Been that person. Right — so how do you actually transform the thing while it's moving?

Katy

Pipeline Builder. Foundry's visual pipeline tool, with a streaming mode. Same interface you'd use for batch, except the logic runs continuously. Record by record, as data lands.

Theo

So nothing waits for a batch to fill up.

Katy

That's the core difference. In batch, your latency is capped by your schedule — hourly job, worst case an hour stale. In streaming, the transform fires on the record. The documented target is sub-second class, rather than minutes.

Theo

But some things you just can't do one record at a time. An average needs more than one number.

Katy

And that's the beautiful constraint. Average trailer temperature — average over WHAT? The stream never ends. So streaming answers with windows. You slice infinity into finite buckets. A tumbling window: every five minutes, back to back, no overlap. Or a sliding one that keeps moving and overlapping.

Theo

So the aggregation isn't over a table. It's over a slice of time.

Katy

Over a slice of time, grouped by a key. Per vehicle. Per sensor. Per warehouse. The system holds state for each key while the window is open, then emits a result when it closes.

Theo

And the truck goes through a tunnel.

Katy

And the truck goes through a tunnel. Loses signal, comes out the other side, dumps ten minutes of readings at once — into a window that already closed. Every streaming system on earth wrestles with that one, and Foundry is no exception. There are documented ways to handle late data, but there's no magic. It's a trade-off between waiting longer and being right.

Theo

Okay. Fast numbers, windowed, grouped. What do I actually DO with them?

Katy

Here's the part that's specifically Palantir. The stream syncs into the Ontology.

Theo

Which, for anyone who skipped the earlier episodes...

Katy

Is Foundry modelling your business as objects instead of tables. Not row four hundred in a database — a TRUCK. A shipment. A turbine. With properties, and links to other objects. When a stream feeds that, the properties update live. The truck's temperature isn't last night's figure. It's now.

Theo

So the dashboard, the alert and the analyst's model are all staring at the same live object.

Katy

That's the architectural bet. One representation, many consumers. And you hang automations on top — rules that watch for a condition and fire an action. Temperature over threshold, flag the shipment, notify the depot.

Theo

Who's actually running this?

Katy

The documented patterns are industrial. Sensor telemetry off machinery. Fleet and supply chain tracking. Manufacturing lines, energy, aviation — places where the whole value is noticing something in the next thirty seconds instead of the next quarter.

Theo

I notice nobody's claiming it makes the decision for you.

Katy

No. And that's the criticism worth sitting with. Faster data is not better judgement. Critics have argued for years that slick real-time interfaces manufacture a feeling of certainty the underlying data hasn't earned. A number arriving in four hundred milliseconds is still only as good as the sensor that sent it.

Theo

And the bull case is that the speed IS the moat.

Katy

That's the company's case. Anyone can build a dashboard. Very few can wire live telemetry into a shared model of the business that an operator, an analyst and an automated rule all act on in the same moment. Whether that's worth what customers pay is a completely separate argument.

Theo

So — the truck. How many seconds?

Katy

That's the number the whole architecture exists to shrink. The gap between something happening in the physical world and somebody being able to act on it. The append-only log, the retention window, the tumbling window, the Ontology sync — all of it is engineering in service of that one gap.

Theo

The distance between the world and the screen.

Katy

Nobody writes headlines about retention policies. But when a company tells you it runs the world in real time... that claim is only ever as true as the pipe underneath it.

Sources

Katy and Theo researched this episode from these sources.

  1. Palantir Foundry documentation — Pipeline Builder
  2. Create Streaming Dataset • API Reference • Palantir
  3. foundry-platform-python/docs/v2/Streams/Dataset.md at develop · palantir/foundry-platform-python