Video explainers
The data lakehouse, explained in under a minute each
Nine short animated explainers covering the ideas that come up first when teams move analytics onto a lakehouse — from why OLTP and OLAP split apart, through Parquet, table formats and catalogs, to semantic layers and agentic AI. There is no narration and no audio track at all: every point is on screen, so they read exactly the same muted, in a meeting, or embedded in your own docs.
-
01 OLTP vs OLAP
Why transactional systems store data by row, analytical systems store it by column, and the two workloads end up on separate infrastructure.
- OLTP runs the business, a few rows at a time
- OLAP answers questions about the business
- The shapes conflict, so the systems split
-
02 The Data Warehouse
The warehouse gave analytics governed, modeled, ACID-backed tables. Its limits are format lock-in, structured data only, and a cost curve tied to one vendor's engine.
- One governed place to ask questions
- Structured data only, in someone else's format
- Storage, compute and format bound together
-
03 The Evolution to the Data Lakehouse
Warehouse, then data lake, then lakehouse: the path from a closed governed store to open files on object storage that behave like real tables.
- The warehouse: reliable, closed, expensive
- The data lake: open and cheap, but not a table
- A table format turns files into a table
- Warehouse behaviour on storage you own
-
04 Data Lakes and Apache Parquet
Object storage is the cheap open substrate; Apache Parquet is the columnar file format — column chunks, encodings and footer statistics — that makes scanning it fast.
- Object storage is cheap, durable and open
- Parquet stores data by column, not by row
- Footer statistics let engines skip data
-
05 Open Table Formats
A table format is the metadata layer that turns a directory of Parquet into a real table — and the three open ones each solve it a little differently.
- A directory of Parquet is not a table
- Metadata names the exact set of files
- Database guarantees on open files
- Same problem, three designs
-
06 Catalogs and Apache Polaris
A catalog holds the current pointer for every table and makes commits atomic. Apache Polaris is an open-source, vendor-neutral implementation of the Iceberg REST Catalog spec.
- A catalog is the table's source of truth
- The Iceberg REST Catalog is an open API
- An open-source Iceberg REST catalog
-
07 The Lakehouse and Agentic AI
Agents query data without an analyst in the loop, which moves access control, definitions and auditability down to the table and the catalog.
- Agents query data on their own
- Permissions have to hold at the table
- One copy, many engines, one meaning
- Agents need context, not just access
-
08 Semantic Layers
A semantic layer holds the endorsed metrics, dimensions and joins every consumer resolves against. Apache Ossie (incubating) is the open standard for moving those definitions between tools.
- Every tool redefines the same metric
- One governed place for meaning
- Definitions written in one vendor's dialect
- An open standard for semantic interchange
-
09 What Is Apache Arrow?
Arrow is the in-memory counterpart to Parquet: one agreed columnar layout so systems can hand data to each other with no serialization step at all.
- Every hop used to re-encode the data
- One columnar layout, agreed in advance
- Zero-copy interchange, vectorized execution
Silent by design — no narration, no music. Hover to preview, click to play, or download any MP4 and reuse it.