749 tutorials · free to read

A working handbook for data and software engineering.

Hands-on tutorials by Alex Merced on Apache Iceberg, the data lakehouse, pipelines, agentic AI, and the languages and tools that hold it all together.

Browse by topic

All topics

Latest tutorials

Page 2 of 125
  1. 01The Iceberg Table Properties That Actually MatterThe Iceberg table properties that decide file count, pruning, write amplification, retention, and metadata growth, by workload.
  2. 02Geospatial Data in Apache Iceberg: Geometry, Geography, and GeoParquetHow Iceberg v3 geometry and geography types, bounding boxes, and native Parquet types give spatial data first-class standing.
  3. 03Inside the Puffin File FormatThe Puffin file format inside out, byte by byte, covering Theta sketches for distinct values and deletion vectors.
  4. 04Local Iceberg Development Environments: Docker, MinIO, and In-Memory Catalogs for CILocal Iceberg development environments: in-process catalogs, a Docker Compose stack with MinIO, and CI configurations that run either.
  5. 05The Lakehouse Ingestion Tool Landscape: Fivetran, Airbyte, dlt, and CDC vs BatchHow Fivetran, Airbyte, dlt, and CDC and streaming tools land well-behaved Apache Iceberg tables, and how to choose and maintain them.
  6. 06Metadata Platforms in 2026: DataHub, OpenMetadata, Atlan, and Catalog ConvergenceHow the technical catalog and the metadata platform are converging in 2026, and how to arrange the two layers for a lakehouse.