Tag: streaming
All the articles with the tag "streaming".
- 21 MIN READ•Aug 4, 2026
How Iceberg V3 Deletion Vectors Fixed Merge-on-Read for Streaming Tables
How Iceberg V3 deletion vectors replaced accumulating positional delete files and made merge-on-read viable for streaming and CDC tables.
Apache IcebergIceberg V3Deletion Vectors - 21 MIN READ•Aug 4, 2026
Reading the Apache Iceberg V4 Proposals Before They Land
A field guide to the Apache Iceberg V4 proposals: adaptive metadata trees, single-file commits, typed statistics, column families, and what is safe to build on today.
Apache IcebergIceberg V4Metadata - 31 MIN READ•Jul 28, 2026
Why Iceberg V4 Wants to Retire Equality Deletes, and What Streaming Teams Should Do About It
Equality deletes made streaming upserts into Iceberg practical at the cost of read performance. V4 proposes retiring them in favor of deletion vectors with an async conversion path.
Apache IcebergStreamingData Engineering - 31 MIN READ•Jul 28, 2026
Apache Fluss and Kafka Solve Different Problems in an Iceberg Pipeline
Fluss puts a columnar, indexed hot tier between Kafka and Iceberg. Here's what it changes structurally, what Kafka still does better, and how to benchmark the comparison yourself.
Apache IcebergApache FlussKafka - 31 MIN READ•Jul 28, 2026
Serving Sub-Second Queries Over an Iceberg Lakehouse With a Hot Tier
A lakehouse cannot serve sub-second queries over seconds-old data. A hot tier in front solves it, with consequences for consistency, governance, and operational surface.
Apache IcebergStreamingData Serving - 31 MIN READ•Jul 25, 2026
Freshness Is a Contract, Not a Note on a Dashboard
Data freshness needs to become an engineering contract with a measurable value, an owner, and consequences. How to decompose lag, make freshness queryable, and keep agents honest.
data freshnessdata qualityapache iceberg - 30 MIN READ•Jul 6, 2026
The State of Streaming to Apache Iceberg in July 2026: Every Path, Its Latency, and What to Do When Seconds Are Not Fast Enough
Every path for streaming data into Iceberg in 2026 — Flink, Spark, Kafka Connect, broker-native, managed pipelines — with honest latency numbers and sub-second hybrid architectures.
Apache Icebergstreamingdata engineering - 5 MIN READ•Feb 18, 2026
Batch vs. Streaming: Choose the Right Processing Model
"We need real-time data." This is one of the most expensive sentences in data engineering – because it's rarely true, and implementing it when it's not neede...
data engineeringbest practicesbatch processing - 3 MIN READ•Jul 29, 2025
Optimizing Compaction for Streaming Workloads in Apache Iceberg
Learn how to design fast, incremental compaction strategies in Apache Iceberg to support high-throughput streaming pipelines without disrupting freshness or performance.
Apache IcebergData OptimizationStreaming - 6 MIN READ•Oct 16, 2024
Data Lakehouse Roundup 1 - News and Insights on the Lakehouse
What's Going on in the Data Lakehouse Space
data lakehousedata engineeringstreaming - 14 MIN READ•Oct 4, 2024
Change Data Capture (CDC) when there is no CDC
Handling Synching Changing Data Across Systems
data lakehousedata engineeringstreaming