Skip to main content

2 posts tagged with "Streaming Lakehouse"

View All Tags

From Kafka to Fluss: How Rednote Migrated a Core Real-Time Indexing Pipeline

Han Liu
Han Liu
Head of Kafka & Fluss at Rednote

From Kafka to Fluss: Xiaohongshu's production-grade migration of its core indexing pipeline

A production case study in columnar streaming, cold-data isolation, and lakehouse integration Presented at Flink Forward Asia 2026.

Rednote (Xiaohongshu) is a lifestyle community platform centered on content discovery and sharing. During the FIFA World Cup, it becomes a hub for live coverage, pre-match analysis, trending posts, and fan discussions.

At that scale, hundreds of millions of users can hit live streaming, search, and feeds at once. Content must refresh in moments while advertising, search, and recommendation services stay stable—powered by a real-time indexing pipeline that continuously ingests data and updates indexes.

That pipeline is critical to Rednote's content distribution and user experience, but its Kafka-based wide-table architecture was reaching cost and stability limits. To scale without compromising real-time performance, Rednote introduced Fluss into its core production path and began progressively migrating index data from Kafka to Fluss.

Fluss × Iceberg (Part 1): Why Your Lakehouse Isn’t a Streamhouse Yet

Mehul Batra
PMC member of Apache Fluss
Luo Yuxia
PMC member of Apache Fluss

As software and data engineers, we've witnessed Apache Iceberg revolutionize analytical data lakes with ACID transactions, time travel, and schema evolution. Yet when we try to push Iceberg into real-time workloads such as sub-second streaming queries, high-frequency CDC updates, and primary key semantics, we hit fundamental architectural walls. This blog explores how Fluss × Iceberg integration works and delivers a true real-time lakehouse.

Apache Fluss represents a new architectural approach: the Streamhouse for real-time lakehouses. Instead of stitching together separate streaming and batch systems, the Streamhouse unifies them under a single architecture. In this model, Apache Iceberg continues to serve exactly the role it was designed for: a highly efficient, scalable cold storage layer for analytics, while Fluss fills the missing piece: a hot streaming storage layer with sub-second latency, columnar storage, and built-in primary-key semantics.

After working on Fluss–Iceberg lakehouse integration and deploying this architecture at a massive scale, including Alibaba's 3 PB production deployment processing 40 GB/s, we're ready to share the architectural lessons learned. Specifically, why existing systems fall short, how Fluss and Iceberg naturally complement each other, and what this means for finally building true real-time lakehouses.

Banner