All selected work
Streaming data engineeringCASE STUDY / 03

Real-time streaming data pipeline

Event ingestion and continuous processing with Kafka, PySpark and Flink, with fault tolerance and observability in the design.

Allstate · Jul 2024 — Dec 2025ENGINEERING CASE STUDY
KafkaPySparkFlinkAWSStreaming
01

Overview

Streaming engineering connects event producers to processing and analytical delivery. The work uses Kafka and Flink alongside Spark-based data processing.

02

Problem

Continuous workloads require deliberate handling of transformation, delivery and failures. Producing a stream is only one part of a dependable data system.

03

My role

Work on event ingestion, streaming transformations and integration with AWS data platforms, connecting processed events to downstream data consumers.

04

Architecture

The conceptual diagram places Kafka between event producers and PySpark / Flink processing. Transformation and enrichment prepare events for storage and analytical consumption.

  1. 01Event producers
  2. 02Kafka
  3. 03PySpark / Flink
  4. 04Enrichment
  5. 05Data platform
  6. 06Analytics
05

Technical decisions

Keep ingestion, processing and delivery boundaries explicit. Consider fault tolerance and observability with the transformation logic.

06

Challenges

Streaming requires attention to event continuity, processing failures and downstream delivery.

07

Data validation

Schema and completeness checks are useful at stream boundaries. Validation should make malformed or incompatible events visible.

SOURCE → VALIDATE → TARGET
  • Schema
  • Row count
  • Partitions
  • Null values
  • Duplicates
  • Data types
  • Reconciliation
08

Performance

Evaluate processing behavior and resource use against the actual event workload.

09

Outcome

Streaming work links event ingestion, continuous transformation and analytical delivery.

10

Lessons learned

Treat continuous processing as a full lifecycle: ingestion, transformation, recovery, delivery and monitoring all matter.