All selected work
Cloud modernizationCASE STUDY / 01

Enterprise data platform modernization

Moving legacy Airflow and EMR workflows into Databricks-native orchestration, governed storage and repeatable PySpark processing.

Zillow Group · Jan 2026 — PresentENGINEERING CASE STUDY
DatabricksPySparkAirflowAWSUnity CatalogDelta Lake
01

Overview

Platform modernization brings orchestration, processing and storage into a Databricks-native workflow. The work includes migrating Airflow DAGs and AWS EMR processing, refactoring PySpark and adapting storage access to Unity Catalog Volumes.

02

Problem

Legacy workflows couple orchestration dependencies, platform-specific processing and S3 storage assumptions. A migration must preserve those dependencies and the resulting data while reducing operational complexity.

03

My role

Workflow analysis, dependency mapping, DAG migration, PySpark refactoring, configuration management, source-to-target validation and Spark performance optimization form the core engineering work.

04

Architecture

Map legacy Airflow dependencies to Databricks Workflows. Run refactored PySpark tasks against governed Unity Catalog storage and process Delta and Parquet outputs. The diagram shows the conceptual migration path rather than a proprietary production topology.

LEGACY
Airflow
AWS EMR
PySpark
S3
MODERNIZED
Databricks Workflows
PySpark
Unity Catalog
Delta Lake
05

Technical decisions

Use configuration-driven execution and reusable utilities. Separate orchestration from transformation logic, manage secrets securely and validate each migrated workflow before treating it as equivalent to the legacy path.

06

Challenges

Dependency mapping, runtime differences and legacy storage assumptions require careful analysis. Data parity and operational reliability remain the criteria for evaluating a migration—not simply whether a job completes.

07

Data validation

Compare source and target schemas, record counts, partitions, null values, duplicates and data types. Reconciliation makes migration testing repeatable and exposes differences that successful task execution alone cannot detect.

SOURCE → VALIDATE → TARGET
  • Schema
  • Row count
  • Partitions
  • Null values
  • Duplicates
  • Data types
  • Reconciliation
08

Performance

Inspect Spark execution plans, partitioning, cache use and shuffle behavior. Tune processing around the observed workload rather than assuming that moving platforms automatically improves execution.

09

Outcome

The modernization is ongoing. Its intended outcomes are maintainable orchestration, governed data access, reusable processing utilities and repeatable migration validation.

10

Lessons learned

A useful migration principle is to preserve the data contract while changing the execution platform. Treat dependency mapping, configuration and validation as engineering deliverables alongside the migrated code.