Overview
Platform modernization brings orchestration, processing and storage into a Databricks-native workflow. The work includes migrating Airflow DAGs and AWS EMR processing, refactoring PySpark and adapting storage access to Unity Catalog Volumes.
Problem
Legacy workflows couple orchestration dependencies, platform-specific processing and S3 storage assumptions. A migration must preserve those dependencies and the resulting data while reducing operational complexity.
My role
Workflow analysis, dependency mapping, DAG migration, PySpark refactoring, configuration management, source-to-target validation and Spark performance optimization form the core engineering work.
Architecture
Map legacy Airflow dependencies to Databricks Workflows. Run refactored PySpark tasks against governed Unity Catalog storage and process Delta and Parquet outputs. The diagram shows the conceptual migration path rather than a proprietary production topology.
Technical decisions
Use configuration-driven execution and reusable utilities. Separate orchestration from transformation logic, manage secrets securely and validate each migrated workflow before treating it as equivalent to the legacy path.
Challenges
Dependency mapping, runtime differences and legacy storage assumptions require careful analysis. Data parity and operational reliability remain the criteria for evaluating a migration—not simply whether a job completes.
Data validation
Compare source and target schemas, record counts, partitions, null values, duplicates and data types. Reconciliation makes migration testing repeatable and exposes differences that successful task execution alone cannot detect.
- Schema
- Row count
- Partitions
- Null values
- Duplicates
- Data types
- Reconciliation
Performance
Inspect Spark execution plans, partitioning, cache use and shuffle behavior. Tune processing around the observed workload rather than assuming that moving platforms automatically improves execution.
Outcome
The modernization is ongoing. Its intended outcomes are maintainable orchestration, governed data access, reusable processing utilities and repeatable migration validation.
Lessons learned
A useful migration principle is to preserve the data contract while changing the execution platform. Treat dependency mapping, configuration and validation as engineering deliverables alongside the migrated code.
