Overview
Data quality automation supports Databricks platform modernization through source-to-target validation of Delta and Parquet processing.
Problem
A migrated workflow can execute successfully while producing a different schema, record count or partition layout. Job success cannot stand in for data correctness.
My role
Develop repeatable validation and reconciliation using Python and PySpark, integrating checks with migration testing and configuration-driven pipelines.
Architecture
A validation engine compares source and target data. Schema, record count, null, duplicate, data type and partition checks feed reconciliation and a quality report.
- 01Source data
- 02Validation engine
- 03Reconciliation
- 04Quality report
Technical decisions
Express checks through reusable utilities and configuration. Keep results explicit so a difference can be investigated rather than hidden behind a generic pass state.
Challenges
Different storage representations and processing paths can make comparisons difficult. Validation needs to distinguish expected changes from unexpected differences.
Data validation
Combine structural checks with record-level and aggregate comparisons appropriate to the dataset. Treat reconciliation as part of migration readiness.
- Schema
- Row count
- Partitions
- Null values
- Duplicates
- Data types
- Reconciliation
Performance
Use distributed comparisons where suitable and consider the cost of scans. Choose checks that provide useful evidence without unnecessary repeated processing.
Outcome
Ongoing quality work supports repeatable migration validation and more visible source-to-target differences.
Lessons learned
Trustworthy data requires evidence. Build validation into the workflow instead of treating it as a final manual review.
