Lakehouse engineering and ML at scale: Unity Catalog governance, declarative pipelines, streaming and MLOps.
Unity Catalog Governance Rollout
01Migrated multiple workspaces from Hive metastore to Unity Catalog with fine-grained access control.
Challenge
Inconsistent permissions across workspaces and no lineage made audits slow and risky.
Solution
- Catalog/schema design aligned to domains and environments
- Automated table migration with UCX and validation reports
- Row filters and column masks for PII protection
- System tables for lineage, audit and usage reporting
- Workspaces unified
- 5
- Tables governed
- 12k+
- Lineage coverage
- 100%
- Unity Catalog
- UCX
- System Tables
- Terraform
- Azure Databricks
Streaming CDC with Lakeflow Pipelines
02Near real-time CDC from operational databases into Delta using declarative pipelines.
Challenge
Nightly full extracts strained source systems and left downstream data a day stale.
Solution
- Change data feeds landed via Auto Loader into streaming tables
- APPLY CHANGES for SCD Type 1/2 handling
- Expectations for data quality with quarantine tables
- Deployed with Databricks Asset Bundles and CI/CD
- Data latency
- 24h → 5m
- Source load
- -70%
- Pipeline SLA
- 99.9%
- Lakeflow / DLT
- Auto Loader
- Structured Streaming
- Delta
- Asset Bundles
Churn Prediction MLOps
03Production ML lifecycle combining my data science roots with platform engineering.
Challenge
Models were retrained manually in notebooks, with no versioning or drift monitoring.
Solution
- Feature tables in Unity Catalog shared across teams
- MLflow tracking, model registry and champion/challenger aliases
- Scheduled retraining with Databricks Workflows
- Model Serving endpoint plus Lakehouse Monitoring for drift
- Retention lift
- +18%
- Auto retraining
- Weekly
- Deploy time
- Hours → min
- MLflow
- Feature Engineering
- Model Serving
- Workflows
- scikit-learn
Platform Cost Optimization
04Cut Databricks spend while increasing workload volume.
Challenge
Rapid adoption led to oversized all-purpose clusters and unpredictable monthly bills.
Solution
- Cluster policies and job clusters replacing interactive compute for pipelines
- Photon and serverless SQL warehouses for BI workloads
- Liquid clustering and OPTIMIZE / VACUUM schedules
- Cost dashboards built on billing system tables
- Monthly DBU spend
- -38%
- Query speed
- 2x
- Tagged spend
- 100%
- Photon
- Cluster Policies
- Databricks SQL
- Liquid Clustering
- System Tables