DQ

Portfolio / Databricks

Databricks projects

Lakehouse engineering and ML at scale: Unity Catalog governance, declarative pipelines, streaming and MLOps.

Unity Catalog Governance Rollout

01

Migrated multiple workspaces from Hive metastore to Unity Catalog with fine-grained access control.

Challenge

Inconsistent permissions across workspaces and no lineage made audits slow and risky.

Solution

  • Catalog/schema design aligned to domains and environments
  • Automated table migration with UCX and validation reports
  • Row filters and column masks for PII protection
  • System tables for lineage, audit and usage reporting
Workspaces unified
5
Tables governed
12k+
Lineage coverage
100%
  • Unity Catalog
  • UCX
  • System Tables
  • Terraform
  • Azure Databricks

Streaming CDC with Lakeflow Pipelines

02

Near real-time CDC from operational databases into Delta using declarative pipelines.

Challenge

Nightly full extracts strained source systems and left downstream data a day stale.

Solution

  • Change data feeds landed via Auto Loader into streaming tables
  • APPLY CHANGES for SCD Type 1/2 handling
  • Expectations for data quality with quarantine tables
  • Deployed with Databricks Asset Bundles and CI/CD
Data latency
24h → 5m
Source load
-70%
Pipeline SLA
99.9%
  • Lakeflow / DLT
  • Auto Loader
  • Structured Streaming
  • Delta
  • Asset Bundles

Churn Prediction MLOps

03

Production ML lifecycle combining my data science roots with platform engineering.

Challenge

Models were retrained manually in notebooks, with no versioning or drift monitoring.

Solution

  • Feature tables in Unity Catalog shared across teams
  • MLflow tracking, model registry and champion/challenger aliases
  • Scheduled retraining with Databricks Workflows
  • Model Serving endpoint plus Lakehouse Monitoring for drift
Retention lift
+18%
Auto retraining
Weekly
Deploy time
Hours → min
  • MLflow
  • Feature Engineering
  • Model Serving
  • Workflows
  • scikit-learn

Platform Cost Optimization

04

Cut Databricks spend while increasing workload volume.

Challenge

Rapid adoption led to oversized all-purpose clusters and unpredictable monthly bills.

Solution

  • Cluster policies and job clusters replacing interactive compute for pipelines
  • Photon and serverless SQL warehouses for BI workloads
  • Liquid clustering and OPTIMIZE / VACUUM schedules
  • Cost dashboards built on billing system tables
Monthly DBU spend
-38%
Query speed
2x
Tagged spend
100%
  • Photon
  • Cluster Policies
  • Databricks SQL
  • Liquid Clustering
  • System Tables