From POC to production
Technical architecture for a Databricks lakehouse with 100+ TB and 120+ datasets. Defined ingestion, modeling, and governance standards used by engineering and analytics teams.
EXPERIENCE
At Veeva, I have helped evolve an enterprise data platform from proof of concept to a production Databricks lakehouse—then extended that foundation into enterprise integrations and applied AI.
January 2025–Present
I design and evolve enterprise data platforms—from lakehouse architecture and integration standards to governed analytics and applied AI. At Veeva, my work has helped take a Databricks platform from proof of concept to production.
Technical architecture for a Databricks lakehouse with 100+ TB and 120+ datasets. Defined ingestion, modeling, and governance standards used by engineering and analytics teams.
End-to-end architecture for batch and near-real-time integrations across SaaS, REST APIs, AWS Kinesis, and SFTP, using a bronze/silver/gold medallion design.
Photon, incremental processing, partition pruning, and Spark tuning; governance and reliability patterns using Unity Catalog, data contracts, lineage, RBAC, and automated data quality.
LLM/RAG workflows for document processing, structured extraction, metadata enrichment, and vector search. Lead architecture reviews and mentor engineers on Spark, data modeling, and production design.
CURRENT FLAGSHIP INITIATIVE
I am the technical owner for the data engineering side of an ongoing enterprise CRM migration. My scope includes migration, source-to-target mapping, reconciliation and data quality, integration architecture, and ingestion from the new CRM.
It also includes operational and analytical reporting, downstream modeling, permissions and governance, and coordination with business and technical stakeholders.
View the ownership overviewPROFESSIONAL IMPACT
≈35%
Reduction in Spark-related processing cost.
≈39.6M
Historical records processed in a major ingestion initiative.
100+ TB
Platform scale: the production Databricks lakehouse I helped evolve.
120+
Platform scale: datasets across the enterprise lakehouse.
Cost reduction and historical ingestion are approximate outcomes across my professional work, separate from the ongoing CRM migration. Volume and dataset counts describe platform scale, not sole individual accomplishment.
CAREER PROGRESSION
Databricks · Apache Spark · Delta Lake · Unity Catalog · Snowflake · dbt Cloud
Python · SQL · PySpark · Structured Streaming · REST APIs · AWS Kinesis · SFTP
Apache Airflow · Databricks Jobs · Git · CI/CD · Data contracts · Lineage · RBAC · Data quality
LLMs · RAG · Document processing · Structured extraction · Embeddings · Vector search · AI agents
Texas Tech University · December 2021
Texas Tech University · December 2020
CONTACT
Connect with me about enterprise data platforms, technical leadership, or applied AI.