Professional focus
I design and evolve enterprise data platforms—from lakehouse architecture and integration standards to governed analytics and applied AI. At Veeva, my work has helped take a Databricks platform from proof of concept to production.
Experience
Senior Data Engineer · Veeva Systems
January 2025–Present
- From POC to production: Technical architecture for a Databricks lakehouse with 100+ TB and 120+ datasets. Defined ingestion, modeling, and governance standards used by engineering and analytics teams.
- Integration reference design: End-to-end architecture for batch and near-real-time integrations across SaaS, REST APIs, AWS Kinesis, and SFTP, using a bronze/silver/gold medallion design.
- Performance & trust: Photon, incremental processing, partition pruning, and Spark tuning; governance and reliability patterns using Unity Catalog, data contracts, lineage, RBAC, and automated data quality.
- Applied AI & technical leadership: LLM/RAG workflows for document processing, structured extraction, metadata enrichment, and vector search. Lead architecture reviews and mentor engineers on Spark, data modeling, and production design.
Earlier experience
Data Engineer · Veeva Systems
January 2022–January 2025
- Built and operated production batch and streaming pipelines with Databricks, Spark, Delta Lake, AWS Kinesis, and REST APIs.
- Standardized orchestration, CI/CD, testing, monitoring, and alerting with Apache Airflow, Databricks Jobs, and Git. Developed ETL/ELT models in PySpark, SQL, and dbt.
Software Developer Intern · Texas Tech University Health Sciences Center
January 2021–December 2021
- Designed Snowflake data models and incremental ingestion with Apache Airflow and Snowflake Tasks for near-real-time reporting.
Software Engineering Intern · Discover Financial Services
June 2021–August 2021
- Developed ETL workflows for financial transaction data, including profiling, validation, and anomaly detection to support SOX-compliant analytics.
Current initiative: enterprise CRM migration
Technical owner for data engineering. Scope includes migration into the new CRM; source-to-target mapping; reconciliation and data quality; integration architecture and ingestion; operational and analytical reporting; downstream modeling; permissions/governance; and business and technical stakeholder coordination.
Selected professional impact
- Approximately 35% — Reduction in Spark-related processing cost.
- Approximately 39.6M — Historical records processed in a major ingestion initiative.
Approximate figures across professional work; separate from the ongoing CRM migration.
Technical scope
Lakehouse & analytics: Databricks, Apache Spark, Delta Lake, Unity Catalog, Snowflake, dbt Cloud
Languages & integration: Python, SQL, PySpark, Structured Streaming, REST APIs, AWS Kinesis, SFTP
Production engineering: Apache Airflow, Databricks Jobs, Git, CI/CD, Data contracts, Lineage, RBAC, Data quality
Applied AI: LLMs, RAG, Document processing, Structured extraction, Embeddings, Vector search, AI agents
Public engineering work
DataNepal: canonical geographic modeling, provenance enforced at export, and a static data publishing architecture using dlt, DuckDB, dbt, Parquet/JSON, and DuckDB-WASM.
Education
M.S. Software Engineering · Texas Tech University
December 2021
B.S. Computer Science · Texas Tech University
December 2020
Credentials listed on my resume
- Databricks Certified Data Engineer Associate
- AWS Certified Cloud Practitioner