SENIOR DATA ENGINEER · DAYTON, OH

Pipelines that survive the audit.

Five-plus years engineering governed lakehouse platforms for payers and health systems at Fifth Third Bank, The Travelers, and Molina Healthcare. Sub-15-minute streaming latency. 99.8% freshness SLAs. Six-figure cost cuts.

LAT — LON — 00:00:00 EST
SCROLL
APACHE SPARKKAFKADELTA LAKEDATABRICKSSNOWFLAKEAWSAZUREAIRFLOWdbtTERRAFORM APACHE SPARKKAFKADELTA LAKEDATABRICKSSNOWFLAKEAWSAZUREAIRFLOWdbtTERRAFORM
01 · IMPACT

Numbers that held up in production.

0
DAILY CLAIMS & POLICY RECORDS PROCESSED
0
REPORTING LATENCY, DOWN FROM 24 HOURS
0
DATA FRESHNESS SLA ACROSS GOLD TABLES
0
ANNUAL CLOUD COST REDUCTIONS DELIVERED
02 · EXPERIENCE

Three companies, one thread: trustworthy data.

From regulatory claims pipelines at Molina to real-time fraud signals at Fifth Third, the work has stayed close to healthcare and financial data that has to be right, on time, and provable.

Senior Data EngineerFifth Third Bank
MAY 2024 — PRESENT
Cincinnati, OH
  • Architected end-to-end PySpark pipelines on AWS EMR processing 2M+ daily insurance policy and claims records with exactly-once semantics, cutting batch processing time by 40%.
  • Designed a bronze/silver/gold Delta Lake medallion architecture on S3 consolidating 10 siloed sources, enabling self-serve analytics for 200+ business users.
  • Engineered real-time CDC ingestion with Kafka, Kinesis, and Debezium, cutting reporting lag from 24 hours to under 15 minutes with zero-downtime schema evolution.
  • Deployed Great Expectations and Monte Carlo observability with 100+ validation rules, cutting data incidents by 60%.
  • Implemented Microsoft Purview lineage across 400+ datasets, passing three consecutive HIPAA audits.
  • Optimized Delta Lake via Z-ordering and VACUUM scheduling, cutting Redshift compute costs by $110K annually; mentored three engineers.
PySparkAWS EMRDelta LakeKafkaPurviewTerraform
Data EngineerThe Travelers Inc.
JUL 2022 — JUN 2023
Atlanta, GA
  • Delivered Azure Data Factory pipelines processing 1.5M+ daily records from mainframe, Oracle, and Salesforce into Synapse Analytics, improving analytical availability by 70%.
  • Tuned Spark jobs on Databricks with broadcast joins and partitioning, reducing runtime by 43% and compute spend by $90K per year.
  • Built Kafka event streams on Event Hubs with Debezium CDC, achieving sub-200ms alert latency with exactly-once semantics.
  • Architected a lakehouse on ADLS Gen2, consolidating eight siloed stores and reducing storage costs by 30%.
  • Established Jenkins CI/CD with pytest and dbt test gates at 82% code coverage, replacing 25+ cron jobs with Airflow DAGs.
Azure DatabricksSynapseEvent HubsdbtAirflow
Junior Data EngineerMolina Healthcare
MAR 2020 — JUN 2022
Long Beach, CA
  • Built Azure Data Factory pipelines ingesting claims, member, and provider data from six source systems, enabling consolidated analytics for 400K+ member records.
  • Built PySpark ETL for claims adjudication data on Databricks, reducing nightly batch time by 35%.
  • Authored 80+ Great Expectations rules, catching 12+ critical quality incidents before downstream impact.
  • Contributed to HEDIS and CMS Star Ratings pipelines covering 3M+ annual claims across 22 quality measures with 100% on-time submission.
  • Built Apache Atlas lineage mapping for 120+ regulatory datasets, supporting NCQA and CMS audit traceability.
Azure Data FactoryPySparkGreat ExpectationsApache Atlas
03 · PROJECTS

Selected builds, not just tickets closed.

Platforms and frameworks built beyond the day-to-day roadmap, each one closing a specific reliability or governance gap.

01
FIFTH THIRD BANK · 2025 — PRESENT
Real-Time Claims & Fraud Signal Streaming Platform
Event-driven platform ingesting 2M+ daily claims and policy events through Kafka and Kinesis into a Spark Structured Streaming layer, cutting anomaly detection latency from hours to under 15 minutes.
KafkaKinesisSpark StreamingDelta LakeDatadog
02
THE TRAVELERS · 2022 — 2023
Pharmacy & Care Management 360 Data Integration
Unified pharmacy claims, prior authorization, and care management data from seven source systems into one member view, with incremental Delta Live Tables auto-loading 150GB+ of daily clinical data.
ADFDelta Live TablesEvent HubsMLflow
03
FIFTH THIRD BANK · 2025 — PRESENT
Enterprise Data Quality & Lineage Governance Framework
DataOps framework combining Great Expectations, Monte Carlo, and dbt tests with OpenLineage tracking across every gold-layer table, cutting undetected data incidents by 60%.
Great ExpectationsMonte CarloPurviewOpenLineage
04
MOLINA HEALTHCARE · 2020 — 2022
Healthcare Claims Data Warehouse Migration
Led migration of legacy stored-procedure reporting to a PySpark and dbt pipeline on Databricks, redesigning star schema fact tables with actuarial and finance stakeholders.
DatabricksdbtSynapseTerraform
04 · STACK

The toolchain, end to end.

Streaming & Big Data

Apache Spark · Structured Streaming
Apache Kafka · Kafka Connect · Streams
Debezium · CDC · Schema Registry
Delta Lake · Delta Live Tables
Hive · HDFS · Parquet · Avro

Cloud & Warehousing

AWS · EMR · Glue · Redshift · Kinesis
Azure · Databricks · Synapse · ADLS Gen2
Snowflake · Unity Catalog
Star Schema · Data Vault 2.0
Medallion Architecture

Orchestration & Quality

Airflow · dbt Core · dbt Cloud
Great Expectations · Monte Carlo · Soda
Microsoft Purview · Apache Atlas
Terraform · Docker · Kubernetes
GitHub Actions · Jenkins · Datadog
05 · EDUCATION & CERTIFICATIONS

Formal training, kept current.

Master of Science, Computer Science
University of Dayton · Dayton, OH
Bachelor of Technology, Computer Science Engineering
Parul University · Vadodara, India
AWS Certified Data Engineer, Associate (DEA-C01)
Databricks Certified Data Engineer Associate
dbt Certified Analytics Engineer
Google Cloud Professional Data Engineer, in progress, target Q3 2026
AVAILABLE FOR SENIOR DATA ENGINEERING ROLES

Let's build something
worth auditing.