Software
Programmer

Architecting scalable data platforms, analytics pipelines, and research infrastructure.

Projects

Systems built around messy, high-volume scientific data.

Harvard T.H. Chan School of Public Health

Clinical Outcomes Data Pipeline

Python and SQL pipelines that clean 500K+ daily patient records, standardize VA and EHR datasets, and prepare data for regression analysis and clinical outcome modeling.

Python SQL EHR Clinical Modeling

Caltech

Distributed Spatial Transcriptomics Processing

Apache Spark and AWS workflows for efficient analysis of petabyte-scale spatial transcriptomics data, paired with Python ingestion services and retry-aware batch orchestration.

Apache Spark AWS Batch Petabyte-scale Genomics

UC Santa Cruz Genomics Institute

Electrophysiology Batch Processing

Parallel Python pipelines for large electrophysiology datasets and clinical records, reducing runtimes from roughly 6 hours to under 2 hours while supporting 3TB+ workloads.

Parallel Python Statsmodels SciPy 3TB+

Experience

Selected Work

Harvard T.H. Chan School of Public Health Aug 2021 - Present Boston

Programmer II

  • Built Python and SQL pipelines to clean 500K+ daily patient records for regression analysis and clinical outcome modeling.
  • Developed parallel batch-processing workflows for 1,000+ automated batches in hours.
  • Processed and standardized VA and EHR datasets to support reliable downstream analysis.
  • Refactored 20+ data processing scripts to improve reliability and support robust statistical inference.
Caltech Aug 2020 - Aug 2021 Pasadena, CA

Senior Software Engineer

  • Built and maintained a scalable Python data analysis pipeline, reducing processing time by 30-50 percent.
  • Developed Apache Spark pipelines for petabyte-scale spatial transcriptomics analysis on AWS.
  • Built Python and AWS Batch ingestion services with retry logic and parallel orchestration, processing 10,000+ files per week with under 1 percent failure.
Citibank Mar 2020 - Aug 2020 New York City Metro

Quantitative Software Engineer

  • Automated metadata ingestion and transformation in Python, reducing manual handling time by roughly 40 percent.
  • Streamlined integration workflows with structured SQL procedures, reducing manual intervention by roughly 30 percent.
  • Applied machine learning in Python to infer and assign column names, cutting data lineage documentation effort by roughly 25 percent.
UC Santa Cruz Genomics Institute Nov 2018 - Mar 2020

Data Engineer

  • Built parallel Python pipelines for large electrophysiology datasets, reducing processing time by roughly 40 percent and supporting 3TB+ workloads.
  • Implemented parallelized batch-processing that cut runtime from about 6 hours to under 2 hours for 50,000+ clinical records per run.
  • Performed dataset validation with Statsmodels and SciPy before downstream analytics.

Skills

Tools I use to move data from raw to reliable.

Languages

Python, SQL

Data Systems

Apache Spark, AWS Batch, parallel processing, ingestion pipelines

Research Data

EHR, VA datasets, spatial transcriptomics, electrophysiology, clinical records

Analysis

Regression workflows, clinical outcome modeling, Statsmodels, SciPy, machine learning

Contact

Open to research software, data engineering, and computational health collaborations.

Connect on LinkedIn