Harvard T.H. Chan School of Public Health
Clinical Outcomes Data Pipeline
Python and SQL pipelines that clean 500K+ daily patient records, standardize VA and EHR datasets, and prepare data for regression analysis and clinical outcome modeling.
Architecting scalable data platforms, analytics pipelines, and research infrastructure.
Large-scale pipelines for cleaning, standardizing, and preparing health data for research analysis.
Spatial Transcriptomics Petabyte-scale research dataDistributed processing patterns for high-volume biological datasets and computational workflows.
Quant Engineering Financial data systems at CitibankMetadata ingestion, SQL workflows, and machine learning for structured financial data operations.
Projects
Harvard T.H. Chan School of Public Health
Python and SQL pipelines that clean 500K+ daily patient records, standardize VA and EHR datasets, and prepare data for regression analysis and clinical outcome modeling.
Caltech
Apache Spark and AWS workflows for efficient analysis of petabyte-scale spatial transcriptomics data, paired with Python ingestion services and retry-aware batch orchestration.
Citibank
Automated metadata ingestion, SQL integration workflows, and Python machine learning to infer column names and reduce manual data lineage documentation effort.
UC Santa Cruz Genomics Institute
Parallel Python pipelines for large electrophysiology datasets and clinical records, reducing runtimes from roughly 6 hours to under 2 hours while supporting 3TB+ workloads.
Experience
Skills
Python, SQL
Apache Spark, AWS Batch, parallel processing, ingestion pipelines
EHR, VA datasets, spatial transcriptomics, electrophysiology, clinical records
Regression workflows, clinical outcome modeling, Statsmodels, SciPy, machine learning
Contact