Real-World Data Scientist
Eight-plus years wrangling and analyzing large-scale EHR and administrative claims data, including hands-on analysis at a 150-million-patient scale, using SQL, R, and Python. I scope analysis requests, deliver against agreed deadlines, and am my team's Truveta subject-matter expert: I independently mastered the platform, including Truveta Studio and its proprietary Prose coding language, within roughly three months of gaining early access, and have since trained five-plus colleagues and built Jupyter- and Databricks-based analyses across multiple disease areas.
- 150M+
- patient EHR scale analyzed
- 3 mo
- to independently master Truveta Studio and Prose
- 5+
- colleagues trained, junior to senior level
- 6
- disease areas addressed on the platform
Core expertise
- Real-world data and EHR analysis at scale. Wrangling and analyzing large-scale EHR and administrative claims data using SQL, R, and Python to scope and deliver on deadline.
- Truveta platform mastery. Truveta Studio and its proprietary Prose language, learned independently and then taught to colleagues at junior and senior levels.
- Notebook and pipeline development. Jupyter- and Databricks-based analyses, including regex query writing and Spark/PySpark, across disease areas spanning C. diff, COVID-19, RSV, influenza, healthcare-associated infections (HAIs), and UTIs.
- Record linkage and data quality. Patient-identity matching across sources (surveillance to Medicare, trial to registry), linkage quality classification, deduplication and incident-episode logic, and continuous-presence census logic.
- Multi-platform data engineering. ETL pipeline design and data modeling across Truveta, MarketScan, CMS claims, PointClickCare, Premier/PHD, Epic, and Cerner.
Honest scope note: my Truveta and Databricks work has been analytic and platform-enablement focused, not machine learning model development, and I have not worked with de-identification or quasi-identifier methods on Truveta data.
Experience
- Directly utilize the Truveta real-world data platform, integrating it with MarketScan, CMS, PointClickCare, and Premier to evaluate COVID-19 vaccine effectiveness, safety, and breakthrough infections across a multi-million-patient EHR database.
- Became the team's Truveta subject-matter expert after being among the first granted platform access, independently mastering Truveta Studio and Prose within roughly three months.
- Built Jupyter- and Databricks-based analyses, including regex query writing and Spark/PySpark, addressing research questions on C. diff, COVID-19, RSV, influenza, HAIs, and UTIs.
- Trained 5+ colleagues at junior and senior levels to work independently within the Truveta platform.
- Perform patient-identity matching and record linkage across data sources, including CDC surveillance to CMS Medicare claims, with deduplication, incident-episode, and stay-reconciliation logic.
- Build analytics and visualizations to detect fraud, waste, and abuse in Medicaid (T-MSIS) and Medicare claims, programming multi-year workflows in SAS (PROC SQL, macros) and Python (pandas).
- Work hands-on in Databricks, Snowflake, and CDC's 1CDP/DCIPHER platform (Contour, Code Workbook, Quiver, Workshop).
- Programs statistical analyses in R and SAS producing reproducible, well-documented code against agreed-upon deadlines.
- Develops evidence synthesis reports, publication-quality tables, figures, and reusable documentation translating complex platform-specific data concepts into clear guidance.
Built CDISC-compliant SDTM datasets and data validation workflows, and programmed SAS and R workflows for automated data quality checks and reconciliation.
Maintained epidemiologic case data within state surveillance systems, managing 150+ weekly investigations to support real-time public health decision-making.
Extracted and cleaned analysis-ready datasets from EHR pulls at a 2,000+ patient scale, applying multivariable regression and Power BI visualization for clinical stakeholders.
Stack
Independently mastered a proprietary coding language (Prose) and platform within about three months, then became the person my team turned to for it.
View the GitHub portfolio