Economics · Statistics · Machine learning

I turn messy records into evidence.

Thirty years of FOMC transcripts, a million protein decoys, a stack of payer policy PDFs. Different domains, same job: figure out how to measure the thing before you model it.

I finished a BA in Mathematics and Economics at Hamilton College in May 2026 (Honors in Economics, Soper Research Prize) and started an MS in Statistics at Rutgers this fall. I am open to leaving the program for the right full-time research or analyst role, and I am equally happy to talk about Summer 2027.

My main work so far is a paper with Prof. Ann Owen, a former Fed economist, that uses LLMs to read every FOMC participant's policy stance out of the transcripts and then asks who actually moves the room. I also built a graph transformer for protein structure scoring, shipped an iOS app, and ported a Stata econometrics command to Python because I needed it and it did not exist.

When I am not at a keyboard I am usually at a piano or a basketball hoop.

Benjamin Zhao
Princeton, NJ → Clinton, NY → New Brunswick, NJ
Now
MS Statistics, Rutgers (regression analysis, probability theory this term)
Looking for
Data analyst, research analyst, and quant roles, especially in health and finance. Start now or Summer 2027.
Based in
New Jersey. Open to NYC, Boston, DC, Philadelphia, and elsewhere.
Latest
FOMC paper submitted to the Journal of Money, Credit and Banking

Research

Working paper and a benchmark submission
CASP17 submission Jan 2026 – Mar 2026 Advisor: Prof. Shawn Chen, Hamilton CS

DProQV2: protein structure quality assessment with a gated graph transformer

Research assistant, Hamilton College Computer Science.

A 12-layer gated graph transformer that predicts per-residue lDDT for protein decoys, combining ESM-C sequence embeddings with Rosetta energy terms as node and edge features. Trained on more than a million decoy structures on Hamilton's HPC cluster (multi-GPU, NVIDIA A100) and submitted to the CASP17 benchmark.

My part was the experiment design and the training infrastructure: data loaders that did not choke on a million graphs, checkpointing, and keeping runs reproducible across nodes.

Model
12-layer gated graph transformer, rotation-invariant features
Scale
1M+ decoys, multi-GPU training
Stack
PyTorch, PyTorch Geometric, Slurm, ESM-C, Rosetta

Projects

Things that shipped or that other people can run

xtscc

2026

Python port of Stata's xtscc: Driscoll-Kraay standard errors for panels with cross-sectional and serial dependence (Hoechle 2007). Written because I needed it for the FOMC paper and could not find a faithful implementation. Validated against Stata output and against Bobrov et al. (2025).

pythoneconometricspanel-data
GitHub

MyNextRead

May 2026 – now

iOS app that finds books through real endorsements rather than ratings. SwiftUI front end on a FastAPI backend running a multi-stage GPT-4.1 pipeline: parallel agents search for endorsements, a verifier agent plus deterministic filters reject false positives. Persistent caching cut repeat-query cost to near zero; adds cost telemetry, rate limiting, and server-side auth. Monetized with AdMob and Amazon affiliates.

swiftfastapillm-agents
Code

Pulls CPT and ICD-10 procedure and diagnosis pairs, with their medical-necessity status, out of semi-structured payer policy PDFs. LangChain, GPT-4, and Pydantic schemas for typed output; deployed as a Gradio app on Hugging Face. The kind of document a health insurer's analytics team spends a lot of hours reading by hand.

langchainhealthcarepydantic
GitHub

HamPlan

Nov – Dec 2025

RAG course recommender for Hamilton. An ETL pipeline harvests course APIs, department pages, and PDF syllabi (session-authenticated) into one dataset; a GPT-4 Turbo recommender sits on top with rule-based PDF chunking, custom vector search, and conversation memory. Demoed live to faculty, 3.8/5 satisfaction.

ragetlembeddings
GitHub

Bi-LSTM sentiment model with a capacity-scaling study (256 to 2048 dims), 91% on IMDB 50K. UNet++ segmentation with deep supervision, a hybrid loss, and CBAM attention, 80.9% mIoU on Oxford-IIIT Pet. Whisper fine-tune on Hugging Face.

pytorchnlpvision
GitHub

Thirty-odd repos in total, including a job-hunting agent built on the Anthropic SDK, healthcare provider (NPI/PECOS) queries, and an NHS symptom scraper. All of it is on GitHub.

Experience & education

Experience

Jan 2026 – Aug 2026
Co-author, FOMC deliberation paper
with Prof. Ann Owen, Hamilton College Economics
LLM measurement pipeline, panel econometrics, revisions
Jan – Mar 2026
Research Assistant, Computer Science
Hamilton College, Prof. Shawn Chen
Deep learning for protein structure quality; HPC workflows
May – Jun 2025
Summer Research Assistant, Economics
Hamilton College
Automated ETL and quality checks for UN, IMF, and World Bank trade and aid flow data
Aug 2024 – May 2025
Tutor, Quantitative & Symbolic Reasoning Center
Hamilton College
Intro Economics and Macroeconomic Theory

Education & honors

2026 – 2028
MS Statistics
Rutgers University, New Brunswick
In progress. Would leave for the right role.
2022 – 2026
BA Mathematics and Economics
Hamilton College, Clinton NY
GPA 3.71, Honors in Economics, Dean's List. Real Analysis, Probability and Statistical Inference, Data Structures and Algorithms, Deep Learning, Econometrics, Financial Economics
May 2026
Soper Research Prize
Hamilton College Economics Department
Mar 2023
2nd place, ASA Undergraduate Statistics Class Project Competition
American Statistical Association
Aug 2026
BlueDot Impact, Future of AI course
Certificate

Toolkit

Daily drivers first

Languages

  • Python
  • R
  • SQL
  • Stata
  • Linux shell
  • C++, Swift, MATLAB

Statistics & econometrics

  • Panel methods, fixed effects
  • Driscoll-Kraay, clustered SEs
  • Regression, inference
  • Time series
  • Experiment design

ML & LLMs

  • PyTorch, PyTorch Geometric
  • Graph transformers, GNNs
  • Claude & OpenAI APIs, LangChain
  • RAG, embeddings, vector search
  • Hugging Face, TensorFlow

Data & cloud

  • AWS (S3, EC2, Athena)
  • Snowflake, Databricks, Azure
  • ETL, data quality checks
  • Tableau, Power BI
  • Git, REST APIs, Slurm

Outside of work

The non-quantitative parts
Piano

Started at five, and studied with Rita Shklar from 2014 through high school. If I get one piece, it is Ginastera's Danzas Argentinas. In 2020 I placed 3rd at the ENKOR International Competition and took 1st and 2nd places in the Great Composer Competition series, and I have played in master classes with Natasha Paremski and Daniel Epstein.

Princeton Symphony master class write-up · Recordings

Paint

I painted through high school, mostly oil and gouache, mostly ordinary scenes: breakfast, a barbershop in Princeton, my dad reading at night. Fish Market won a Scholastic Art & Writing Gold Key. I have not picked up a brush since, but a few of the pieces still hang around, so they are on the right.

Full portfolio

Basketball

Co-president of the Hamilton College Basketball Club, 2025 to 2026. Not a competitive player; I just like the repetition of shooting, and running a club taught me more about scheduling people than any course did.

Languages

English, Mandarin (fluent speaking, still working on reading and writing), and intermediate Spanish, maintained mostly through a Spanish playlist I started during lockdown.

Fish Market, oil painting
Fish Market, oil. Scholastic Gold Key.
The Continental Barbershop in Princeton
The Continental Barbershop, Princeton
My Young Dad, Light in the Night
My Young Dad, Light in the Night

Get in touch

If you are hiring for a data or research analyst role and any of the above looks useful, I would like to hear from you. I am a US citizen and do not need sponsorship. I reply to email.