Dave Liu — Portfolio

Dave Liu

ML Research Engineer · Senior Data Scientist · Systems Builder

Resume (PDF) View Resume Get in Touch

About Me

I knew I wanted to work in data before I knew what a data scientist was. At ten years old I noticed a pattern: civilizations chased gold and spices, then oil, and now—in the information age—data. It struck me that data is the modern gold, and that understanding the world at any real depth requires massive amounts of it. That idea has guided every career decision I've made since.


I believe the world is far more complex and nuanced than most people think—or care to see. Beneath every decision, every behavior, every interaction lies a web of signals that, if you look closely enough, tells a richer story than the surface ever could. To me, understanding those nuances isn't just intellectually satisfying; it's how we build a world that works better for everyone. Predicting the future starts with genuinely understanding the past and present, and I think the best version of our future begins by understanding people.


That philosophy has taken me across biotech, e-commerce, healthcare, recruiting, and fintech over the past decade. At Shipt I lead personalization for millions of grocery shoppers. Before that I helped build a cancer detection model at Freenome, designed a task-matching system that saved Change Healthcare $7M a year, and built talent-ranking models at Riviera Partners. On nights and weekends I run AutoTrader, a fully autonomous stock prediction system I built from scratch.


What ties all of that together is a belief that the hardest part of data science is rarely the model—it's getting the right data, defining the right metric, and making sure the thing actually ships. I spend as much time asking “should we be measuring this at all?” as I do writing training loops. It's an approach that's led to production systems at every company I've been part of, and one I don't plan on changing. In 2026 it became an independent research program—three self-published papers on multi-agent synthesis, model selection under distribution shift, and LLM-grounded feature extraction—all built on the same rule I use in production: the system that produces a number is never the one that grades it.

Experience

September 2024 – Present
Senior Data Scientist — Shipt (Personalization Team)

Lead data scientist on Shipt's Personalization team, architecting recommender and eligibility systems behind personalized shelves for millions of grocery shoppers. Mentored 4 data scientists, including 2 direct reports.

  • Built and launched Deals For You V2, replacing matrix factorization with two-tower retrieval and precision reranking. A CUPED-corrected A/B test drove +91% shelf engagement, +1.76% order incidence, and +1.41% orders per user.
  • Audited pre-treatment imbalance and a variant activation gap before SVP planning, correcting the annualized estimate from +83K to ~50K incremental orders (95% CI: 20K–80K) and ~$3.7M directional GMV. The analysis preserved a real win while retiring an inflated topline.
  • The validated result prompted leadership to re-plan H2 around Deals as a potential ~$6M contributor to Shipt's $20M incremental-GMV target, prioritizing scale-up, holdback measurement, and reuse of the eligibility platform.
  • Architected a reusable real-time Eligibility platform: a store-scoped hard pre-filter before approximate nearest-neighbor retrieval with swappable content tables. Redesigned an initial 448M-row footprint down by ~95% to ~20M store-specific rows plus a compact global set, with zero-downtime refreshes and p99 <750 ms serving.
  • Across the broader personalization portfolio, delivered up to 23% Personalized GMV lift across A/B-tested shelves and a cold-start strategy with +47% engagement and +8% first-time orders.
  • Created the Personalization Interaction Score (views → clicks → add-to-carts → purchases) and an automated Shelf Attribution pipeline. Challenged an irreproducible 5% attribution claim, established the defensible 1–4% range, and raised the organization's measurement standard.
  • Partnered with Engineering to replace CSV-based recommendation delivery with Kafka pipelines across legacy recommenders. During the team's first four months, served as the sole data scientist while maintaining all 16 Discovery Science repositories.
  • Designed a Retrieval-Augmented Generation system over Shipt's retailer catalogs and built the business case that secured investment in an agentic-AI framework.
November 2020 – June 2024
Machine Learning Research Engineer — Freenome

Core ML engineer at a genomics company developing a blood test for early-stage cancer detection. Worked at the intersection of infrastructure and research, building the systems that scientists depend on daily.

  • Key contributor to Freenome's core product: a multiomics cancer detection model that predicts cancer stage (1–4) from blood-draw data. Built data abstractions to handle petabyte-scale genomic datasets, unblocking cross-analyte feature development and meaningfully accelerating training and evaluation cycles.
  • Built a model comparison system that tracks research versus production model performance side-by-side—a requirement for FDA audit compliance and a tool that gave the team confidence that production models stayed aligned with research intent.
  • Designed and built large portions of Freenome's distributed ML training and serving platform, used daily by 30+ scientists and researchers. Key contributions included scaling CPU-bound training across O(100) machines, supporting multiple evaluation strategies (leave-one-out, K-fold), and building model artifact storage that made reproducibility straightforward.
  • Led the adoption of PyTorch, MLFlow, and RayTune across the ML team, replacing legacy tooling to enable GPU acceleration, experiment tracking, and hyperparameter tuning at scale.
  • Built a cloud cost monitoring system that surfaced the biggest storage and compute expenses in GCP. The visibility alone drove optimizations that saved the company over $10M annually—making it one of the highest-ROI projects I've worked on.
  • Recognized with a Servant Leadership Award, elected by managers and peers across the engineering organization.
January 2020 – November 2020
Sr. Machine Learning Engineer — Change Healthcare

Worked on ML systems for health insurance claims processing—a domain where model accuracy translates directly into operational cost savings.

  • Designed a ranking model that matches human workers to claims processing tasks based on skill, history, and task complexity. The model delivered $7M in annual value by reducing the volume of manual task assignment and the need for additional hires.
  • Built a classification model for partitioning sensitive patient documents, using image and text data to route health insurance claims to the correct processing workflow.
  • Developed internal AWS cloud tooling and production API infrastructure to support the ML team's deployment pipeline.
  • Led a cross-functional tiger team to prototype a conversational chatbot using Rasa and HuggingFace's NLP library for internal claims inquiry workflows.
January 2019 – December 2019
Data Scientist — Riviera Partners

Built ML models for an executive recruiting firm, working across the full pipeline from data collection to model serving.

  • Developed a suite of models: a classifier to estimate job-departure likelihood, a regression model to predict team sizes from resume features, and a ranking model to surface and match top candidates to open roles using a custom NDCG listwise loss function.
  • Built an end-to-end framework for rapid model prototyping, training, evaluation, and serving—enabling the team to iterate on new models without re-engineering infrastructure each time.
  • Wrote data collection scrapers to harvest structured candidate data from public sites and APIs.
January 2017 – December 2018
Undergraduate Researcher — UC Berkeley

Two concurrent research positions exploring ML applications in energy and neuroscience.

  • California Institute for Energy and Environment (CIEE): Built a recurrent neural network for predicting building energy usage, exploring how temporal patterns in consumption data can inform smarter grid management.
  • Bengson Research Lab, Sonoma State: Applied ML models to EEG data to computationally predict individualized occipital lobe activation patterns. The work showed early feasibility for brain-computer interface applications.
Earlier Roles

Where the foundation was built.

  • Data Science Intern, Castlight Health (2017) — Designed an entity matching and deduplication pipeline using gradient-boosted classifiers with hard negative mining. Achieved 85–95% precision/recall across hospital, facility, and practitioner entity types.
  • Data Science Contractor, Riviera Partners (2016) — Built a team size prediction model from public data and a Python wrapper for survival model time-series analysis. Set up Flask model-serving infrastructure.
  • URAP, Berkeley Institute of Data Science (2016) — Helped map UC Berkeley course progression through different majors by computationally organizing class taxonomies and running deduplication.
  • Data Science Intern, Doximity (2015) — Built a gradient-boosted classifier to identify malformed web-scraped articles and used reverse geocoding with fuzzy string matching to link doctor names in news articles to facility profiles.
Education
University of California, Berkeley — Class of 2018
  • BS in Computer Science and Data Science (Dual Degree)
  • Berkeley Institute of Data Science Undergraduate Research Apprenticeship (2015)

Technical Skills

Languages

Python SQL Bash Java C/C++

ML & Data

PyTorch XGBoost Scikit-learn Pandas MLFlow FAISS Gensim NLTK SpaCy

Cloud Platforms

GCP AWS Azure

Data Infrastructure

PostgreSQL Snowflake MySQL Spark Kafka

Orchestration

Airflow Flyte Metaflow GitHub Actions

Infrastructure

Docker Kubernetes Git CI/CD

What Colleagues Say

"Dave is one of the most thorough and thoughtful engineers I've worked with. He doesn't just build models—he builds the systems around them that make sure they actually work in production."

Former Colleague

Freenome

"What sets Dave apart is his willingness to ask the hard questions—about metrics, about assumptions, about whether we're solving the right problem. That intellectual honesty makes everyone around him better."

Former Manager

Shipt

"Dave taught me git, and somehow made it make sense. He has a rare ability to explain complex technical concepts in a way that doesn't make you feel stupid for not already knowing them."

Research Scientist

Freenome

Featured Projects


Other Projects

XGBoost Visual Guide (2026)

An interactive visual textbook explaining XGBoost and gradient boosting from first principles. 10 sections covering decision trees, ensemble methods, step-by-step gradient boosting with animated residual reduction, learning rate effects, XGBoost-specific innovations (histogram splits, sparsity handling), overfitting/early stopping, feature importance, and a hyperparameter cheat sheet. Built with D3.js and Chart.js.

Depression Classifier (2018–2022)

A personal project that grew out of curiosity about whether lifestyle patterns could predict mental health outcomes. Built an ML model trained on behavioral data (sleep, exercise, social activity, diet) that achieved a K-fold cross-validated AUC of 90% (n=110) at classifying depressive episodes. The model's feature importances were eye-opening enough to change some of my own habits—including joining a running club that I'm still part of.

Sentic Python Package (2017)

An open-source Python library for multi-dimensional sentiment analysis. Goes beyond positive/negative polarity to capture mood, attention, sensitivity, aptitude, and pleasantness across 20+ languages. Built on the SenticNet4 knowledge base. Available on PyPI.
(GitHub)

Myndful.us (2018)

An ML-powered habit-tracking web app designed to help users build healthier routines. I managed a team of 8 through design, development, and launch. The app analyzes journaling entries and activity logs to surface patterns and suggest personalized behavioral nudges.


ClimateChase (2016)

A strategy game built with Flask and React where players manage a country's energy portfolio, balancing investments across nuclear, solar, wind, and fossil fuels while responding to economic shocks and policy changes. A fun way to explore the tradeoffs in energy transition.

PDF-To-Audiobook Converter (2016)

A tool that chains together Google's Tesseract OCR engine with macOS text-to-speech to turn any PDF—including scanned documents—into listenable audio files. Built it because I wanted to "read" textbooks while running.

XRP Trade Algorithm (2014)

My first foray into algorithmic trading: a Python wrapper and terminal interface for the Ripple (XRP) API that combined real-time sentiment analysis with automated trade execution. Won 3rd place in the Ripple API Contest. The earliest sign that I'd eventually build something much bigger.

Get in Touch

Always happy to chat about data science, ML systems, or interesting problems.

50685071@proton.me

LinkedIn

GitHub