RLHF for frontier LLMs
3+ years producing preference and correction data for frontier models across math, coding, and computer science.
Machine Learning & Computer Vision Engineer
Three years training frontier models from the inside. Now I build the agents and workflows that put them to work.
3+ years producing preference and correction data for frontier models across math, coding, and computer science.
Each agent owns one stage: sourcing, signals, or enrichment. Retrieval grounds every step, writes are validated before they land, and a person approves anything that goes out.
The OGR capstone. Four pose models benchmarked across four platforms, scored with per video holdout rather than window level splits: F1 of 0.97 on people the model had already seen, 0.89 on people it had not.
Sole developer of a data heavy online store. The catalog lives in Neon PostgreSQL and refreshes daily from distributor feeds, with REST APIs behind both the admin dashboard and the storefront.
The chatbot above. It answers only from a written record of my work, and a second model checks every answer before it reaches your screen. Nothing personal is hardcoded, so one profile file and an API key turn it into someone else's.
This assistant answers only from Elliot's real background. It drafts an answer, then a second model checks it against the source documents before showing it, so it declines rather than guesses.
An eval is a fixed set of test questions with known answers, rerun after every change to see whether the chatbot got better or worse. On the latest run, the right source is found for all 28. Each row below is a change the numbers decided.
| Finding | Measure | Before | After |
|---|---|---|---|
|
01 A test caught a rule the chatbot was dropping adopted |
Said it was confidential n=4 |
75% | 100% |
| Search test n=24 |
79% | 79% | |
|
02 Saving tokens cost facts rejected |
Found a right source n=24 |
100% | 83% |
| Share of right sources found n=24 |
90% | 60% | |
|
03 The grader was wrong, not the chatbot corrected |
Refused when it should n=11 |
64% | 100% |
| Right fact in the answer n=19 |
90% | 100% |
Method: a free offline search test runs on every change, and an answer test calls
the live model. Scores are recomputed from stored answers. Limits: facts match as
phrases and refusals by a marker list, so scores are triage and transcripts are read.
In 03, before is the original grade of the same stored answers. Run files are in
evals/results/.
Every line joins a role or project to a skill it used. Scroll to walk the career in order and watch the skills build up.
How to read it
Bands sit along the top in time order. Lines run down from each band to the skills that work used.
2022–2026
University of Mississippi, Data Science emphasis
Machine Learning, Computer Vision, Data Science, C++, Python, SQL, pandas, Java
2023–2026
RLHF training data for frontier LLMs
RLHF, LLMs, Model Eval, Model Experimentation, Prompt Engineering, Python, Statistics
2024–2025
iOS client for an insurance platform
Swift, JavaScript, REST APIs, SwiftUI, Django, MongoDB, React
2024–2026
Sole developer, catalog and storefront
PostgreSQL, REST APIs, Data Engineering, SQL, JavaScript, Stripe
2026
Real time gesture recognition under a latency budget
PyTorch, OpenCV, MediaPipe, FPS Benchmarking, Pareto Analysis, Pose Estimation, TCNs, CNNs, Model Training, Data Augmentation, Edge Deploy, Image Classification, Python, NumPy
2026
Autonomous M&A deal sourcing, three agents
LangChain, RAG, Embeddings, LLMs, Prompt Engineering, Python, Git, Docker, FastAPI, Streamlit
2026
Grounded resume chatbot, open source template
FastAPI, LangChain, RAG, Embeddings, Docker, JavaScript, Python, Data Visualization, Google Cloud
All of it
Every engagement at once. Hover a line to see which skill it is.
The same skills, grouped by discipline. Pick a discipline to pull it out of the web.
Computer Science graduate (B.S., Data Science emphasis, May 2026) with hands on experience in applied machine learning and computer vision. Built end to end gesture recognition systems benchmarked across CPU and GPU platforms under a latency budget, and trained temporal CNN models on custom video datasets. More than 3 years producing RLHF training data for frontier LLMs. Seeking roles in ML engineering, AI, or computer vision.
Python, a frontier LLM, Streamlit, SQLite, Gmail MCP, Git.
Engagement details, pipeline composition, and deal figures are confidential and are not published here. Happy to discuss the architecture and engineering decisions.
Python, PyTorch, OpenCV, MediaPipe, TCNs.
Python, FastAPI, LangChain, FAISS, Groq, Docker, Google Cloud Run, JavaScript, three.js.
SwiftUI, Django, MongoDB, React.
Leave a message and an email address or phone number, and Elliot will get back to you. It goes straight to Elliot's inbox. The assistant never sees it, and nothing is stored on this site.