SERVICES
Consulting where medicine, software, and data meet
I help healthtech teams and medical organizations move faster by assessing what's feasible, getting the data and evidence right, and building the software to back it up.
Clinical GenAI & LLM validation
Independent safety, bias, and accuracy benchmarking for medical LLM and GenAI systems, with rigorous evaluation that stands up to clinical and regulatory scrutiny. I design the test sets, run the model comparisons, and quantify performance the way peer reviewers and regulators expect.
Medical AI feasibility
Assess whether an AI idea is clinically and technically viable, and worth the investment, before you commit engineering time to it.
Health data strategy
Turn messy clinical and health data into a usable, well-governed foundation for products, models, and research.
Scientific literature review
Ground your product, claims, and roadmap in the current evidence base so decisions hold up to clinical and regulatory scrutiny.
Technical & medical writing
Clear documentation, manuscripts, and regulatory-ready writing that translate complex work for the right audience.
Software development
Build the prototypes, tools, and integrations that move a healthtech project from idea to something real.
Is One Run Enough? Reproducibility of Flagship LLMs in Biomedical Text Processing
JAMIA, Journal of the American Medical Informatics Association (2026), 33(6):1179-1184. Lead author.
Quantified run-to-run reproducibility of GPT-5.2 and Gemini 3 Flash on 250 oncology trial abstracts across temperature and reasoning settings, measuring how stable model outputs are under identical inputs.
Evaluating LLMs for Oncology Clinical Trial Text Mining
JCO Clinical Cancer Informatics (2025). Peer-reviewed: PubMed 41197109. Lead author.
Benchmarked GPT-4o, o1-preview, and GPT-5 on extracting structured information from oncology trial text, reporting F1 scores with confidence intervals for a head-to-head comparison.
Show Your Work: Verbatim Evidence Requirements for Medical LLMs
Biomedical text-processing study (2026). Lead author.
Tested whether requiring GPT-5.2, Gemini 3 Flash, and Claude Opus 4.5 to cite mechanically-verifiable quotes improves the auditability of trial-eligibility classification.
Not sure which of these you need?
Tell me where the project is stuck and I'll point you to the right starting place, even if that's not working together.
Book an intro call