SERVICES

Consulting where medicine, software, and data meet

I help healthtech teams and medical organizations move faster by assessing what's feasible, getting the data and evidence right, and building the software to back it up.

WHAT I OFFER

How I help healthtech teams

Clinical GenAI & LLM validation

Independent safety, bias, and accuracy benchmarking for medical LLM and GenAI systems, with rigorous evaluation that stands up to clinical and regulatory scrutiny. I design the test sets, run the model comparisons, and quantify performance the way peer reviewers and regulators expect.

Medical AI feasibility

Assess whether an AI idea is clinically and technically viable, and worth the investment, before you commit engineering time to it.

Health data strategy

Turn messy clinical and health data into a usable, well-governed foundation for products, models, and research.

Scientific literature review

Ground your product, claims, and roadmap in the current evidence base so decisions hold up to clinical and regulatory scrutiny.

Technical & medical writing

Clear documentation, manuscripts, and regulatory-ready writing that translate complex work for the right audience.

Software development

Build the prototypes, tools, and integrations that move a healthtech project from idea to something real.

PUBLISHED RESEARCH

The validation work is grounded in peer-reviewed research

Safety, bias, and accuracy benchmarking for medical LLMs is the core of the validation service, and it draws on peer-reviewed studies I lead. These papers design the test sets, run model comparisons, and report performance using the methods clinical and regulatory reviewers look for.

Is One Run Enough? Reproducibility of Flagship LLMs in Biomedical Text Processing

JAMIA, Journal of the American Medical Informatics Association (2026), 33(6):1179-1184. Lead author.

Quantified run-to-run reproducibility of GPT-5.2 and Gemini 3 Flash on 250 oncology trial abstracts across temperature and reasoning settings, measuring how stable model outputs are under identical inputs.

Evaluating LLMs for Oncology Clinical Trial Text Mining

JCO Clinical Cancer Informatics (2025). Peer-reviewed: PubMed 41197109. Lead author.

Benchmarked GPT-4o, o1-preview, and GPT-5 on extracting structured information from oncology trial text, reporting F1 scores with confidence intervals for a head-to-head comparison.

Show Your Work: Verbatim Evidence Requirements for Medical LLMs

Biomedical text-processing study (2026). Lead author.

Tested whether requiring GPT-5.2, Gemini 3 Flash, and Claude Opus 4.5 to cite mechanically-verifiable quotes improves the auditability of trial-eligibility classification.

Not sure which of these you need?

Tell me where the project is stuck and I'll point you to the right starting place, even if that's not working together.

Book an intro call