Dec 2025 – Feb 2026
RACE
RAG-Optimized Clinical Reasoning Engine
Clinical language models hallucinate facts that can endanger patients, and the models capable of real reasoning are too large to run anywhere near the point of care.
Skills used
I take models from research notebooks to constrained hardware — and I engineer against hallucination as a first-class risk, not an afterthought.
Read the case study View codeImpact
- 5.5 GB
- Compressed model size
- <1s
- Evidence retrieval latency
- 8 GB
- VRAM deployment target met
- 0.1%
- Parameters trained via QLoRA















