11/06/2026
BloodGPT scored 100% on Stanford’s clinical AI benchmark. The first and only in the world.
Claude Opus 4.5 got 92%. The Stanford ML team, who built the test, got 98% using an “errors memory” approach, not on a first run.
We got every single task right. On both versions of the benchmark (easy&hard). Five times each, every time reset for a first run. 3,000 tasks total. Using light models like Gemini Flash and Claude Haiku.
Why it matters: clinicians around the world have reasonable skepticism about the safety of medical AI, and they are right to. At 92% accuracy, 8 times out of 100 something goes wrong - a missed lab value, a wrong dose, a medication interaction nobody flags. And there’s no way to tell which 8. The confident answer looks identical to the correct one. We are solving that problem, with proof.