Research ethics are not à la carte.
In June, a NYU Langone research team demonstrated that, according to their research, frontier AI models (Chat-GPT, Claude, Gemini) beat out OpenEvidence on a series of medical benchmarks. In response, OpenEvidence lashed out, accusing the researchers of releasing biased research. The AI giant claimed that the Langone team had failed to disclose that OpenEvidence had previously refused an API request from NYU.
Only a week later, OpenEvidence backed a consortium study which demonstrated it outperformed frontier models on a set of questions fielded from clinicians. What they did not disclose: OpenEvidence executives supplied the data for researchers, co-developed the data collection plan, and recruited and paid survey respondents.
New models, new accuracies
Now, OpenEvidence is pushing forward with more closed book research. On September 3rd, OpenEvidence launched a new family of models: Osler (answers in about five seconds), Sackett (thirty seconds), Snow (its deepest production model, five minutes), and Darwin, a “research preview” restricted to institutional partners.
CEO Daniel Nadler explained that “every model in the family is held to the same standard of clinical accuracy. What varies is time: how long a model thinks, and how deep it searches.” We’ve been asked to take OpenEvidence’s accuracy claims in good faith.
And the accompanying accuracy scores are striking. Darwin posts a perfect 100% on MedQA. It leads the field on MedXpertQA (72.8%), HealthBench Professional (82.7%), and NOHARM (87.2%) – ahead, OpenEvidence says, of Claude and Gemini. The company also claims more American physicians now use its platform than every competing AI tool combined. No citation accompanies it.
Closed loop
None accompanies the benchmarks either: these are OpenEvidence's own tests, scored by OpenEvidence, published by OpenEvidence.
There is a particular irony in naming your reasoning model after David Sackett – credited with founding evidence-based medicine – and then supplying, as your only evidence, a scorecard you wrote and graded yourself.
And, recall for a moment OpenEvidence’s “proof” of the Nature study’s impropriety: the researchers had scored frontier models against known benchmarks developed by OpenAI. HealthBench Professional is a known benchmark developed by OpenAI.
Who's watching?
On the same day, the FDA's Technology-Enabled Meaningful Patient Outcomes for Digital Health Devices (TEMPO) pilot picked up its fourth participant of the summer. TEMPO lets selected generative-AI-enabled devices reach patients before formal marketing authorization, on the theory that real-world deployment will generate the evidence premarket review used to demand up front. The plan is to scale to forty participants across four clinical use areas.
The claim isn't that OpenEvidence is a participant in TEMPO, or that they even encompass the same regulatory category. But, hold the expansion of TEMPO side-by-side with OpenEvidence’s self-benchmarking. In the same week, both institutions decided that old proof was redundant, and that new proof could be assembled on real patients or by a company grading its own test. In a word, both decided independent verification is optional.
CHARGE doesn't assert that Darwin's benchmarks are fake or that TEMPO is reckless by design. Clinical accuracy, however, is not clinical utility, and a claim is not evidence because a vendor itself is confident.
OpenEvidence is totally confident in its ability to manipulate its own research narrative. The question is the same as always: who, besides the vendor, checked any of this?
References
[1] https://www.businesswire.com/news/home/20260903878517/en/Introducing-the-OpenEvidence-Model-Family
[3] https://www.govinfosecurity.com/fda-pilot-gives-ai-health-tools-real-world-test-bed-a-32772



