Key Takeaways
- Speed without credibility is expensive and risky. When generation outpaces validation, the result is uncertain evidence produced faster than anyone can evaluate it.
- The failure point isn’t the algorithm. It’s the data feeding it, the study design behind it, and the documentation around it — all more demanding under AI-assisted methods, not less.
- A credibility reckoning is coming. The teams building validation into their workflows now will be the ones whose evidence holds up when it arrives.
There is a version of the AI-in-HEOR story that goes like this: AI compresses timelines, reduces costs, and puts rigorous real-world evidence within reach of teams that couldn’t previously afford to generate it. That story is true, but here is the part the field hasn’t reckoned with yet.
The Productivity Argument Has Already Won, and That’s the Problem
AI adoption in HEOR has moved from experimentation to operational reality — ISPOR ranked AI its top trend for 2026–2027. Generative AI now runs literature review workflows at major CROs. Causal machine learning tackles comparative effectiveness questions that traditional epidemiological approaches struggled to answer. LLMs assist health economic model construction.
The efficiency gains are real.
But efficiency is a means, not an end. The end is evidence that changes decisions correctly. Right now, AI produces outputs faster than the field can verify them.
The Gap Isn’t the Algorithm
Data quality is non-negotiable. An AI method applied to incomplete or miscoded patient-level data doesn’t generate uncertain evidence — it generates confidently wrong evidence. The model has no mechanism to flag missing specialist visits, undercoded diagnoses, or patients who left the dataset because they switched payers. It produces clean-looking outputs, but the flaws stay buried in the input.
Study design still requires human judgment. AI executes a study design with precision and speed. It cannot originate one. Decisions about patient population, comparator, outcome definition, follow-up window, and confounding require judgment specific to the disease area, the question, and the regulatory context. Delegate those decisions to AI — even partially — without robust expert oversight, and even well-resourced teams start making mistakes. That’s happening right now.
Documentation standards haven’t caught up. The ELEVATE-GenAI guidelines and ISPOR’s working group outputs are meaningful progress. But most AI-assisted RWE packages being produced today do not include the methodological transparency — on AI tool use, validation steps, expert review — that would allow an independent reviewer to evaluate the evidence. Regulators and payers are filling that gap with ad hoc credibility judgments. Eventually, one of those judgments will produce a visible failure.
The Reckoning Is Coming
The pattern is familiar. A method gains adoption faster than the field develops frameworks to evaluate it. Outputs are trusted by default. Then a high-profile failure surfaces — a payer reversal, a regulatory rejection, a submission built on AI-assisted evidence that doesn’t hold up under scrutiny. The field overcorrects.
Whatever form the AI-RWE reckoning takes, the damage will extend beyond the specific failure. It will cast doubt on AI-generated evidence broadly including evidence produced by teams that did the work carefully.
What the Leading Teams Are Doing Now
The organizations that will be best positioned as this conversation becomes unavoidable are building in validation as a design principle, not retrofitting it as compliance.
In practice, that means three things:
- Control the data foundation before methods are applied.
- Keep expert judgment at the study design stage, not just the review stage.
- Document AI tool use and validation in terms that anticipate reviewer scrutiny.
None of this slows evidence generation if built in from the start. All of it becomes expensive — in time, credibility, and competitive position — if it has to be rebuilt after the reckoning.
The productivity argument for AI in HEOR is made. The credibility argument still needs to be built. The teams doing both simultaneously are the ones building durable capability.
Built for the Reckoning, Not Against It
At Komodo Health®, we built Marmot™ around the failure points that put AI-generated evidence at risk: incomplete data, weak methodological control, and poor documentation.
Marmot is grounded in the Healthcare Map®, with 1T+ real-world records, 330M+ de-identified patients, 15 years of longitudinal history, and daily refreshes designed to reveal gaps other datasets miss. Because AI cannot compensate for bad data, we start by making the data harder to get wrong.
Study design remains with the HEOR team. You define the cohort, comparator, outcomes, and confounding strategy; Marmot executes the design precisely, supports controlled sensitivity analyses, and reproduces the same result months later.
Every output includes a visible audit trail of the SQL, cohort logic, and validation checks. Independent validation agents assess results against clinical coding standards and published literature.
Grounded. Explainable. Reproducible. Verified. Governed. That is the standard Marmot applies to AI-assisted RWE— the standard regulators and payers should expect.
About the Author
Ashis Das, MD, MPH, PhD is Senior Director of Evidence Intelligence at Komodo Health, where he works at the intersection of clinical medicine, AI, and real-world evidence. A physician and thought leader, he focuses on how the healthcare industry generates and applies evidence to drive better decisions across biopharma and payer stakeholders.
This piece is part of The RWE Standard, Komodo’s monthly column on HEOR and real-world evidence in the age of AI.