The MVP Was the Easy Part: What It Really Took to Build Trustworthy Healthcare AI

by John Wollman, Head of Revenue Strategy, Komodo Health

Almost every one of our Marmot™ clients has, at some point, attempted building their own AI-based analytic solution in-house, and they all arrived at the same conclusion. It goes like this: a small team points a foundational model at a data warehouse or data lake, wires up a few agents, and within a couple of weeks, there’s a working demo, a ”minimally viable product” (MVP) that looks great and produces answers to real questions. People lean in on the Zoom demo. It looks impressive. And it raises an obvious question: if a small team can build this in a few weeks, why license a platform instead of building our own?

It’s a fair question. The honest answer is that, given enough time and skilled resources, some of these companies could likely build something that looks and feels a lot like Marmot. The effort is genuinely valuable. It proves the concept, builds excitement, and shows the shape of what generative and agentic AI can do. We know this because it is where we started. Komodo Health®’s Marmot MVP came together in about four weeks and demoed beautifully. Seeing it, I was giddy with excitement about the possibility of truly democratizing complex analytics.  

The trouble is that an MVP still has to earn the trust that healthcare requires. Trust is an imperative in healthcare, from day one. The real cost of building a trustworthy, agentic, analytic AI solution shows up long after those early weeks. It shows up in the slow, unglamorous work of turning a confident-sounding prototype into something executives, clinicians, brand teams, and even regulators can rely upon, and something compliance teams can endorse.

What Separates an MVP From a Platform

An MVP can produce answers that seem plausible but are completely wrong (been there, unfortunately, in demos for senior executives). It can polish an output enough to wow a room, but can’t guarantee that the answer is correct, that it can be reproduced, or that anyone can trace how it arrived at the conclusions.

Healthcare demands better. An answer that guides a commercial launch, identifies care gaps, or finds sites and investigators for a clinical trial has to be something an executive can stand behind, a clinician can validate, and a compliance officer can defend. That means taking a foundational LLM model (all of which are intentionally non-deterministic by design) and making it behave deterministically when the situation calls for it. It also means showing the work: leaving a visible, auditable trail behind every result so it can be inspected and proven.

Komodo bridges the gap between a demo/MVP and a trustworthy platform by adhering to five core pillars of trust. To be relied on in healthcare, an AI system needs to be:

  • Grounded in a verifiable source of truth, so every answer traces back to its source rather than a model’s best guess.
  • Explainable, so every output is transparent, and every analytic step can be inspected, not just the final conclusions.
  • Reproducible, so the same question follows the same approach and yields a consistently derived answer every time.
  • Verified, so results are checked against clinical standards and real-world knowledge before the user or stakeholder sees them.
  • Governed, so access, privacy, and policy are enforced at the platform level, every interaction is logged, and business rules can be enforced at multiple levels.

Most MVPs meet one or two of these pillars, at best. The hard work, the work that separates a prototype from a product, is achieving all five and keeping them current as models and data keep changing. Meeting each pillar implies substantive work that a do-it-yourself team owns forever. Grounded means building and maintaining ontologies that turn fragmented, inconsistent healthcare data into something a model can reason over correctly. Explainable means instrumenting every analytic step for audit, not merely displaying a response. Reproducible means forcing deterministic output from non-deterministic models and reusing validated cohort definitions and codesets. Verified means (first) running sustained, clinician-led validation cycles and building/maintaining agentic eval suites to assert validity (and fix things when the LLMs go off the rails). Governed means building role-based access, logging, and policy enforcement into the platform itself. These things take more time and resources than most people anticipate.

What Healthcare AI Relies On

The 16-Month Gap

Komodo knows the size of this gap because we lived it. The effort to take that three-week MVP to a first commercial release was around six months. Ten months later, we released a fully “leveled-up” version. That effort required a near-total rewrite, informed by input from roughly 40 live customers, including six of the top 10 biopharma companies.

Leveling up took many false starts and a lot of learning along the way. Reproducibility, taking a non-deterministic model and forcing it to follow the same approach every time, was the single hardest problem the team solved. Verification took the longest. Komodo ran what we called “Promptapaloozas,” sessions where 40 clinician colleagues fired prompts at Marmot, manually flagged what was right and wrong, and re-tested after the team fixed the underlying agents. Eventually, accuracy landed within the bands that a real product requires. Along the way, the team solved problems that sound trivial until you try them, like getting an LLM to resolve drug names reliably or understand the current date in context.

A team developing an MVP isn’t a few weeks behind Marmot; it’s more than a year behind. And the target keeps moving as the models, the data, and the end-user expectations all keep evolving. Some particularly interesting elements that support the five pillars of trust that we built (or rebuilt) for the leveled-up release include:

  1. Deterministic reproducibility. A rewrite built around a “virtual file system” that reuses validated artifacts, codesets, cohort definitions, and the methodology behind each analysis, so the same question follows the same governed path every time. 
  2. Scheduled analytics. Customers can schedule recurring analyses, weekly launch tracking, or ongoing alerts for physicians with trial-eligible patients, and trust that each run follows the identical process.
  3. Agentic verification. Before any analysis runs, five critique agents check the methodology, catch issues, and revise. Results are cross-referenced against clinical coding standards, sources like PubMed, and fourteen years of methodological best practice before a user ever sees them. 
  4. Full execution transparency. Every tool call, SQL query, and data filter is visible in a step-by-step stream, a graphical process view, and a browsable file system. Users can even branch from the point where a model went off track and recover without losing prior work.
  5. Open, tiered architecture. Marmot now separates tiers so a customer can, for example, use their own backend data infrastructure, run on Komodo’s infrastructure, a customer’s own, or a hybrid of both, and stay open to first-party, third-party, and public data.

Each of these capabilities would require a DIY team to design, validate, and maintain on its own. On a trustworthy foundation, they’re enhancements. Without one, they’re additional sources of risk.

What 16 Months Could Buy Instead

Deploying resources and cost to build something already proven in the market is a strong argument against going the DIY route. A stronger argument is the opportunity cost of the time spent.

The organizations winning in pharma AI share one trait: they spend their engineering hours on questions the business needs answered, instead of rebuilding trust every time a foundational model releases a new version (which happens quite frequently). Scarce clinical and analytical talent is worth far more analyzing markets and finding patients than it is validating an internal tool for the third time this year.

And the maintenance never ends. When a model jumps a version and quietly breaks something, or when a new capability makes last quarter’s architecture obsolete, a DIY team has to notice, diagnose, and retrofit before getting back to the actual work. Every month spent closing that gap is a month diverted from the next commercial question or the next insight the business is waiting on.

The Real Question

Marmot isn’t a chatbot or a prompt runner. It’s an enterprise platform, built for healthcare, that orchestrates and validates 14 years of analytical intelligence on top of the most comprehensive views of U.S. healthcare, augmented by a customer’s own data.

The question isn’t whether a pharma team can build an MVP. It clearly can, given enough time and resources. The real question is whether it wants to spend the next year, probably more like a year and a half, architecting trust into that MVP and maintaining it forever as the models keep shifting, or start with a proven product and get value on day one.

Many Marmot customers who start down the DIY path reach the same conclusion: immediate value beats re-creating what already works.

Build vs. Buy: Lessons from a Top 10 Pharma on their AI Journey