← The evidence log
AI in Clinical Trials  ·  Clinical Data Management

AI in Clinical Trials: Show Your Work, or Stay Out

In June 2025, the FDA did something the industry it regulates had spent a decade refusing to do. It gave its scientists an AI.

The tool was called Elsa. It rolled out agency-wide ahead of schedule, and in an internal message reported by STAT News, Commissioner Marty Makary told staff to use it to "expedite clinical protocol review and reduce the overall time to complete scientific review."

Seven weeks later, CNN was reporting, based on interviews with current and former FDA employees, that Elsa had cited studies that don't exist. HHS pushed back on parts of that story. But the agency's own head of AI, Jeremy Walsh, didn't pretend the risk wasn't real. Elsa, he said, was "no different" from other generative AI. And generative AI can hallucinate.

Even the FDA's head of AI acknowledged it: generative AI systems can hallucinate.

Jeremy Walsh, FDA · to CNN, July 2025

Think about what that means for a second. The FDA reviews more clinical documents than any organization on the planet. It had every reason to want this to work. And within two months, it was dealing with the exact problem the clinical trials industry has been warning about for years.

So if you work in clinical research and you've been the person saying "not yet, not like this" every time an AI vendor shows up? You weren't behind. You were right.

I run AccuraTrials, and we build AI that reads protocols and builds studies. Which means I'm telling you the skeptics were right about my own category. They were. Here's why, and here's what we think actually fixes it.

The four weeks nobody outside the industry sees

Before a single patient is enrolled, someone has to turn the protocol into a working study. If you've never seen this up close, it's humbling. The protocol runs two hundred pages and change. The visit schedule in one section quietly disagrees with the assessment table in an appendix, and a human has to decide which one the database believes. The edit checks live in a spreadsheet that keeps growing. UAT starts Monday. Then an amendment lands and touches the eligibility criteria that were locked last week.

This takes weeks of skilled work, and mistakes are expensive in a way most industries can't imagine. One Tufts analysis from 2016 put the median cost of implementing a substantial Phase 3 protocol amendment at $535,000, and guess what causes many amendments: ambiguous eligibility criteria and mis-specified schedules. The exact stuff that gets misread during setup. Meanwhile, recent research estimates that a quarter or more of total trial cost goes into source-data verification. Checking, essentially, that the data really says what it claims to say.

Now hand that job to a language model and watch what happens in the room. The demo is impressive for about ten minutes. Then someone asks the question this industry always asks: where did this come from? Which page? Which sentence? And the vendor starts talking about confidence scores.

That moment, the one everyone in clinical data management has lived through, is really four separate problems wearing one trench coat. Click through them:

01It makes things up+
A language model can write an inclusion criterion that reads beautifully and exists nowhere in the protocol. In most software that's a bug you fix later. In a trial, it flows into screening decisions, into the data, into what a regulator eventually reviews. The scary part isn't that the output looks wrong. It's that it looks right.
02It can't show its source+
Clinical research runs on one question above all others: where did this come from? Decades of documentation rules exist because "trust me" has never once worked on an auditor. Most AI tools can't answer that question, and an answer without a source is unusable here, even when it happens to be correct.
03It won't give the same answer twice+
Run the same protocol through the same model on Tuesday and Thursday and you can get two different builds. Validation in this industry means proving a system behaves consistently, on the record, against a specification. "Usually gives the same answer" is not something you can write a test script for.
04Nobody signed it+
In a trial, somebody signs. A data manager signs off on the build. An investigator signs the forms. When an AI generates an artifact, whose name is on it? For most tools out there, the honest answer is nobody's. And an unsigned artifact has never once satisfied an auditor.

Here's the part the AI industry took years to understand: a smarter model fixes none of these. Cut the error rate in half and you've still got occasional fabrication, still no sources, still different answers on different runs, still no signature. The vendors kept promising better accuracy. The industry kept asking a completely different question. That's why so little has stuck. Not stubbornness. A category error.

So we gave the AI a smaller job

At AccuraTrials, we didn't try to build a model that never hallucinates. Nobody can. We changed what the model is allowed to do.

The AI gets exactly one job: read the protocol and suggest what's in it. Visit schedules, eligibility criteria, the fields the electronic case report forms (eCRFs) will collect, the timing windows. That's it. It suggests. It doesn't build.

And every single suggestion has to come with proof: the actual sentence from the protocol it was read from, tied to where it sits in the document. Not a confidence score. A quote. That rule is enforced by the system itself, not politely requested in a prompt. If the model produces something it can't tie back to real protocol text, that suggestion doesn't get shown with a little warning icon. It gets pulled aside entirely and handed to a human. We can't stop a model from inventing things. What we can do is make sure an invention has nowhere to go.

Then a person takes over. A human reviews the interpretation clause by clause, with the protocol quote sitting right next to each item, and approves it on the record. There's an audit trail of who saw what and when. The build has a name on it again. A real one.

Only after that approval does anything get built. And here's the piece that matters for anyone who's ever run validation: the AI isn't the one building it. The forms, the edit checks, the study configuration all get generated by ordinary software working from the approved, human-signed specification. Ordinary software is something this industry knows exactly how to test.

Read, prove, sign, build. The model drafts, a human approves, and regular code does the assembly. Each of those four walls above runs into a gate designed for it.

Funny thing: after Elsa's rough summer, the FDA kept going, and by 2026 Elsa 4.0 shipped with a heavy emphasis on humans verifying the AI's work. The most document-burdened regulator on earth looked at its own AI problem and landed on the same answer: let the machine draft, make the human sign. The industry's instinct was never anti-AI. It was pro-evidence. Build for that instinct, and the door that's been closed for a decade starts to open.

Why we're doing this at all

Because the cost of building studies by hand isn't just a line item. It's a tax on every drug program that exists, and it crushes the ones with the thinnest margins first: rare disease studies, small biotechs, academic research. Every point that tax comes down, a trial that didn't pencil out becomes possible. A patient population that was going to be skipped gets studied.

That's the actual prize. And it's why the industry's caution deserves respect instead of eye-rolling. Nobody in clinical research was afraid of progress. They just refused to gamble a trial on a tool that couldn't show its work.

Fair. So we built one that has to.

What we're building

My co-founder James, who builds this system, wrote a companion piece on the engineering behind it, the checks, the failures that shaped them, and five questions to ask any AI vendor: The Audit Trail Is the Product.

AccuraTrials is the auditable intelligence layer for clinical trials. The AI drafts from the protocol. Humans review and sign, with the evidence next to every clause. Ordinary code assembles the study, so the model never touches what ships, and every output traces back to a quoted source. Making AI safe for clinical trials isn't a feature we added. It's the entire company.

See it on your own protocol

Send us a protocol and walk through a build end to end: quotes, holds, and human sign-off included. Or just tell me where you think this falls apart. I answer both kinds of email.