I have sat through a lot of clinical AI demos in the last two years, and almost all of them work. That is the problem. The demo is the part that is easy now. A competent developer can wire a model into a clinical-looking interface in a week, and the result will summarise a note, draft a code, and look convincing to anyone who has not watched one of these systems meet a real facility.
The capability gap is invisible at that stage. It only shows up in the part nobody demos: what the system does when the model is wrong, when the payer API times out at 2am, when a specialty's charting convention breaks an assumption nobody wrote down, or when your volume triples and the token bill arrives. Evaluating a clinical AI partner means finding ways to see that part before you sign.
Credentials are the floor, not the answer
Start with the honest version of what credentials are worth, including ours. AST is Certified in the Claude Partner Network. Our engineers hold Claude certifications, and that tells you something real: we have put people through formal training on the platform we build with, and we build with Claude, powered by Anthropic, inside clinical systems rather than treating model work as a side experiment.
It does not tell you that the system holds. No certification does — not ours, not SOC 2, not HITRUST. We have argued this before in evaluating healthcare vendors beyond certifications, and model work does not get an exemption from our own argument. A credential is a snapshot of training. The questions below are about behaviour under load, which is the only thing that decides whether a clinical AI project survives its first year.
So use credentials the way you would use a medical licence when choosing a surgeon. It filters out the people who should not be in the room. It does not pick the surgeon.
The five things that actually separate teams
1. Evaluation, with real cases and a number attached
The single sharpest question you can ask is: how do you know it works? A serious team has a held-out set of real clinical cases — messy ones, with incomplete demographics and mixed payer rules — and can tell you accuracy by category, not in aggregate. Aggregate accuracy hides the thing you care about. A system at 94% overall can be at 71% on the one specialty that drives your volume, and the average will never tell you.
Ask what happens to that suite when they change a prompt. If every change is re-run against the full set before it ships, you are talking to a team that treats model behaviour as software. If evaluation is something they did once before launch, you are the evaluation suite.
2. Grounding, so every output can be traced to its source
In a clinical setting an output that cannot be traced is an output that cannot be reviewed, and a reviewer who cannot check the work quickly will stop checking it. That is the real failure mode — not a dramatic hallucination, but a review step that quietly becomes a rubber stamp because verifying each item takes longer than redoing it.
So the requirement is mechanical: every drafted code, every summary line, every claim-affecting assertion should point back to the words that justify it. Ask to see the reviewer's screen, not the output. If checking one draft means reading the whole encounter again, the design has pushed the cost onto your staff.
3. Versioning, so you can explain last month's output
This is the one most teams have not thought about, and it is the one that hurts in an audit. A payer or a regulator asks why a claim from four months ago was coded the way it was. Answering that requires knowing which prompt version, which model version, which rule set, and which input produced it — and models get deprecated and replaced on the provider's schedule, not yours.
Ask directly: if the model you use today is retired next year, what happens to our ability to explain decisions made on it? A team that has thought this through logs the model and prompt version with every output and keeps the evaluation results for each version. A team that has not will tell you the model does not change much. It changes.
4. Designed behaviour on the failure path
Non-determinism is the part of this technology that traditional healthcare software engineering does not prepare you for. The same input can produce different output, so "we tested it" means something weaker than it used to. Good teams respond by constraining the surface: structured outputs rather than free text where it matters, validation against a deterministic rule set before anything moves, and an explicit decision about what happens when the model returns something unusable.
That last one is the tell. Ask what the system does when the model times out, returns malformed output, or produces low confidence. The answers that reassure me are boring: it queues for a human, it falls back to the existing manual path, it fails visibly. The answer that worries me is that it has not come up.
5. Cost and latency at your actual volume
Token economics do not scale the way licence fees do, and a pilot tells you very little about production. A workflow that costs a few cents per encounter is fine at 200 encounters a day and a budget line at 20,000. Ask for cost per encounter at your projected volume and what drives it — context length is usually the answer, and a team that has optimised for it will say so immediately.
Latency has the same shape. Ambient documentation can absorb a few seconds. A verification check that a front-desk clerk is waiting on cannot, and a workflow that adds eight seconds to patient check-in will be abandoned regardless of how accurate it is.
The healthcare-specific questions
Everything above applies to any serious AI build. These are the ones that only matter because there is PHI involved, and they are where generalist teams tend to fall over:
- Does your BAA cover the model provider? If PHI reaches a third-party model, the chain of agreements has to reach there too. Ask to see it, not to hear about it.
- Where do inference requests go? Region matters for data residency, and it matters more if you operate across the UAE, Saudi Arabia, the EU and the US on one platform.
- Is our data used for training? The answer should be an unambiguous no, in writing, with the configuration that enforces it.
- What is the minimum necessary set actually sent? The convenient implementation sends the whole chart. The correct one sends what the task requires.
- How is the model prevented from acting outside scope? Drafting is not submitting. The boundary should be enforced in the system, not in the prompt.
- Who signs off before anything reaches a payer or the record? If the answer involves the word "eventually", find out what triggers it.
These connect directly to HIPAA compliance architecture, and they are not satisfied by a model provider's own compliance posture. The provider's certifications cover the provider. Your workflow is yours.
| What you see | What it usually means | What to ask next |
|---|---|---|
| A polished demo on curated data | The happy path works | Run it on one of your own messy encounters, unrehearsed |
| Accuracy quoted as a single number | Category-level performance is unknown or unflattering | Break it down by specialty, payer and document type |
| "The model handles that" | No deterministic validation layer | What checks the output before it moves? |
| No answer on model deprecation | Versioning was not designed | How do we explain a decision after the model is retired? |
| An eval suite with failures in it | The team measures honestly | Which failures are you fixing next, and why those? |
How I would evaluate a partner this month
- Give them your worst data Not a clean sample. An encounter with missing demographics, a mixed payer situation, and the charting habits of your busiest specialty. Watch it live, unrehearsed.
- Ask for the eval report Accuracy by category, the held-out set, and what changed between the last two versions. A team that measures will hand this over readily.
- Sit with the reviewer workflow Time how long it takes a clinician to verify one draft. If it is slower than doing the work, the system will not be used, whatever it scores.
- Break the dependencies Have them show you the behaviour when the model is slow, the EMR is unavailable, and the payer endpoint returns nothing useful.
- Trace one output end to end From source words to drafted output to approval to destination, with the logs that prove each hop and the model and prompt version attached.
- Price it at real volume Cost per encounter at your projected throughput, and what happens to that number when context length grows.
Six exercises, most of which can be done in a week. Together they tell you more than any credential, including ours, and more than any reference call — because a vendor's references were chosen by the vendor, and an unrehearsed encounter was not.
Why we build this way
We apply the same list to ourselves, which is where Medexa came from. It is our own platform for clinical documentation and claims, and it is in pilot — I mention it here as an illustration of the design argument rather than a product pitch, because the constraints shaped it directly. Codes surface with the words that justify them, so review is fast. Every decision names the payer rule it applied, so it can be explained later. Agents move from shadow to assist only when measured agreement with reviewers supports it, rather than on a launch date.
None of that is exciting. It is what the five questions above produce when you take them seriously, and it is what we would want to see from anyone we were buying from. If you are evaluating partners for AI clinical documentation, hold us to the same list.
The firms worth working with will welcome this list, because it is how they evaluate their own work. The ones that get uncomfortable when you ask for the eval report have told you what you needed to know, and they have told you early, which is the cheapest time to find out.
Put a partner through this list
If you are choosing a team to build clinical AI, start with one bounded workflow and run the six exercises above. We are happy to be on the receiving end of them.





Comments
Comments are warming up. Live, no-sign-in discussion will appear here shortly.
Have a question now? Email info@allstartech.net.