Buyer Guides

How to Run a Clinical AI Reference Check

Vendor-supplied references are selected to be positive. That does not make them useless. It means asking different questions than the ones they were prepared for.

The short answer

Vendor-supplied references are chosen because they will say good things, so do not test enthusiasm. Test specifics: how long implementation actually took against the estimate, what the vendor did when something went wrong, what happened at the first model update, and what they would do differently. Then find at least one reference the vendor did not supply.

Explained at three levels

1 Plain English

Every vendor gives you happy customers. The trick is not to ask whether they like it, because you already know the answer. Ask what surprised them, what took longer than promised, and what they would change. Then ask that customer who else they know running it, and call that person too.

2 Informed buyer

The highest-value reference is not the most enthusiastic one. It is the one most similar to you: comparable size, comparable EHR and version, comparable integration model, and live long enough to have been through at least one model update and one support escalation.

3 Technical and professional detail

Ask to speak to the integration analyst and the informatics support contact, not only the clinical sponsor. The sponsor experienced the benefit. The analyst experienced the build, and the support contact experienced everything that went wrong afterward. They answer different questions and the second two are rarely offered.

Who to ask for

Request references matched on four dimensions: organization size and type, EHR and version, integration model, and time live. Ask specifically for one that has been running at least a year, because that is the minimum to have experienced a model update, a support escalation, and a renewal conversation.

Also request three roles rather than one: the clinical sponsor, the integration analyst who did the build, and whoever supports it day to day. Vendors default to offering the sponsor.

Questions that produce useful answers

Implementation

  • How long did it take, start to production, against what was originally estimated?
  • How many hours did your own team put in, and which team was it?
  • What was in scope that you did not expect, or out of scope that you assumed was in?
  • What would you do differently if you started again?

When something went wrong

  • What is the worst problem you have had with it?
  • How did the vendor respond, and how long did it take?
  • Did you escalate, and did escalation work?
  • Has your support contact changed, and did the relationship survive it?

The first question is the most valuable in the whole call. A reference who says there have been no problems either has not been live long or is not engaging, and either way the call needs redirecting.

Model updates

  • Has the model been updated since you went live?
  • Were you told in advance, and with what detail?
  • Did behavior change noticeably?
  • Did you revalidate, and how much work was it?

This set is the most diagnostic and the least frequently asked. See post-deployment validation requirements.

Adoption and outcomes

  • What proportion of eligible clinicians actually use it now, versus at launch?
  • Which specialties or departments dropped off, and did you find out why?
  • What did you measure, and what did you find?
  • Did the result match what you expected when you approved it?

Adoption decay is the quiet failure in most clinical AI categories and it does not appear in a vendor case study.

Commercial

  • Did the cost match the original model?
  • What was not in the quote?
  • How did renewal go, and did pricing move?

Finding references the vendor did not choose

Ask every supplied reference the same closing question: who else do you know running this?

It is a normal question inside professional networks, it is rarely refused, and the second-degree reference was not prepared. Regional health system associations, informatics professional groups, and EHR user communities are the other routes, and they produce a different distribution of experience than a vendor list does.

One unprepared reference is usually worth more than three prepared ones. Not because prepared references lie, but because they were selected, and selection is the whole mechanism.

Where this goes next

More in Buyer Guides