Audio and data handling
- Is encounter audio retained? For how long, and can we set that to zero?
- Are transcripts retained separately from audio, and under what schedule?
- Is any of it used to train or improve models, and can that be contractually disabled?
- How is patient consent handled, and what does the vendor supply to support it?
- Where is processing performed, and which subcontractors are involved?
- What happens to everything at contract termination?
This set carries more weight in ambient documentation than in any other clinical AI category, because the raw input is an unredacted clinical conversation.
Note generation and clinician effort
- What proportion of notes are signed without substantive edit, measured how, at which customers?
- How long does the clinician spend reviewing, and against what baseline?
- Can note structure be configured by specialty, by service line, by individual?
- How does it handle interruptions, multiple speakers, family members, and mid-visit topic changes?
- What happens with a non-English encounter or an interpreter present?
- Does it distinguish patient-reported history from clinician assessment?
That last question separates products more than any accuracy metric. A note that attributes a patient statement as a clinical finding creates a downstream problem that is tedious to unwind.
EHR integration
- Does the draft write into the record, or is it produced outside it?
- In-context launch, or a separate application the clinician switches to?
- Which EHR, which version, and is EHR vendor program participation required?
- What integration work falls on our team, in hours and elapsed weeks?
- What breaks on an EHR upgrade, and who fixes it?
- Does it interact with coding, orders, or problem list, or only the note?
Specialty and setting coverage
- Which specialties is performance strongest and weakest in, in the vendor’s own data?
- Inpatient, ambulatory, procedural, telehealth: which are supported and which are adapted?
- Reference customers in our weakest-fit specialties, not only the strongest.
Ask for the weak cases specifically. Every vendor has them, and one that claims not to is answering a different question.
Pricing behavior
- Per clinician, per encounter, per site, or flat?
- What happens to unit price at enterprise scale, and is the tier structure documented?
- Are part-time clinicians and residents priced the same as full-time attendings?
- What is included in implementation, and what is billed separately?
- What exactly converts at the end of a pilot, and at what price?
See ambient AI cost models for where the unbudgeted lines usually sit.
Change management and monitoring
- How often is the underlying model updated, and are we notified in advance?
- Can we defer an update?
- What performance data do we get continuously, and can we measure independently?
- What is the support model after implementation, and who owns the relationship?
See post-deployment validation requirements.
One question worth asking last
Ask what happens to the product when clinicians stop using it. Adoption decay is the quiet failure mode in this category: strong initial uptake, real satisfaction, and then steady drift back to old habits in the specialties where the fit was weakest.
A vendor who has seen it will describe what they do about it. A vendor who has not will tell you it does not happen.