The four contractual requirements
Model change notification
Before, not after. With a description of what changed and what class of behavior change to expect. With the ability to defer where the deployment is clinically sensitive, or at minimum to be told early enough to prepare.
A vendor unwilling to commit to advance notification is telling you that the product you validate is not the product you will run, and that you will not be told when it changes. That is a governance answer, not a technical one.
Independent monitoring access
The ability to measure performance yourself, on your own patients, against your own outcomes. Vendor dashboards are useful and are not a substitute, because the vendor is measuring its reference rather than your reality.
Drift detection and alerting
Both input drift, meaning the data feeding the model has changed, and performance drift, meaning the outputs have. With thresholds, and with a named recipient for the alert. See algorithmovigilance.
Stated obligation on degradation
What the vendor does when performance falls below the agreed threshold: investigation, timeline, remediation, and what happens if remediation fails. Without this, monitoring produces a finding nobody is obligated to act on.
What a monitoring plan contains
- Metrics. Technical performance, and at least one operational or clinical measure. A model can hold its statistics while the workflow around it quietly degrades.
- Baseline. Captured before go-live. Without it, later measurements have nothing to be compared against, and this is the single most common omission.
- Cadence. More frequent early, settling into a routine. Reset after any model change.
- Thresholds. The value that triggers investigation, agreed in advance rather than argued about afterward.
- Owner. Named individual. See board documentation.
- Escalation. Who is told, in what order, on what timeline.
- Withdrawal procedure. How the tool is switched off, who can do it, and what the workflow reverts to.
The withdrawal procedure is what separates a control from a report. A plan that cannot turn the thing off is a monitoring exercise with no consequence attached.
Measure use, not only performance
Two failure modes are invisible to accuracy metrics and common in practice.
Alert fatigue and override. If clinicians dismiss the output at a high rate, the model can be performing exactly as validated while delivering nothing. Override rate is worth tracking from day one, and a rising one is an early signal.
Automation bias. The opposite failure. Clinicians defer to the tool in cases where they would previously have looked harder. This is difficult to measure directly and worth watching for through case review rather than through a dashboard.
Why this decides renewals
Health systems increasingly build an AI model inventory that includes tools bought years earlier, and then go back to those vendors asking for validation population details, update history, and monitoring support.
Vendors who can answer quickly keep the business. Vendors who cannot answer at all have lost renewals over it, not because the product stopped working, but because the organization could no longer demonstrate that it had not.