When the Model Is Right and the System Is Wrong

Technical accuracy does not guarantee safe clinical implementation. Most of the difficulty arrives after the model is correct.

A diagnostic model reports an area under the curve of 0.94 on a held-out set. It is better than the comparison cohort of clinicians on the same images. It is cleared, procured, and integrated.

Two years later the outcome data shows no improvement, and in one subgroup a small deterioration. The model has not changed. It still performs at 0.94 on the held-out set.

This is not an unusual story, and it is not really a story about the model.

The performance describes conditions that no longer apply

Validation cohorts are assembled. Images are complete, correctly labelled, acquired on known equipment, and accompanied by the clinical context needed to interpret them. Ambiguous cases are frequently excluded, because a ground truth could not be established — which is to say, the hardest cases were removed from the measurement.

Deployment supplies none of that. Inputs arrive incomplete and out of order. The patient has a device that degrades the study. The prior imaging is at another trust. The history is three lines because the patient is unwell and the history was taken in a corridor.

A model does not need to be biased or brittle to behave differently here. It only needs to have been measured on a distribution that deployment does not reproduce, which is very nearly all of them.

The recommendation arrives before the judgment

The more consequential effect is on the clinician.

A recommendation that appears after a clinician has formed an impression is information: it can be weighed, and disagreement is a considered act. A recommendation that appears before — on the worklist, at the top of the record, as a pre-populated field — is an anchor. The clinician is no longer forming an impression and then comparing it; they are evaluating whether to overturn one.

Those are different cognitive tasks with different error profiles, and the second is more susceptible to automation bias: the well-documented tendency to under-detect a system’s errors when it is usually right. Usually right is the precondition for the failure, not protection against it.

The interface decision determining which of these happens is typically made by an integration team, on grounds of screen real estate, and is not part of the evaluation.

A model may perform well in isolation while producing worse outcomes once introduced into a workflow. Nothing about the model changed. The system it joined did.

Override authority is not oversight

The standard reassurance is that a clinician remains responsible and may override. As a control, this deserves more scrutiny than it usually gets.

Consider what the override rate actually tells you. If a system’s alerts are overridden 96% of the time, it is not functioning as a safeguard; it is generating work, and the clinician has developed a reflex to clear it — which will also clear the 4% that mattered. If alerts are almost never overridden, the question is whether that reflects genuine agreement or the absence of realistic opportunity to disagree: no time, no visibility of the reasoning, and an institutional expectation that deviation must be justified.

Neither extreme describes meaningful oversight. Both are commonly reported as evidence of it.

The useful measurements are unglamorous and rarely collected: the rate, the time available per decision, the proportion of overrides that were subsequently correct, and whether anyone reviews them.

The measurement gap

The model is measured on its outputs. The thing that matters is the outcome. Between the two sits a chain that is rarely instrumented end to end:

Technology → workflow → human behaviour → dependency → risk → governance → outcome.

Most evaluation stops at the first arrow. Most consequences appear after the third. An implementation can therefore be simultaneously well-validated, correctly procured, compliant, and harmful, with no individual step having been done badly.

What follows

Evaluate the pathway, not the component. The unit of assessment is the clinical process with the tool in it, compared against the same process without.

Treat the interface as a clinical decision. When and where a recommendation appears determines whether it informs or anchors. That belongs in the safety case.

Instrument the human side. Override rates, time-to-decision, and changes in what clinicians independently verify are the leading indicators. They are also the ones that degrade first and are noticed last.

Ask what stopped happening. The most reliable question in this whole area is not whether the tool is accurate. It is which independent check quietly ceased once the tool became trusted, and whether anyone decided to remove it.