Perspectives

Why most clinical LLM triage evaluations are measuring the wrong thing

Written by Dr Annabelle Painter | Aug 27, 2026, 1:46:52 PM

In short:

  • Triage is about managing risk, not getting the diagnosis right. Safe triage means recognising red flags, managing uncertainty and directing patients to the right level of care- not necessarily ranking the eventual diagnosis first.

  • Benchmarks ≠ real world performance. Exam scores, diagnostic accuracy and clinician comparisons capture only part of what determines whether an LLM will be safe and effect in real-world triage tasks.

  • “Doctor-level performance” treats a distribution as a fixed target. Clinicians vary substantially in their assessments and decisions, particularly under the uncertainty inherent in triage.

  • Both under-triage and over-triage matter. Evaluations need to measure the trade-off between missed escalation and unnecessary healthcare utilisation, rather than treating under-triage as the only safety failure.

  • LLMs introduce different failure modes from clinicians. Fluent, confident language can disguise poorly grounded reasoning and influence patient or clinician behaviour in ways conventional accuracy measures miss.

  • Clinical readiness needs system-level, real-world evaluation. Effective testing should examine escalation, uncertainty calibration, behavioural effects, longitudinal performance, human override and post-deployment outcomes.

Clinical LLM evaluation is increasingly focused on medical exams, diagnostic reasoning benchmarks, and comparisons with clinicians. These methods provide useful information about reasoning ability, but they often measure only a narrow aspect of what determines real-world clinical impact.

Evaluating triage requires assessment beyond controlled, exam-like conditions. Triage takes place among real patients with incomplete information, inside a workflow, and under time pressure and uncertainty. A model can perform well on the test and still behave unsafely in practice, because the test does not measure the factors that make triage safe.

This piece sets out the assumptions built into current triage evaluations, the problem with each one, and what a more realistic evaluation would measure instead.

The clinician isn't a fixed reference point

Many evaluations assume that clinician performance represents a stable reference standard. In reality, clinical assessments are often highly variable. Present the same patient to multiple clinicians and there may be substantial differences in differential diagnoses, risk assessments, investigation plans, and disposition decisions. This variability is particularly pronounced in triage, where information is incomplete, uncertainty is high, and multiple management approaches may be clinically reasonable.

This matters because benchmarking against clinicians can create the impression that clinical performance is a fixed target, when in reality it is often a distribution. Understanding where a model sits within that distribution may be more informative than simply measuring whether it matches a particular reference answer.

Getting the diagnosis right isn't the same as triaging safely

Another common assumption is that diagnostic accuracy is a good proxy for triage quality.

Triage tools are often scored on whether the final diagnosis appears on their differential list and how near the top it sits, with the implicit assumption that higher is better.

This is misleading for two reasons. Firstly, triage is not primarily a diagnostic task. It is a risk stratification and escalation task. Its core function is not to identify the most likely condition, but to determine the most appropriate next step given uncertainty.

Secondly, differential diagnosis lists are inherently probability-weighted. Common conditions should dominate early reasoning not because rare conditions are ignored, but because they are, by definition, more likely explanations of the presenting complaint. Rare but serious conditions are managed through explicit exclusion strategies, red-flag recognition, and escalation rules, rather than by appearing at the top of a ranked list.

For example, a sudden-onset headache is far more likely to be a primary headache disorder such as migraine than a subarachnoid haemorrhage. As such, migraine should sit near the top of a differential diagnosis list in a case of a true subarachnoid haemorrhage. A subarachnoid haemorrhage should also be actively considered and ruled out where clinically indicated, even if it is far less probable. Safe triage does not require elevating a correct final diagnosis to the top of the ranked differential; it requires ensuring it is not missed when red flags are present.

This mirrors everyday practice. In primary care, the focus is risk management, not diagnostic closure: ruling out the serious, safety-netting the uncertain, and directing the patient to the appropriate level of care. That decision depends on far more than diagnosis generation — patient context, co-morbidity, communication, healthcare utilisation, and whether the patient actually acts on the advice they are given all shape whether triage is safe.

A large share of presentations never reach a definitive diagnosis at first contact: patients arrive with missing, conflicting, or ambiguous information, and the pathway begins with observation, referral, or investigation rather than resolution. Triage therefore operates upstream of diagnostic certainty, in exactly that space of uncertainty where the goal is safe navigation rather than diagnostic closure.

In practice, this means a system can produce a clinically appropriate triage decision while still having a non-exhaustive or imperfect differential. Conversely, a highly accurate ranked differential is not sufficient to guarantee safe triage if it fails to prioritise risk or escalation.

Evaluating triage through the lens of differential diagnosis accuracy therefore conflates two fundamentally different clinical objectives.

Under-triage and Over-triage are both unsafe

Both under-triage and over-triage create safety risks.

Under-triage can delay care and cause direct patient harm. Over-triage increases healthcare utilisation, contributes to overcrowding, burdens clinicians, and can reduce access for patients with greater need. Real-world triage involves balancing these competing risks.

This matters for evaluation because many benchmarks implicitly treat only under-triage as a safety risk. A model tuned to eliminate under-triage can look impressively "safe" while generating system-level harm elsewhere. A quality evaluation must weigh both risks, not reward the elimination of one at the expense of the other.

AI errors aren't comparable to human errors

Current evaluations often assume AI errors are comparable to human errors. They are not.

Clinicians operate within professional, organisational, and accountability structures. Their uncertainty is often visible through behaviours such as escalation, consultation, investigation, or safety-netting.

LLMs generate outputs through pattern prediction. They can produce fluent and convincing explanations even when poorly grounded, making errors harder to detect and potentially more persuasive. LLMs can describe uncertainty without truly operating under it. Their apparent confidence is often a feature of language generation rather than a reliable estimate of risk.

This matters in triage because tone, framing, and apparent certainty shape how both patients and clinicians perceive risk, urgency, and how they act on it. Even when the content is broadly correct, reassuring phrasing can delay escalation, leading to behaviour-shaping error. Conversely, LLMs are sensitive to how a patient describes a problem: the same clinical situation, phrased in reassuring or minimising terms, can shift the output in ways a clinician's judgement may not.

Evaluations that compare models only to human performance may therefore overlook novel risks that emerge when AI systems are deployed at scale. Users may interpret confident language as evidence of reliability when the underlying probabilities are poorly calibrated. Uncertainty therefore needs explicit evaluation and validation rather than reliance on a model’s self-reported confidence.

Static benchmarks miss dynamic reality

Many evaluations rely on one-shot clinical questions with fixed answers. Real clinical encounters are iterative and evolve over time.

Safety-relevant behaviours often emerge only across multiple interactions:

Recent multi-turn evaluations, including systems such as Google’s AMIE tool, represent progress. However, they still tend to focus on conversational performance rather than downstream clinical and safety outcomes for individual patients and the wider system.

Dynamic evaluation is necessary, but it does not fully bridge the gap between model performance and real-world clinical safety.

Safety relies on the whole system, not just the model

Most evaluations treat the model as the unit of analysis. In practice, deployed systems also include:

Safety emerges from interactions between these components, not from model outputs alone.

A more capable model can still produce unsafe outcomes if surrounding systems fail to constrain, interpret, or override its recommendations appropriately. Clinical AI should therefore be evaluated as a system property rather than solely a model property.

What should we measure instead?

A more realistic evaluation framework would focus on system-level safety properties:

  1. Escalation behaviour under uncertainty – Does the system appropriately escalate risk when required?
  2. Uncertainty calibration – Is uncertainty communicated in a way that supports safe decisions?
  3. Behavioural influence – Does the system unintentionally persuade users toward unsafe actions?
  4. Longitudinal consistency – Does reasoning remain coherent across multiple interactions?
  5. Failure mode analysis – What types of errors occur and what situations make this most likely?
  6. Human override dynamics – When do clinicians disagree with the system, and why?
  7. Post-deployment monitoring – How does performance change across population subgroups and over time?

These measures focus on whether systems behave safely under realistic conditions rather than whether models simply produce correct answers.

Conclusion

Clinical LLMs are often evaluated as though they are exam candidates. In reality, they function as components within complex, safety-critical systems.

Benchmark performance remains useful, but it is not sufficient evidence of clinical readiness. The central challenge is not only improving reasoning ability, but ensuring that systems behave safely under uncertainty, over time, and within real healthcare workflows.

The key question is no longer whether LLMs can achieve doctor-level performance on tests. It is whether healthcare systems can deploy them safely.