Beyond Accuracy: The Conditions for Equitable AI in Healthcare
AI could widen access to specialist expertise, reduce administrative burden and support earlier identification of disease. Yet it may also reproduce the unequal conditions from which healthcare data arise. This review outlines how clinicians and health-system leaders can make equity a measurable requirement of AI quality, from problem selection to post-deployment monitoring.
Why this matters
Health systems are exploring AI in clinical decision support, diagnostic imaging, risk prediction, patient communication, remote monitoring and administration. The attraction is understandable: demand is rising, workforces are constrained, diagnosis may be delayed and access to specialist care differs sharply across geography, wealth and digital connectivity. Machine-learning systems can process images, laboratory results, physiological signals and electronic health records at scale. Generative systems may summarise records, draft documentation and assist communication.
But scale is not the same as fairness. Healthcare data record not only disease but also the history of who reached services, who was tested, whose symptoms were documented, who was referred and who received treatment. An algorithm may therefore look accurate while reproducing unequal care. The central practical question is whether a defined tool, used for a particular task and population in a specified workflow, creates a fair distribution of benefit and harm.
Implementation cannot be separated from local infrastructure. Imaging protocols, sensor availability and acute stroke pathways differ between organisations. So do existing imaging and decision timelines. A model developed using one scanner vendor, magnetic resonance imaging (MRI) field strength, dataset or patient mix may not transfer reliably to another setting. Wearable data may be unavailable for patients without appropriate devices or stable connectivity. Retrospective performance is an important starting point, but it does not show that a tool will fit local workflows, provide timely support, preserve equitable access to follow-up or improve outcomes. In neurology and other specialties, this makes transferability a clinical, rather than solely technical, question.
The World Health Organization (WHO) frames equity, inclusiveness, transparency, accountability, autonomy, well-being, safety and sustainability as central principles for AI in health [1]. The review argues that these principles need operational expression. Equity should be treated as a clinical performance requirement that is designed, tested and monitored throughout the AI lifecycle.
Evidence at a glance
| Evidence item | Design | Population or setting | AI-related issue | Main reported finding | Key limitation | Relevance to equity |
|---|---|---|---|---|---|---|
| WHO guidance [1] | Guidance | AI in health | Ethics and governance | Places inclusiveness, equity, transparency and accountability among core principles | Not an intervention study | Provides a governance foundation |
| Obermeyer et al. [2] | Algorithm analysis | Population-health management | Proxy target bias | Cost-based prediction underestimated health need among Black patients | One specific use case | Demonstrates why target choice matters |
| Racial-bias review [3] | Systematic review | Healthcare AI studies | Bias mechanisms | Identified recurring data, systemic, design and deployment mechanisms | Heterogeneity limited synthesis | Shows that risk spans the lifecycle |
| Commercial-AI scoping review [4] | Scoping review | Product validation studies | Subgroup reporting | Demographics were more often reported than subgroup performance | Does not establish product effectiveness | Identifies a transparency gap |
| Mature-model analysis [5] | Open-access analysis | Demographic and geographic groups | Dataset transfer | Apparent AI advantage was less clear out of distribution | Limited detail in supplied source | Supports local evaluation |
Study-by-study clinical interpretation
The WHO guidance [1] does not provide comparative clinical outcomes for a specific algorithm. Its value is normative and organisational: it establishes that equity should sit alongside safety, accountability and transparency. For a health system, this means that a technically promising product still requires scrutiny of who can access it, who may be missed and what happens after its recommendations are generated.
Obermeyer and colleagues [2] offer a particularly important warning about labels. The population-health algorithm sought to identify people likely to benefit from extra care, using future healthcare expenditure as its prediction target. The source reports that Black patients were substantially sicker than White patients at the same predicted risk, because lower expenditure reflected unequal access and treatment. When the target was changed to a direct measure of health burden, the disparity was greatly reduced. Clinically, the lesson is direct: utilisation, spending, prior referral and documented treatment may answer operational questions, but they cannot automatically stand in for need. Their validity must be examined in the care system that produced them.
The systematic review of racial bias in healthcare AI [3] found a heterogeneous literature, limiting quantitative synthesis. Even so, it identified a consistent pattern of concern across biased datasets, historical and systemic inequities, algorithmic design and unfair deployment. This matters because it rejects a narrow view of bias as a coding error. Inequity can begin when a problem is selected around an institutional objective rather than a patient need, continue through data generation and labelling, and become clinically significant only after deployment in a resource-constrained or digitally inaccessible pathway.
The 2026 scoping review [4] focuses on commercially available medical AI products and finds that demographic information was often reported more readily than performance by demographic subgroup. This is not a minor reporting detail. A cohort can be diverse while a model still has clinically important differences in false negatives, false positives or calibration. Without disaggregated results, clinicians and procurers cannot meaningfully assess whether a system is likely to work fairly in the people it is intended to serve. The absence of subgroup reporting is an evidence gap, not reassurance.
The 2026 mature-model analysis [5] is consistent with this concern. The source reports that the strongest apparent AI advantage was often seen in populations similar to development datasets, with less clear superiority in out-of-distribution groups. It does not establish a universal rule for all medical AI. It does reinforce that internal validation is not evidence of dependable local performance where disease prevalence, equipment, coding, clinical practice or access to services differ.
Finally, reporting frameworks can improve visibility of these issues. TRIPOD+AI covers reporting of clinical prediction models using regression or machine-learning methods, including participants, outcomes, performance, uncertainty and study design [6]. CONSORT-AI and SPIRIT-AI address clinical trials of AI interventions [7]. They do not guarantee that a model is fair or effective, but they can make limitations easier to identify and challenge.
Why the studies should or should not be compared directly
Direct comparison of these sources would be inappropriate. They address different questions and use different forms of evidence. The WHO document is guidance. Obermeyer and colleagues analyse a specific population-health target. The systematic review examines a varied literature, the scoping review assesses reporting practices and the mature-model analysis considers demographic and geographic variation. Reporting guidance addresses transparency rather than effectiveness.
What links the sources is not a shared endpoint or effect estimate, but a common message: equity cannot be inferred from a headline accuracy figure. The sources support examining the provenance of data, the clinical meaning of labels, subgroup performance, transferability and the consequences of deployment. They do not identify one metric that should govern every use case. For screening, false negatives may matter most; for risk scores guiding treatment thresholds, calibration and availability of treatment may be more relevant; for triage, unequal delays and missed emergencies may be the central harms.
What this means for clinical practice
The starting point should be the care problem, not the availability of a product. Before selecting a system, teams should define the intended users, affected population, prediction target, expected action after each result and plausible benefit. A positive result needs a realistic route to confirmation, referral, treatment or monitoring. An uncertain result requires human review. A negative result should not be treated as a reason to stop evaluating concerning symptoms.
Clinicians should ask what the system predicts and how that target was defined. They should know where it was developed, whether local patients and workflows are comparable, and how errors vary by relevant characteristics such as age, sex, ethnicity, language, disability, socioeconomic position, geography and comorbidity. These questions are not peripheral to clinical utility. They determine whether the output can be interpreted safely.
Summary metrics remain useful, but they are insufficient on their own. An overall area under the receiver-operating-characteristic curve, accuracy or F1 score can obscure a higher false-negative rate in a smaller group. Calibration also requires attention: a predicted probability should correspond to observed risk in the relevant population. Poor calibration may create unequal treatment thresholds even when discrimination appears acceptable. Metrics should be linked to the absolute consequences of errors and to the capacity of downstream services.
Data provenance is central. Health systems should seek documentation of source institutions, geography, referral pathways, disease spectrum, diagnostic methods, treatment patterns and missingness. Missing information may be a signal of access barriers or clinical practice rather than random noise. Synthetic data and oversampling may help development datasets, but they do not create independent patient examples or substitute for testing in real under-represented populations.
External validation assesses whether a model works outside development-like conditions; temporal validation tests later cohorts as practice and data systems change. Silent prospective validation can provide a cautious intermediate stage: the tool runs in the background, allowing review of local data quality, calibration, subgroup performance, alert burden and reliability without changing decisions. When a system is expected to change care, prospective evaluation should assess decisions, patient outcomes, access, workload and potential distributional effects against usual care.
Human oversight needs substance. A clinician must have the time, authority and information needed to challenge an output. Teams need defined escalation routes, documentation and incident processes. Patients should have understandable information and a route to human review and redress when AI materially influences diagnosis, triage, treatment or access decisions. Accountability must remain identifiable across developer, vendor, health organisation and clinician roles.
Digital inclusion should be designed into the pathway. Remote monitoring, patient portals and automated messaging can require a reliable device, broadband, digital literacy, privacy at home, accessible interfaces and language support. The source supports preserving alternatives including in-person care, telephone contact, interpreters, professional messaging and paper or offline options where appropriate. A digital system should add a route to care, not become a gatekeeper that leaves non-users waiting longer.
What remains uncertain
The evidence supplied does not demonstrate that a particular local service will achieve equitable clinical benefit. Local validation is still needed to determine whether imaging and record data, scanner and workflow conditions, and available follow-up are comparable to the evidence base. Governance for imaging and wearable data, documentation, auditability and accountability must also be established locally. A registry that records purpose, version, data inputs, validation sites, performance, ownership, updates and review dates can support oversight, but it must be coupled with authority to act. Recalibration, increased human review, restriction, suspension or withdrawal may be necessary if performance changes or harm emerges.
Representativeness remains a major limitation. The source notes convenience data, small subgroup samples, incomplete demographics and varying definitions of race, ethnicity, deprivation and disability, with evidence concentrated in high-income academic settings. Under-represented age, ethnicity and comorbidity groups need meaningful subgroup assessment, including calibration where relevant. Safeguards should also address language, disability, connectivity and access to follow-up. Prospective research is required to show whether AI changes diagnosis, treatment, outcomes, workload and access without widening disparities. Retrospective evidence of prediction performance cannot establish these effects.
Conclusion
AI may help extend specialist support, identify missed disease, improve communication and reduce administrative burden, but these possibilities do not ensure equitable care. The reviewed evidence supports making equity a condition of clinical quality across problem definition, data governance, target selection, validation, workflow design, procurement and post-deployment monitoring. A model can be technically accurate while using biased proxies, producing unequal errors or directing patients to services they cannot access. Health systems should therefore favour targeted augmentation of established care, retain meaningful human oversight and provide non-digital alternatives. Transparent subgroup reporting, local and prospective evaluation, accountable governance and the ability to restrict or withdraw harmful systems are necessary. The appropriate standard is meaningful clinical benefit distributed fairly across the populations served.
Educational content only. It does not replace clinical assessment, current guidelines, or patient-specific professional advice.