YOUR AI PASSED VALIDATION. BUT IS IT STILL THE SAME AI? ....Model Drift, Continuing Assurance and the Governance of Change in Healthcare AI

AI validation is not permanent assurance. As models, data, patient populations, workflows and human reliance change, healthcare organisations need continuing evidence that yesterday’s validation still supports today’s use.

Share
YOUR AI PASSED VALIDATION. BUT IS IT STILL THE SAME AI? ....Model Drift, Continuing Assurance and the Governance of Change in Healthcare AI

Dr Alwin Tan, GAICD, MBBS, FRACS, EMBA (Melbourne Business School)
Senior Surgeon | Governance Leader | HealthTech Co-founder | Founder of the Institute for Systems Integrity (ISI) | Harvard Medical School—AI in Healthcare | University of Oxford—Sustainable Enterprise | Bastas Academy for Healthcare Leadership – Triple Scholar

Vidoula Uckiah, LL.M, GAICD, Accredited Mediator, AFCHSM, CHM
Healthcare Strategy & Governance Advisor | Multinational, Government and Institutional Reform Specialist | Workforce & Workplace Transformation and Mediator | Founder of Vision Consulting & Mediation | Graduated over 20 years ago as a lawyer from the University of Amsterdam and admitted to the Supreme Court of the Northern Territory.


INSTITUTE FOR SYSTEMS INTEGRITY

Executive Summary

As artificial intelligence becomes more deeply embedded in healthcare, considerable attention is rightly being given to the evidence required before an AI system enters clinical use. Validation, regulatory assessment, local evaluation and implementation assurance all contribute to determining whether a technology is sufficiently safe, effective and appropriate for its intended purpose.

The difficulty is that healthcare does not remain static after that decision has been made.

Patient populations evolve, clinical practice changes, data sources and technical infrastructure are modified, and workflows adapt. Vendors release updates and, in some systems, the AI itself may change. Clinicians may also interact differently with technology as familiarity and confidence develop.

An AI system can therefore remain operational while the conditions supporting its original validation become progressively different from those in which it is currently being used.

Emerging medical-AI literature increasingly recognises dataset shift, model drift and the need for lifecycle evaluation and post-deployment monitoring. The governance challenge is broader. Relevant change may occur in the model, data, patient population, clinical environment, workflow, intended purpose or patterns of human reliance.

This raises an important question for healthcare organisations:

When has enough changed that yesterday’s validation no longer provides sufficient assurance for today’s use?

This paper contributes to that evolving policy conversation by considering validation as conditional evidence rather than permanent status.

It proposes the concept of a Validation Half-Life for consideration: not as a mathematical expiry period, but as a way of recognising that the assurance provided by historical validation may weaken as the distance grows between the circumstances in which an AI system was validated and those in which it currently operates.

It argues that material change should not automatically require complete revalidation, but should trigger proportionate assurance according to what has changed, its significance, potential consequences and the extent to which existing evidence remains applicable.

The objective is not perpetual revalidation.

It is continuing assurance: maintaining sufficient evidence to justify continued confidence in AI as both the technology and the healthcare environment around it evolve.


This paper forms part of a series by the Institute for Systems Integrity (ISI) exploring the governance implications of artificial intelligence in healthcare.

The series contributes to an evolving policy conversation rather than prescribing a single governance or regulatory model. Earlier papers have considered trustworthy systems and shared accountability, accountability alignment where AI influences clinical decisions, the governance response when AI-related harm occurs, and decision rights where intervention, suspension and reauthorisation may be required.

This paper considers a related but less visible point in the AI lifecycle:

What happens to assurance while an AI system continues to operate and the conditions around it change?

Its focus is not the technical mathematics of model drift or the establishment of particular validation methodologies. Rather, it considers the organisational and governance implications of change, and how healthcare organisations might determine whether evidence supporting an earlier decision to deploy AI remains sufficient to support its continued use.


Validation is a moment.

Healthcare is not.

Healthcare is familiar with approval gates.

Technologies are developed, evidence is generated, performance and risks are assessed, approval is obtained and implementation follows. These processes remain essential, but they can also create an understandable impression that validation represents a completed event.

AI complicates that assumption because it operates within a clinical environment that continues to evolve after deployment.

Patient populations and disease prevalence change; new treatments and clinical guidelines emerge; documentation practices, data sources and technical infrastructure are modified; and clinical pathways are redesigned.

The AI itself may also change through retraining, software updates, configuration changes or modification of underlying models.

Research concerning dataset shift and clinical AI demonstrates that models developed on historical populations and patterns can perform differently as the environment in which they operate changes (Finlayson et al., 2021; Guo et al., 2022; Zhang et al., 2022; Williams et al., 2026).

Importantly, the model does not necessarily need to change for the relevance of its original validation to weaken.

A hospital might introduce a new laboratory assay, modify referral criteria, redesign a clinical pathway or change an upstream data source while the AI remains technically identical.

The system may continue producing plausible outputs without any obvious malfunction, even though the relationship between its inputs, outputs and clinical environment is no longer the same.

This shifts the governance question from whether an AI system was validated to whether the evidence supporting that validation continues to describe the system and environment in which care is now being delivered.


Change extends beyond the model

The language of model drift can inadvertently narrow the governance problem.

Technical literature distinguishes between forms of change such as data or covariate drift, population shift and concept drift, as well as direct changes to the model through retraining or updating. Each can affect whether previous evidence remains applicable, but healthcare organisations may also need to consider changes that occur in the way technology is incorporated into care.

Workflow and purpose can gradually shift after implementation.

A system introduced at one point in a clinical pathway may begin to be used earlier or later, or expand into another clinical environment. An output designed to inform professional judgement may gradually acquire greater decisional weight.

Individually, these adjustments may appear relatively minor, yet collectively they can produce a form of use that is materially different from what was originally evaluated.

Human reliance may change in similar ways.

Clinicians may initially approach a new AI recommendation cautiously and interrogate its output closely. If the system repeatedly appears reliable, familiarity and confidence may alter how recommendations are treated.

This is not inherently problematic, but it means that the relevant unit of governance is not simply the algorithm.

It is increasingly a human–AI system operating within a changing clinical environment.

Research concerning responsible AI deployment in safety-critical settings similarly suggests that evaluating model performance alone does not necessarily establish the safety or effectiveness of the combined human–AI system (Morey, Rayo and Woods, 2025).

Version control is therefore necessary but insufficient.

It can identify what changed within the technology, but continuing assurance also requires an understanding of what changed around it.

That requires connections between clinical governance, safety and quality, data governance, digital health, human factors, risk management, operations, procurement and frontline clinical experience.

Different parts of the organisation may each hold only part of the information needed to recognise that the assumptions supporting earlier validation have begun to change.


The Validation Half-Life

This paper proposes the Validation Half-Life as a governance concept for considering this problem.

It is not intended as a mathematical formula or a predetermined expiry period.

Rather, it reflects the proposition that the assurance provided by validation may weaken as the distance grows between the conditions under which an AI system was validated and the conditions under which it is currently being used.

Consider two systems validated at the same time.

The first continues to use the same model for the same purpose, within a comparable population and stable clinical workflow. Its relevant data sources remain consistent, performance remains within expected parameters and no significant safety signals have emerged.

The second has undergone a model update, its patient mix has changed, an upstream data source has been modified, its use has expanded into another clinical environment and clinicians appear to be interacting with its recommendations differently.

Both systems can accurately state that they were validated twelve months ago.

But that statement does not provide the same level of assurance in each case.

The distinction is important because time alone does not necessarily expire validation; material change can weaken the assumptions on which validation depended.

The relevant question is therefore not simply when validation occurred, but whether the conditions and evidence underpinning it remain sufficiently applicable to current use.

This approach also avoids creating arbitrary revalidation periods.

A relatively stable system may retain strong supporting evidence over a considerable period, while significant changes shortly after deployment could justify earlier reassessment.


From revalidation to proportionate assurance

Recognising that validation can weaken does not mean every change should trigger complete revalidation.

Healthcare changes continuously, and an approach that treats every software update, workflow adjustment or fluctuation in data as requiring a new validation exercise would quickly become unworkable.

Excessive controls could consume scarce resources, delay beneficial innovation and create governance processes so burdensome that they become difficult to sustain.

A more practical principle is that:

Material change should trigger proportionate assurance.

The response should reflect what has changed, how significant the change is, what consequences could follow and how confident the organisation remains that its existing evidence applies.

A minor interface modification may require little more than documentation and review, while a material change to a clinical model, expansion into a substantially different population or unexpected deterioration in subgroup performance may require targeted evaluation or broader reassessment.

Changes affecting high-consequence diagnostic or treatment systems would reasonably attract greater scrutiny than changes affecting lower-risk administrative applications.

This also suggests that continuing assurance should combine periodic review with event-triggered review.

Scheduled review remains valuable, but a calendar cannot recognise that a patient population changed in March, an upstream system was replaced in May or a vendor modified an underlying model in July.

Potential triggers may therefore include material model or software changes, deterioration in performance, expansion of intended use, new patient populations, changes to upstream data, unexpected subgroup outcomes, significant workflow changes, patterns of clinician override, emerging evidence or new safety signals.

Recent lifecycle evaluation frameworks similarly recognise the need for reassessment where drift, version changes or safety signals alter the conditions under which an AI system was previously evaluated (Ye et al., 2026).

A trigger should not be interpreted as evidence that the AI is unsafe.

Its purpose is to indicate that something relevant to the assumptions supporting earlier validation has changed sufficiently to warrant proportionate inquiry.


Monitoring must connect to governance

Post-deployment monitoring is increasingly recognised as an important component of healthcare AI governance and algorithmovigilance (Balendran et al., 2024; Trella et al., 2026).

But monitoring alone does not constitute continuing assurance.

An organisation may have sophisticated systems for detecting changes in performance without having clearly established what should happen when those systems identify a concern.

Effective assurance requires monitoring to connect with interpretation, decision-making and authority.

Organisations need to know:

Who reviews an emerging signal?

What thresholds require further investigation?

Who determines whether previous validation remains applicable?

What additional evidence may be required?

Who has authority to restrict or modify use where uncertainty becomes significant?

The decision, and the basis on which it was made, should also be capable of being recorded and reviewed.

This does not necessarily require healthcare organisations to create entirely new governance structures.

Existing arrangements for clinical governance, patient safety, quality assurance, digital health, risk management, procurement and incident management provide substantial foundations.

The challenge is to ensure that these arrangements are sufficiently connected to recognise when changes in technology or the environment around it have altered the basis on which confidence in an AI system rests.


Validation as a living claim

A useful shift may therefore be to understand validation not as a permanent status but as a living claim supported by continuing evidence.

The question moves from:

“Was this AI validated?”

towards:

“What evidence supports its continued use today?”

That evidence may draw on current and subgroup performance, drift monitoring, incident and safety information, clinician feedback, patterns of override, workflow assessment, version history, population characteristics and confirmation that the technology continues to be used for its intended purpose.

The intention is not to reproduce the original validation exercise indefinitely, but to maintain sufficient evidence to support the continuing proposition that reliance on the system remains reasonable within its current operating conditions.

This becomes more complex with generative and increasingly agentic technologies.

The object being governed may extend beyond a single model to include prompts, retrieval sources, connected tools, guardrails, data, workflow and human interaction, any of which may change independently.

Continuing assurance may therefore increasingly need to ask not only which model was validated, but which combination of technology, information, configuration, workflow and human interaction was actually evaluated.


Continuing use should also remain a decision

Lifecycle governance should ultimately include the possibility that an AI system no longer warrants continued use.

Healthcare organisations can become operationally dependent on technologies once they are embedded within workflows, contracts, budgets, training and clinical practice.

Over time, continuation can become the default, and the question can shift from whether a system continues to provide sufficient value and assurance to whether anything sufficiently serious has occurred to justify removing it.

Those are different standards.

An AI system does not need to fail catastrophically before its continued use should be reconsidered.

It may become less accurate, less useful, poorly aligned with contemporary clinical practice, inappropriate for the population being served, superseded by better technology or insufficiently supported by current evidence.

Continuing assurance should therefore encompass not only monitoring and reassessment but, where appropriate, restriction, replacement and retirement.

This is also relevant at board and executive level.

Boards do not need to understand the mathematics of concept drift, but they should be able to obtain assurance about:

which AI systems are operating;

the purposes and populations for which they were validated;

what has materially changed since deployment;

how those changes are detected;

who determines whether existing evidence remains sufficient; and

when continued deployment itself is reconsidered.

These are governance questions rather than purely technical ones.


Contributing to the policy conversation

International regulatory and academic thinking is increasingly moving towards lifecycle approaches to healthcare AI, including post-market surveillance, change-management processes, real-world monitoring and continuing evaluation.

This creates an opportunity to consider how these developments translate into practical governance within healthcare organisations.

In Australia, this direction is evident in the Therapeutic Goods Administration's approach to AI-enabled medical-device software.

Current TGA guidance emphasises risk management across the product lifecycle, consideration of intended populations and clinical environments, and post-market monitoring, including attention to AI-specific risks such as performance degradation and data drift (Therapeutic Goods Administration, 2025).

Internationally, regulatory approaches are similarly developing around lifecycle management and controlled modification of AI-enabled medical devices, including mechanisms for prospectively managing and evaluating certain planned changes.

As these approaches develop, greater consistency may be valuable around what constitutes material change, expectations for post-deployment assurance, communication of changes by technology providers, and the relationship between clinical consequence and the level of assurance required.

These issues become increasingly important where AI systems move between healthcare organisations and jurisdictions, or where healthcare providers depend upon technologies developed, hosted or modified outside their direct control.

This need not require a universal revalidation regime.

Different technologies, applications and risk profiles will require different approaches.

The opportunity may instead lie in developing sufficiently consistent principles to enable organisations to demonstrate why evidence generated at an earlier point remains an appropriate basis for reliance today.


Conclusion — From validation to continuing assurance

Validation remains fundamental to responsible healthcare AI, but it cannot reasonably be interpreted as permanent assurance.

AI operates within dynamic healthcare systems in which models, data, patient populations, clinical practice, workflows, infrastructure and patterns of human interaction can all change.

Those changes may alter the relationship between an AI system and the evidence that originally supported its deployment, even where no obvious technical failure has occurred.

The concept of a Validation Half-Life is offered as one way of considering this challenge.

It does not suggest that validation automatically expires with time.

Rather, it recognises that the strength of historical validation depends upon the continuing relevance of the conditions and assumptions on which it was based.

Where material change occurs, the appropriate response should be proportionate to the nature of the change, its potential consequences and the extent to which existing evidence remains applicable.

The broader shift is from validation as a completed event towards continuing assurance.

This means maintaining the capacity to recognise meaningful change, reassess assumptions when necessary and determine whether additional evidence, restriction, revalidation, replacement or retirement is warranted.

Existing healthcare governance systems provide much of the foundation for this work; the challenge is to connect them effectively around technologies whose behaviour and operating environments may evolve after deployment.

Healthcare has considerable experience in learning that safety and quality cannot rest solely on the decision that permitted something to begin.

AI presents an opportunity to apply that lesson prospectively, while governance arrangements are still developing, rather than waiting for serious harm to expose where assurance has become disconnected from reality.

The question for healthcare organisations may therefore become less whether an AI system was validated, and increasingly:

Whether there remains sufficient evidence to support confidence in its use today.


References

Balendran, A. et al. (2024) ‘Algorithmovigilance, lessons from pharmacovigilance’, npj Digital Medicine, 7, 270. doi: 10.1038/s41746-024-01237-y.

Finlayson, S.G. et al. (2021) ‘The Clinician and Dataset Shift in Artificial Intelligence’, New England Journal of Medicine, 385, pp. 283–286. doi: 10.1056/NEJMc2104626.

Guo, L.L. et al. (2022) ‘Evaluation of domain generalization and adaptation on improving model robustness to temporal dataset shift in clinical medicine’, Scientific Reports, 12, 2726. doi: 10.1038/s41598-022-06484-1.

Morey, D.A., Rayo, M.F. and Woods, D.D. (2025) ‘Empirically derived evaluation requirements for responsible deployments of AI in safety-critical settings’, npj Digital Medicine, 8, 374. doi: 10.1038/s41746-025-01784-y.

Rosenthal, J.T., Beecy, A. and Sabuncu, M.R. (2025) ‘Rethinking clinical trials for medical AI with dynamic deployments of adaptive systems’, npj Digital Medicine, 8, 252. doi: 10.1038/s41746-025-01674-3.

Sittig, D.F. and Singh, H. (2025) ‘Recommendations to Ensure Safety of AI in Real-World Clinical Care’, JAMA, 333(6), pp. 457–458. doi: 10.1001/jama.2024.24598.

Therapeutic Goods Administration (2025) Evidence requirements for software using artificial intelligence. Australian Government Department of Health, Disability and Ageing.

Trella, A.L. et al. (2026) ‘Effective monitoring of online AI decision-making algorithms in just-in-time adaptive interventions’, npj Digital Medicine, 9, 655. doi: 10.1038/s41746-026-02669-4.

Williams, E. et al. (2026) ‘Simulation of covariate and concept drift in machine learning hospital admission prediction from emergency triage’, npj Digital Medicine. doi: 10.1038/s41746-026-02956-0.

Ye, Z. et al. (2026) ‘A five-phase evaluation framework for diagnostic and predictive medical artificial intelligence’, npj Digital Medicine, 9, 678. doi: 10.1038/s41746-026-03155-7.

Zhang, A. et al. (2022) ‘Shifting machine learning for healthcare from development to deployment and from models to data’, Nature Biomedical Engineering, 6, pp. 1330–1345. doi: 10.1038/s41551-022-00898-y.


Intellectual Contribution and Attribution

The governance framing, synthesis and concepts proposed in this publication form part of the intellectual contribution of the Institute for Systems Integrity (ISI), informed by the established and emerging research, regulatory approaches and other sources acknowledged in this paper.

This includes the application and development of these ideas within the paper’s broader framework for considering validation, material change and continuing assurance in healthcare AI.

ISI encourages the use, discussion and further development of this work in research, policy and practice.

Where original concepts, frameworks or formulations developed in this publication are reproduced, applied, adapted or built upon, appropriate acknowledgement of the Institute for Systems Integrity and citation of the originating publication is respectfully requested.