Key takeaways
- Healthcare RAG should be treated as an evidence-access system with a language interface, not simply an LLM connected to a vector database.
- Authorization should constrain what evidence can be retrieved before sensitive healthcare data enters model context.
- Clinical retrieval must account for provenance, freshness, conflicting evidence, incomplete records, and synchronization state—not semantic similarity alone.
- Probabilistic reasoning and authoritative action should sit on opposite sides of a deterministic policy and approval boundary.
- Production healthcare RAG should be evaluated for authorization, patient isolation, provenance, conflict detection, abstention, synchronization awareness, and action safety in addition to answer quality.
Contents · 13 sections
- RAG helps. It does not solve healthcare AI.
- Retrieval should begin with identity, not embeddings
- A clinical record is evidence, not just text
- Healthcare AI should preserve disagreement
- Missing evidence is not negative evidence
- Infrastructure state is part of clinical provenance
- The LLM should not own the action boundary
- A production architecture
- The threat model is larger than hallucination
- A healthcare RAG maturity model
- Evaluation needs to test the system, not only the answer
- Start where the system can create value without pretending to be the doctor
- Nigeria has an opportunity to design this layer early
Retrieval-augmented generation is often presented as a straightforward improvement to large language models.
Connect the model to trusted knowledge. Retrieve information relevant to the user's question. Add that information to the model's context. Generate a better answer.
For many applications, that abstraction is useful.
Healthcare is where it starts to break.
A production healthcare system cannot ask only whether the model retrieved something relevant. It must also ask whether the person making the request was authorized to retrieve it, whether the evidence is current, where it came from, whether another record contradicts it, whether important information may simply be missing, and whether the resulting answer is allowed to become an action.
Those are not peripheral compliance concerns.
They are part of the architecture.
Nigeria is reaching a particularly important moment for this discussion.
The country's digital-health direction is increasingly centered on interoperability, connected health records, stronger data governance, and infrastructure capable of moving information between previously fragmented systems.
That creates enormous opportunities for AI.
It also raises the cost of getting the architecture wrong.
The more connected healthcare data becomes, the more powerful retrieval becomes. But the same connectivity also increases the consequences of unauthorized retrieval, incorrect patient matching, stale evidence, cross-facility inconsistencies, and AI systems treating incomplete records as complete truth.
The central argument of this paper is simple:
Production-grade healthcare RAG is not primarily a retrieval problem. It is an authorization, provenance, evidence-quality, and action-boundary problem operating under real clinical and infrastructure constraints.
RAG helps. It does not solve healthcare AI.
The case for RAG in healthcare is legitimate.
Large language models have several limitations that become especially consequential in medicine: their internal knowledge can become outdated, they can produce unsupported information, and they often provide weak provenance for their conclusions.
RAG attempts to reduce these problems by exposing a model to external evidence at inference time.
Recent research supports its potential.
A 2025 systematic review in PLOS Digital Health examined healthcare RAG research across a broad range of applications and architectures. The field is growing rapidly, but the review also identified important gaps, including the absence of standardized evaluation frameworks and limited treatment of ethical considerations.
A separate systematic review and meta-analysis published in the Journal of the American Medical Informatics Association examined RAG use across biomedical applications and considered the conditions required for clinical development.
The direction is promising.
But the right conclusion is not that connecting an LLM to trusted documents automatically creates a trustworthy clinical system.
It does not.
RAG can improve the evidence available to a model.
It cannot, by itself, determine whether that evidence should have been retrieved, whether it is sufficient, whether it conflicts with other evidence, or what authority the generated answer should carry.
Grounding is necessary.
Grounding is not governance.
Retrieval should begin with identity, not embeddings
Consider a clinician asking:
What medications has this patient previously reacted to, and is the drug I am considering contraindicated?
A conventional RAG architecture might take the question, generate an embedding, search a vector database, retrieve the most similar records, and provide them to an LLM.
That may work technically.
But the first production question should occur before any vector search:
Who is asking?
The next questions follow immediately.
What is their role?
Which facility are they operating from?
Are they currently involved in this patient's care?
What encounter are they working within?
What is the purpose of the request?
Which categories of information does that purpose entitle them to access?
A doctor, pharmacist, claims officer, receptionist, laboratory scientist, administrator, and external specialist may all interact with the same patient while legitimately requiring different information.
Their natural-language queries should not create identical retrieval scopes.
This leads to an important principle:
Authorization should constrain retrieval, not merely generation.
A weak architecture retrieves broadly and then tells the model not to reveal information the user should not see.
By then, the sensitive information has already crossed into model context.
The access-control failure happened before the model generated a single token.
A stronger architecture determines the permissible evidence space first and performs retrieval only inside that boundary.
Nigeria's privacy direction reinforces this requirement. The Nigeria Data Protection Commission has specifically emphasized privacy-preserving technologies and privacy-by-design in healthcare.
Privacy therefore cannot exist only as an instruction in the system prompt.
It has to exist in the retrieval architecture.
A clinical record is evidence, not just text
Generic RAG tutorials often treat a knowledge base as a collection of chunks.
Healthcare needs a richer model.
Suppose two retrieved notes contain nearly identical sentences.
One says that a patient reported a possible reaction to penicillin several years ago but could not remember the details.
Another documents clinician-confirmed anaphylaxis following amoxicillin during a recent encounter.
Embedding similarity may consider both highly relevant to the same query.
Clinically, they do not carry the same evidentiary weight.
Retrieval therefore needs to reason about more than semantic similarity.
A piece of clinical evidence should carry information such as:
- source
- patient
- encounter
- author
- professional role
- facility
- timestamp
- clinical status
- version
- access classification
- confidence, where applicable
- relationship to newer or superseding evidence
The system can then ask not only:
How relevant is this?
but also:
How authoritative is it?
How current is it?
Has it been superseded?
Does another source disagree with it?
Was it recorded as an observation, patient report, diagnosis, or confirmed event?
This suggests a different conceptual model for clinical retrieval.
Semantic relevance remains important.
But it should operate alongside provenance, clinical authority, freshness, and patient context.
The vector database remains useful.
It simply stops being the entire evidence model.
WHO's SMART Guidelines initiative illustrates why structured clinical knowledge matters. Its approach is designed to translate evidence-based recommendations into standards-based, machine-readable, adaptive, and testable components that digital health systems can implement consistently.
Clinical knowledge is not simply text to be embedded.
Its structure matters.
Healthcare AI should preserve disagreement
Large language models are excellent synthesizers.
That becomes a liability when synthesis removes clinically meaningful disagreement.
Imagine a patient's longitudinal record contains the following progression.
In 2024:
No known drug allergies.
In 2025:
Possible penicillin reaction. Patient uncertain.
In 2026:
Amoxicillin-associated anaphylaxis confirmed.
A model may summarize all three as:
The patient has a penicillin allergy.
For many purposes, that may be useful.
But it has removed the history of how that conclusion emerged.
More importantly, some medical records contain conflicts that have not been resolved.
A laboratory result may disagree with a previous diagnosis.
Two clinicians may document competing assessments.
A newer record may contradict an older one without explicitly replacing it.
Information received from another facility may be incomplete.
A healthcare RAG system should therefore represent disagreement as a first-class state rather than treating every retrieval set as material to be compressed into one coherent narrative.
Clinical RAG should surface disagreement before attempting to resolve it.
This is not simply a prompting technique.
The evidence layer has to preserve the relationships that allow the model—and ultimately the clinician—to see that a conflict exists.
Missing evidence is not negative evidence
There is an even more fundamental problem in fragmented health systems.
The system searches the available records and finds no history of a condition.
That does not necessarily mean the patient has no history of the condition.
Perhaps the relevant treatment happened at another facility.
Perhaps older records were never digitized.
Perhaps a facility has not synchronized its data.
Perhaps identity matching failed.
Perhaps the current system does not have access to the source containing the information.
The safe statement is:
No evidence of the condition was found in the records available to this system.
That is meaningfully different from:
The patient has no history of the condition.
This distinction becomes particularly important as Nigeria moves from fragmented facility-level systems toward greater interoperability.
During that transition, AI will sometimes reason over partial state.
The architecture must know that.
More importantly, the user must know that.
Infrastructure state is part of clinical provenance
Much of the global conversation around healthcare AI quietly assumes continuously available cloud infrastructure and synchronized records.
That assumption should not be built into a Nigerian healthcare architecture.
A useful production system should be able to distinguish at least three operating conditions:
Online and synchronized
Online but partially synchronized
Offline or operating from local state
Suppose a clinician asks for a longitudinal patient summary while a hospital is temporarily unable to reach an external health-information exchange.
The facility's local record may still be useful.
But the system should not present it as though it represents every available record across the patient's care history.
Instead, infrastructure state becomes part of the answer's provenance.
A response may need to carry information equivalent to:
Facility-local evidence available. External exchange unavailable. Last successful synchronization: 09:43. Response generated from local records only.
This is not merely an infrastructure concern.
It changes what the AI actually knows.
That makes synchronization state part of evidence quality.
For health systems operating across environments where connectivity cannot always be assumed, this distinction should be designed into the system rather than discovered during an outage.
The LLM should not own the action boundary
The next architectural question is what happens after the model produces an answer.
A healthcare AI system might reasonably:
- retrieve a hospital policy
- summarize recent encounters
- locate laboratory history
- prepare HMO documentation
- draft a discharge summary
- retrieve information from an approved care plan
These are informational or assistive functions.
The risk changes when the same system can:
- issue a prescription
- change medication
- modify a diagnosis
- approve an insurance claim
- alter an authoritative medical record
- order a procedure
- trigger another consequential workflow
At that point, generated language has crossed into authority.
That transition should never happen implicitly.
WHO's guidance on generative AI in healthcare identifies inaccurate output, automation bias, privacy, security, and governance among the risks that need to be managed around these systems. It also emphasizes clearly defined tasks, human responsibility, stakeholder involvement, and post-deployment evaluation.
A production architecture should therefore separate interpretation from execution.
The LLM may interpret evidence.
But a state-changing action should cross a deterministic boundary.
That boundary can require:
- typed intent
- policy evaluation
- role permissions
- clinical or monetary limits
- required human approval
- a recorded policy version
- an auditable authorization event
The principle is straightforward:
Let the model interpret evidence. Do not let interpretation silently become authority.
This becomes even more important as healthcare systems evolve from conversational copilots toward agents capable of acting across EMRs, claims systems, pharmacies, scheduling platforms, and patient-communication systems.
Planning and execution should not share the same trust boundary.
A production architecture
A production-grade healthcare RAG system should look less like:
Question → Vector Search → LLM → Answer
and more like:
The flow begins with the authenticated actor and their current context.
That context determines the scope of evidence the system is permitted to search.
The retrieved evidence is then normalized around properties such as:
- provenance
- freshness
- clinical authority
- conflict
- supersession
- synchronization state
Only then should that evidence move into the reasoning layer.
The output should subsequently be classified according to the authority it requires.
An informational response may be displayed directly.
Clinical decision support should preserve the professional's judgment.
A state-changing action should cross a deterministic policy and approval boundary before anything is written back into an authoritative system.
Continuous audit and observability should span the entire flow.
One implication is easy to miss:
The closer healthcare AI gets to production, the less the architecture is about the model itself.
The difficult engineering moves outward into identity, policy, interoperability, data quality, provenance, synchronization, observability, evaluation, and workflow control.
The model remains important.
It simply becomes one component inside a much larger trust system.
The threat model is larger than hallucination
Hallucination receives disproportionate attention because it is intuitive and visible.
A production threat model has to go further.
| Failure mode | What happens | Why it matters | Architectural control |
|---|---|---|---|
| Unauthorized retrieval | A legitimate user retrieves records outside their permitted care context. | The answer may be accurate while the access itself is inappropriate or unlawful. | Identity, purpose, care-context, and policy checks before retrieval. |
| Cross-patient leakage | Evidence from one patient enters another patient's context. | Correct information attached to the wrong person can become a severe privacy and clinical-safety incident. | Strong patient identity, retrieval filters, tenant isolation, and adversarial testing. |
| Prompt injection through records | Untrusted text inside notes, referrals, uploads, or external sources attempts to influence model behavior. | Clinical documents become part of the attack surface. | Treat retrieved text as data, isolate instructions, restrict tools, and enforce policy outside the LLM. |
| Stale evidence | Old policies, guidelines, records, or medication information remain highly ranked. | Correct retrieval can still produce clinically inappropriate guidance. | Versioning, freshness metadata, expiration, and supersession rules. |
| Provenance collapse | The system produces a claim that cannot be traced back to supporting evidence. | Clinicians cannot inspect or challenge the basis of the answer. | Claim-level citations, evidence lineage, and source mapping. |
| Conflict suppression | Contradictory evidence is synthesized into one confident narrative. | Clinically meaningful uncertainty disappears. | Explicit conflict detection and presentation. |
| Absence overreach | “No evidence found” becomes “the patient does not have X.” | Partial data is mistaken for complete truth. | Coverage metadata, calibrated language, and abstention rules. |
| Authority escalation | A recommendation becomes a prescription, claim approval, record update, or other state change. | Probabilistic output silently gains operational authority. | Typed actions, deterministic policy gates, and explicit approvals. |
| Privilege drift | Access remains valid after a role, facility, or care relationship changes. | Yesterday's legitimate authorization becomes today's data leak. | Short-lived authorization, revocation propagation, and policy re-evaluation. |
| Synchronization uncertainty | Local records are treated as globally current during integration or connectivity failure. | Clinicians may act on incomplete state without knowing it. | Sync-state metadata, degraded-mode behavior, and explicit disclosure. |
| Identity collision | Records from different people are incorrectly merged or associated. | Correct retrieval from the wrong patient's record can be catastrophic. | Strong identity resolution, ambiguity detection, and human confirmation. |
| Agent tool misuse | A retrieval or assistant system invokes write-capable tools outside the intended workflow. | A read-oriented assistant silently becomes an execution system. | Separate read/write credentials, tool allow-lists, policy enforcement, and audit. |
Several of these failures can happen while the language model itself behaves exactly as designed.
A grounded answer can still be unsafe.
An accurate answer can still be unauthorized.
A clinically sensible recommendation can still exceed the system's authority.
Healthcare-RAG security therefore cannot be reduced to hallucination rate.
A healthcare RAG maturity model
Not every hospital should begin with the same level of AI autonomy.
A safer deployment strategy is to treat healthcare RAG as a capability that matures through stages.
| Level | Capability | Typical use | Primary control requirement |
|---|---|---|---|
| Level 0 | Searchable records | Structured search, keyword retrieval, filters, policy lookup | Data quality, identity, access control, interoperability |
| Level 1 | Evidence retrieval | Natural-language retrieval of permitted records or policy evidence | Retrieval correctness under authorization |
| Level 2 | Evidence-grounded summarization | Patient summaries, encounter summaries, claims documentation, discharge drafting | Provenance-preserving synthesis |
| Level 3 | Clinical decision support | Contraindication alerts, guideline comparison, missing-investigation prompts | Human oversight, conflict detection, abstention, clinical evaluation |
| Level 4 | Policy-bound workflow execution | Preparing orders, claims, refill workflows, or follow-up actions for approval | Typed intent, deterministic policy, approval rules, transactional safety |
| Level 5 | Constrained autonomous operation | Narrowly bounded automated actions under mature governance | Limits, policy versioning, expiry, human override, continuous monitoring |
Level 0 — Searchable records
The organization has digital information and conventional search but no generative AI layer.
Examples include structured record search, keyword retrieval, filters, and policy lookup.
The foundation is data quality, identity, interoperability, and access control.
Level 1 — Evidence retrieval
Users can retrieve permitted healthcare information using natural-language queries.
The system finds evidence but does not synthesize consequential conclusions.
Examples might include:
Show this patient's previous HbA1c results.
or:
Retrieve the encounters associated with this HMO claim.
The primary requirement is retrieval correctness under authorization.
Level 2 — Evidence-grounded summarization
An LLM synthesizes retrieved evidence while preserving citations, uncertainty, and provenance.
Examples include longitudinal summaries, encounter summaries, claims documentation, and discharge-document drafting.
The key requirement becomes provenance-preserving synthesis.
Level 3 — Clinical decision support
The system interprets evidence and produces information capable of influencing professional judgment.
Examples include highlighting possible contraindications, comparing treatment against guidelines, or surfacing potentially missing investigations.
The system informs.
The clinician decides.
Conflict detection, abstention, clinical evaluation, and human oversight become critical.
Level 4 — Policy-bound workflow execution
The AI may prepare or initiate actions, but execution passes through deterministic controls.
Examples might include preparing an order for approval, assembling a claim for authorized submission, triggering an approved refill workflow, or scheduling follow-up based on an established care plan.
This level requires typed intent, policy evaluation, approval rules, transactional safety, and auditable authorization.
Level 5 — Constrained autonomous operation
Only narrowly bounded actions with mature evidence, policy, monitoring, and governance should reach this stage.
Autonomy should remain constrained by factors such as:
- action type
- patient context
- clinical threshold
- monetary threshold
- expiration
- policy version
- human override
- continuous monitoring
The important distinction is that maturity should not be measured by how much the model can do.
It should be measured by how safely the surrounding system can support what the model is allowed to do.
A hospital should not jump from retrieval to autonomy simply because a newer model is more capable.
The trust infrastructure has to mature too.
Evaluation needs to test the system, not only the answer
The lack of standardized evaluation remains one of the important gaps in healthcare-RAG research.
Answer accuracy matters.
Retrieval quality matters.
But those are not sufficient measures for a production clinical system.
A more complete evaluation should test several dimensions.
| Dimension | Core question | Example measure |
|---|---|---|
| Retrieval relevance | Did the system retrieve evidence that actually addresses the request? | Recall/precision against clinician-labelled evidence sets |
| Retrieval completeness | Did it omit evidence that should materially affect the answer? | Percentage of required evidence retrieved |
| Authorization correctness | Did the user receive only information they were entitled to access? | Unauthorized retrieval rate; policy-test pass rate |
| Patient isolation | Can information from one patient enter another patient's context? | Cross-patient leakage rate under adversarial testing |
| Provenance coverage | Can important claims be traced to supporting evidence? | Percentage of clinically significant claims with valid citations |
| Citation correctness | Does the cited evidence actually support the generated claim? | Citation entailment or clinician-review score |
| Freshness correctness | Did the system appropriately prefer current evidence? | Stale-source selection rate |
| Supersession handling | Can newer authoritative evidence correctly replace or contextualize older evidence? | Accuracy on supersession cases |
| Conflict detection | Does the system surface meaningful disagreement between sources? | Recall on clinician-labelled conflict cases |
| Abstention quality | Can it decline to conclude when evidence is insufficient? | Appropriate abstention rate; false-confidence rate |
| Missing-data awareness | Does it distinguish absence of evidence from evidence of absence? | Error rate on incomplete-record scenarios |
| Clinical correctness | Is the healthcare content accurate and appropriate for the task? | Expert review; guideline concordance; task-specific benchmark |
| Action-boundary correctness | Does advisory output remain advisory unless correct authorization exists? | Unauthorized action-attempt rate |
| Synchronization awareness | Does the system communicate when its evidence may be stale or incomplete? | Degraded-state disclosure accuracy |
| Prompt-injection resistance | Can retrieved content manipulate system policy, instructions, or tools? | Attack success rate across adversarial document sets |
| Latency | Can the system respond within workflow-appropriate time limits? | P50/P95 end-to-end latency |
| Human usefulness | Does the system reduce cognitive or operational workload? | Task completion time, correction rate, user satisfaction, adoption |
This creates an important distinction between model evaluation and system evaluation.
A model can score well on answer quality while the system fails authorization tests.
A retriever can achieve high recall while repeatedly selecting stale evidence.
A system can produce clinically correct answers while hiding uncertainty.
Production evaluation therefore needs failure-oriented test cases.
Do not test only:
Can the system answer this question?
Also test:
- What happens when the relevant record is missing?
- What happens when two records disagree?
- What happens when the newest record is unavailable?
- What happens when the user's permission changes during a session?
- What happens when the facility loses connectivity?
- What happens when malicious instructions appear inside a retrieved document?
- What happens when the same request is made by a clinician, receptionist, pharmacist, and claims officer?
Those cases are much closer to production reality.
Start where the system can create value without pretending to be the doctor
The strongest early healthcare-RAG deployments may not be autonomous diagnostic systems.
There are lower-risk, high-value places to begin.
Longitudinal patient summarization
Help clinicians navigate months or years of records while preserving links to the underlying encounters.
Natural-language clinical record search
Reduce the time needed to locate previous investigations, treatments, observations, and documentation.
Hospital-policy retrieval
Make SOPs, internal policies, protocols, and approved clinical guidance easier for authorized staff to find.
Claims documentation
Retrieve and organize the encounter evidence required for HMO and corporate-payer submissions.
Discharge-document drafting
Prepare evidence-grounded documentation for clinician review.
Post-care support
Retrieve approved care-plan information for reminders, follow-up, and patient engagement without independently altering clinical decisions.
These workflows exploit what RAG does well: making relevant evidence easier to find and use.
They also keep consequential medical authority with the professional and workflow responsible for it.
Capability can increase later.
Trust should increase with evidence.
Nigeria has an opportunity to design this layer early
Nigeria is developing digital-health infrastructure while production AI architecture is still evolving.
That combination creates a significant opportunity.
The country does not have to reproduce an architecture in which AI is attached to healthcare records as an afterthought.
Authorization, provenance, interoperability, privacy, synchronization, and human control can be treated as first-class requirements while the underlying digital-health ecosystem itself is developing.
The Federal Ministry of Health and Social Welfare has already identified interoperability, data protection, cybersecurity, governance, patient confidentiality, and responsible AI as important elements of the country's digital-health direction.
The Nigeria Data Protection Commission is likewise emphasizing privacy-preserving technologies and privacy-by-design around healthcare data.
Those priorities should converge.
Because once interoperable health information becomes broadly retrievable, the important question will no longer be merely whether an AI system can find the data.
It will be:
Should this actor be allowed to retrieve it?
How complete is what was retrieved?
Which evidence should be trusted?
What conflicts exist?
What does the system not know?
What is the model allowed to conclude?
And what, if anything, is the system allowed to do next?
That is the difference between connecting an LLM to healthcare data and building healthcare AI infrastructure.
RAG should therefore be understood less as an LLM feature and more as an evidence-access system with a language interface.
Once framed that way, the architecture changes.
Identity matters before retrieval.
Authorization becomes part of search.
Provenance becomes part of ranking.
Infrastructure state becomes part of confidence.
Disagreement becomes something to preserve.
And action becomes a separate trust boundary.
That is what production-grade healthcare RAG should begin to look like.
Frequently asked questions
What is retrieval-augmented generation in healthcare?
Retrieval-augmented generation, or RAG, allows an AI system to retrieve external healthcare information before generating a response. In clinical environments, that evidence might include patient records, hospital policies, drug information, clinical guidelines, or claims data.
Why is ordinary RAG architecture insufficient for healthcare?
Healthcare systems must determine not only whether information is relevant, but whether the user is authorized to retrieve it, whether the evidence is current and complete, whether sources conflict, and whether the generated response is permitted to trigger an action.
How should authorization work in healthcare RAG?
Authorization should be evaluated before retrieval. A user's identity, role, facility, care relationship, encounter context, and purpose should determine the evidence the retrieval system is allowed to search.
Can healthcare RAG operate when hospital connectivity is unreliable?
Yes, but the system should distinguish between synchronized, partially synchronized, and local-only evidence. Connectivity and synchronization state should form part of the provenance of the generated response.
Should healthcare AI be allowed to take autonomous clinical actions?
Any state-changing action should cross a separate deterministic control boundary with structured intent, policy checks, permissions, approval requirements, and an audit trail. Model-generated interpretation should not silently become operational authority.
References
- Federal Government Moves to Accelerate Digital Health Transformation with National Health Technology and Data Analytics Office. Federal Ministry of Health and Social Welfare, Nigeria (2026)
- Federal Government Reaffirms Commitment to Building an Inclusive, Digital and Resilient Health System. Federal Ministry of Health and Social Welfare, Nigeria (2026)
- NDPC Highlights Privacy-Preserving Technologies as Key Priority for Securing Health Data in Nigeria. Nigeria Data Protection Commission (2025)
- NDPC, MDCN Partner to Secure Sensitive Health Records, Address National Security Risks. Nigeria Data Protection Commission
- Retrieval-augmented generation for large language models in healthcare: A systematic review. Lameck Mbangula Amugongo, Pietro Mascheroni, Steven Brooks, Stefan Doering, Jan Seidel — PLOS Digital Health (2025)
- Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines. Journal of the American Medical Informatics Association — Oxford Academic (2025)
- Ethics and governance of artificial intelligence for health: large multi-modal models. WHO guidance. World Health Organization (2024)
- SMART Guidelines. World Health Organization


