A federated learning hallucination reduction method for a medical large model

CN122840197APending Publication Date: 2026-09-29BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610641132.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-11
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0010]本发明的目的是针对现有医学大模型幻觉消减技术在隐私保护、知识融合与事实锚定方面存在的缺陷和不足,创造性地提出一种面向医学大模型的联邦学习幻觉消减方法,由计算机系统实现,能够有效解决现有医学大模型在分布式临床场景中面临的生成幻觉突出、多中心知识融合不足、隐私保护与性能提升难以兼顾的行业痛点,应用于医疗问答、辅助诊断等核心临床场景

Benefits of technology

[0045]本发明,与现有技术相比,具有以下优点:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840197A_ABST
    Figure CN122840197A_ABST
Patent Text Reader

Abstract

The application discloses a kind of federal learning hallucination reduction methods for medical big model, belong to artificial intelligence medical technical field.This method is based on UMLS unified medical terminology system, automatically constructs "disease-symptom-drug-examination" four-dimensional association knowledge graph.Using Fed-LoRA parameter efficient fine-tuning realizes implicit fusion of cross-agency knowledge, combined with FedRAG retrieval enhancement to complete real-time fact anchoring and evidence checking of generated content in the reasoning stage;Server side updates global model through secure aggregation protocol, and implements personalized optimization based on client feedback portrait.The application reduces hallucination from two-dimensional system of parameter memory and real-time retrieval, and the generated content can be traced back, and the original medical data does not leave the local, which significantly improves the accuracy and safety of the output of the medical big model, and is suitable for medical question answering, clinical decision assistance and other multiple scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of artificial intelligence and medical and health applications. It relates to a collaborative generation technology that enhances generation through federated learning and retrieval. This technology uses a computer system to reduce the illusion generated by large medical models while protecting the privacy of medical data. It is applicable to application scenarios such as medical question answering, clinical decision support, and diagnostic prompts. Background Technology

[0002] In recent years, large language models have been widely used in the medical field, including intelligent triage, assisted diagnosis, and medical report generation. Medical large models, with their powerful language understanding and generation capabilities, can provide reference for clinical decision-making. However, factual errors or fictitious information appearing in the generated content—a phenomenon known as "hallucination"—severely restricts their practical deployment in clinical settings.

[0003] Federated learning, as a distributed machine learning technique, allows multiple clients to collaboratively train models without sharing raw data, providing a feasible solution for privacy-preserving multi-center medical knowledge fusion. Research indicates that combining federated learning with efficient parameter fine-tuning techniques can achieve cross-institutional knowledge flow while protecting privacy. However, existing technologies still have two shortcomings: first, the federated learning process typically only aggregates model parameters, failing to effectively align and fuse knowledge scattered across various clients; second, there is a lack of structured knowledge representation and efficient retrieval methods specific to the medical field. On the other hand, significant progress has been made in medical knowledge graph construction technology. Research shows that using the UMLS (Unified Medical Language System) medical terminology dictionary can extract medical entities from unstructured data such as electronic medical records and construct structured knowledge graphs, while publicly available data such as MIMIC-III provides rich clinical data sources for knowledge graph construction. However, existing knowledge graphs are mostly used for decision support in single-machine environments and have not yet been effectively integrated with federated learning frameworks, making it difficult to achieve collaborative utilization of multi-center knowledge while protecting privacy.

[0004] Existing technical solutions primarily employ local knowledge-based calibration loss constraint LoRA fine-tuning and construct a global conflict graph using differential privacy metadata to suppress conflicting concepts between different clients. Their focus is on knowledge calibration during the training phase and server-side conflict-aware aggregation. While such solutions can alleviate the conflict between illusion and privacy to some extent, they still do not fully address the following issues:

[0005] First, the lack of a unified medical terminology space and multi-source fusion knowledge graph for medical question-and-answer and auxiliary diagnostic scenarios makes it difficult to perform interpretable retrieval of causal, treatment, and evidence paths among "diseases, symptoms, drugs, and examinations".

[0006] Second, the reasoning stage lacks a closed-loop mechanism for entity standardization, graph path recall, vector evidence fragment recall, and answer consistency verification of user queries.

[0007] Third, server-side feedback is mostly focused on the global aggregation level, lacking personalized optimization strategies for individual client knowledge gaps, retrieval biases, and biases in non-independent and identically distributed data.

[0008] Furthermore, existing large-scale medical model hallucination reduction methods are mostly based on single-machine algorithms or simple parameter aggregation, and have not formed a closed loop of computer system-level federated knowledge fusion and real-time fact verification.

[0009] Therefore, there is an urgent need for a technology that can achieve fact anchoring and traceability of content generated by medical big models while protecting data privacy, in order to reduce the illusion phenomenon of medical big models. Summary of the Invention

[0010] The purpose of this invention is to address the shortcomings and deficiencies of existing medical large-scale model illusion reduction techniques in terms of privacy protection, knowledge fusion, and fact anchoring. It creatively proposes a federated learning illusion reduction method for medical large-scale models, implemented by a computer system. This method can effectively solve the industry pain points faced by existing medical large-scale models in distributed clinical scenarios, such as prominent generative illusions, insufficient multi-center knowledge fusion, and the difficulty in balancing privacy protection and performance improvement. It can be applied to core clinical scenarios such as medical question answering and assisted diagnosis.

[0011] This method is implemented using edge computing nodes, servers, GPU acceleration units, and dedicated databases and frameworks within a computer system, with all steps completed automatically by the computer. Its innovations include: starting with the construction of a standardized medical knowledge graph based on multi-source fusion, addressing issues such as fragmented medical knowledge representation, inconsistent terminology, and insufficient accuracy in computer retrieval; using the UMLS (Unified Medical Language System) as a standardized terminology space, and employing computer natural language processing tools to extract entities and standardize their mapping from two core data sources. On one hand, the computer automatically extracts core medical entities and semantic relationships, such as diseases, symptoms, drugs, and examination indicators, from question-answer pairs in the PubMed QA medical question-and-answer corpus; on the other hand, the computer automatically extracts clinical entities, such as demographic information, diagnostic codes, test results, and medication records, from structured electronic medical record data in databases (such as MIMIC-III) and normalizes their terminology. Based on this, the computer automatically constructs a four-dimensional knowledge graph of "disease-symptom-drug-examination", which deeply integrates evidence-based question-and-answer knowledge with real clinical case knowledge to form a dedicated medical knowledge graph covering common and rare diseases. This provides an authoritative and standardized knowledge base for the factual anchoring of subsequent model-generated content, and fundamentally solves the problem of insufficient accuracy in medical knowledge retrieval and matching.

[0012] After the knowledge graph is constructed, this method uses a lightweight large language model pre-trained on medical corpus as the base model and the knowledge graph as the model's exclusive knowledge base. The computer automatically updates the model's knowledge base. At the same time, the lightweight base design is adapted to the computing power constraints of medical edge devices, laying the foundation for distributed federated training.

[0013] Compared to the existing two-stage calibration approach of "local knowledge rooted loss + global conflict graph," this invention further specifies the technical implementation details of the "dual collaborative knowledge enhancement mechanism" and "personalized feedback optimization." The dual collaborative knowledge enhancement mechanism proposed in this invention does not merely introduce knowledge constraints into the training loss; instead, it completes cross-institutional parameterized knowledge pattern learning through Fed-LoRA during the training phase, and uses FedRAG to perform entity standardization retrieval, graph path expansion, vector evidence fragment recall, and answer consistency verification on the local knowledge graph during the inference phase. This ensures that the generated results are simultaneously constrained by both parameter memory and real-time evidence. The personalized feedback optimization mechanism proposed in this invention does not merely suppress parameter updates based on global conflict weights; instead, the computer establishes a task-entity-error type feedback profile for each client. Based on different reasons such as knowledge gaps, insufficient retrieval recall, conflicts between generated content and evidence, and heterogeneous data distribution, it triggers knowledge fragment encryption push, retrieval weight adjustment, FedProx regularization strength adjustment, and LoRA aggregation weight adjustment, respectively.

[0014] The objective of this invention is achieved through the following technical solutions.

[0015] A federated learning illusion reduction method for large medical models addresses the problem of generating illusions in distributed privacy-preserving scenarios through the collaborative work of computer hardware and software. This method is executed by a computer system. The computer system adopts a client-server distributed architecture, consisting of multiple medical client edge computing nodes and a central server, with all nodes communicating via a network. The server is deployed in a data center that meets medical data security standards and is responsible for tasks such as federated parameter aggregation, error analysis, and personalized feedback optimization. Each participating medical institution deploys at least one edge computing device to perform tasks such as local model fine-tuning, knowledge graph construction, and retrieval-enhanced inference.

[0016] This method includes the following steps:

[0017] Step 1: The computer system constructs a targeted knowledge graph for medical question answering and assisted diagnosis.

[0018] The computer system, based on the UMLS medical terminology dictionary, integrates PubMed QA core knowledge and structured entities from databases (such as the MIMIC-III database) to construct a multi-source fusion medical knowledge graph. Among these:

[0019] UMLS: Unified Medical Language System.

[0020] PubMed QA is a high-quality corpus specifically built for Medical Question Answering (MedQA) tasks. It is collected from PubMed summaries and aims to advance machine language understanding and reasoning capabilities in specialized biomedical fields.

[0021] MIMIC-III: Multi-parameter Intelligent Monitoring Database is a free, open-source, public resource database for intensive care unit research.

[0022] Specifically, step 1 may include the following steps:

[0023] Step 1.1: The computer system uses the UMLS medical terminology dictionary as a standardized terminology space and employs natural language processing tools such as MetaMap to perform medical entity recognition and mapping on question-answer pairs in PubMed QA, extracting entities and their semantic relationships. Entities include diseases, symptoms, drugs, and examination indicators.

[0024] Step 1.2: The computer system extracts entities (including demographic information, diagnostic codes, test results, medication records, etc.) from the structured electronic medical record data in the database and maps them to UMLS standard terminology.

[0025] Step 1.3: The computer system constructs a four-dimensional knowledge graph of "disease-symptom-drug-examination", and integrates the question-and-answer knowledge in PubMedQA with the clinical instance knowledge in the database to form a knowledge graph covering common and rare diseases.

[0026] Step 2: The computer system constructs an initial medical language model as the base model, and uses the knowledge graph constructed in Step 1 as the model knowledge base to realize customized updates of the knowledge base.

[0027] The base model is a lightweight model pre-trained on a medical corpus to adapt to the computing power constraints of medical edge devices.

[0028] Step 3: The computer system combines Fed-LoRA (a model heterogeneity personalized federated learning framework) parameter fine-tuning with FedRAG (a federated reinforcement learning framework) retrieval enhancement generation to construct a dual knowledge enhancement mechanism.

[0029] Specifically, step 3 may include the following steps:

[0030] Step 3.1: The client uses Fed-LoRA to fine-tune local parameters.

[0031] The backbone parameters of the base model are frozen on each client, and low-rank adaptation matrices are introduced only in the attention layer for training. The client only uploads the low-rank adapter parameters to the server, and the server uses a federated averaging algorithm to weight and aggregate the LoRA parameters from different clients, thereby achieving implicit fusion of cross-institutional knowledge.

[0032] Step 3.2: Deploy the local search enhancement module on each client.

[0033] Based on the knowledge graph constructed in step 1, an interface for building and retrieving a local vector database is implemented. For example, this can be achieved using frameworks such as LangChain.

[0034] During the model generation process, user queries are matched and retrieved in real time with the local knowledge graph, and relevant knowledge fragments are recalled as the generation context to achieve fact anchoring of the generated content.

[0035] Step 3.3: Set up an implicit knowledge enhancement module and an explicit knowledge enhancement module in the computer system. The two modules interact and work together through the computer's internal bus. The implicit knowledge enhancement module enables the model to learn knowledge patterns from different institutions through parameter updates, while the explicit knowledge enhancement module forces the model to generate content based on the local authoritative knowledge base during the inference phase. The two work together to reduce illusions from the two dimensions of parameter memorization and real-time retrieval.

[0036] Step 4: The server side implements server-side federated aggregation and feedback optimization based on the federated learning framework.

[0037] Specifically, step 4 may include the following steps:

[0038] Step 4.1: The server collects the encrypted LoRA parameters uploaded by each client, performs a weighted average using a secure aggregation protocol, and updates the global model parameters.

[0039] Step 4.2: The server uses differential privacy technology to perform gradient pruning and noise injection on the parameters uploaded by the client to prevent the original data from being inferred from the parameter updates.

[0040] Step 4.3: Based on the performance metrics of each client in tasks such as medical Q&A and assisted diagnosis, the server identifies the source of error. For clients whose performance is degraded due to lack of knowledge, relevant knowledge fragments are pushed through an encrypted channel. For clients whose aggregation bias is caused by heterogeneous data distribution, the FedProx regularization strength is adjusted to alleviate the performance degradation caused by non-independent and identically distributed data.

[0041] Step 5: Deploy the trained global model to various clients and apply it to scenarios such as medical question answering and assisted diagnosis.

[0042] During the reasoning phase, the client receives user input, retrieves relevant knowledge from the knowledge graph, combines it with model parameterized memory to generate the final answer, and provides the authoritative source and tracing path of the answer as needed.

[0043] Furthermore, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the federated learning illusion reduction method for large medical models described above.

[0044] Beneficial effects

[0045] Compared with the prior art, the present invention has the following advantages:

[0046] 1. This invention organically combines implicit Fed-LoRA parameter fine-tuning implemented by computer with explicit FedRAG retrieval enhancement, ensuring the factual accuracy of generated content from two dimensions: parameter memory and real-time retrieval. Furthermore, it suppresses generated content that conflicts with evidence fragments through answer consistency verification, thereby reducing factual errors from the source.

[0047] 2. This invention constructs a targeted knowledge graph based on databases such as UMLS, PubMed QA, and MIMIC-III, covering knowledge of common and rare diseases, and improves retrieval accuracy through standardized mapping of medical terminology.

[0048] 3. This invention enables the collaborative utilization of multi-center knowledge within a federated learning framework, ensuring that the original data does not leave the local machine and meeting the requirements of medical data privacy regulations.

[0049] 4. This invention uses LoRA for efficient parameter fine-tuning, and the client only uploads a small number of low-rank adapter parameters, which significantly reduces the communication bandwidth requirements and adapts to the computing power constraints of medical edge devices.

[0050] 5. This invention establishes client feedback profiles based on tasks, entities, and error types, enabling personalized optimization for knowledge gaps, retrieval biases, and heterogeneous data distribution, thereby improving adaptability when deployed in different medical institutions.

[0051] 6. The content generated by the present invention through computer can be linked to authoritative knowledge sources, providing a traceability path, facilitating review and judgment by clinicians, and complying with medical ethics and regulatory requirements. Attached Figure Description

[0052] Figure 1 This is a flowchart illustrating the method of the present invention;

[0053] Figure 2 This is the loss curve after 150 training rounds of this invention. Detailed Implementation

[0054] To better illustrate the objectives and advantages of this invention, the following description, in conjunction with the accompanying drawings and embodiments, further clarifies the invention. It should be noted that the implementation of this invention is not limited to the following embodiments, and any modifications or alterations made to this invention will fall within the scope of protection of this invention.

[0055] The embodiments described provide a targeted knowledge graph based on UMLS, PubMed QA, and MIMIC-III constructed through a computer system. This organically combines efficient Fed-LoRA parameter fine-tuning with FedRAG retrieval-enhanced generation, forming a dual knowledge enhancement mechanism. While protecting data privacy, it anchors the factual basis of generated content from both parameter memory and real-time retrieval dimensions, reducing the illusion rate of large medical models in question-answering and assisted diagnostic scenarios.

[0056] Example 1: Application of this method in a medical question-and-answer scenario

[0057] In this embodiment, the edge computing device of hospital client A is deployed in the Internet hospital question-and-answer system to answer patients' inquiries about medications and adverse reactions for common chronic diseases. Client A's local data includes 500 anonymized hypertension-related question-and-answer samples, 1200 outpatient follow-up records, and 300 medication records; the knowledge graph pre-contains "hypertension-antihypertensive drugs-adverse reactions" question-and-answer knowledge extracted from PubMed QA, and "diagnosis code-medication record-laboratory indicator" clinical entities extracted from MIMIC-III structured fields.

[0058] The present invention provides a federated learning method for reducing hallucinations in large medical models, implemented by a computer system, comprising the following steps:

[0059] Step 1: The computer system constructs a question-and-answer knowledge graph.

[0060] Client A uses UMLS to map "hypertension," "high blood pressure," and "elevated blood pressure" to the same disease concept, "enalapril," and "ACEI" to drug and drug category concepts, and "dry cough" to the concept of adverse reaction. It then establishes a graph path of "hypertension - ACEI drugs - enalapril - dry cough adverse reaction - recommendation to evaluate and change drug category." Simultaneously, it maps local follow-up records of test indicators such as serum creatinine, serum potassium, systolic blood pressure, and diastolic blood pressure to standard test nodes.

[0061] Step 2: Local Fed-LoRA fine-tuning.

[0062] Client A freezes the backbone parameters of the medical foundation model and trains only the low-rank adaptation matrix in the attention layer. Training samples follow a "user question - standard answer - evidence path" format. For example, the question is "I have a persistent dry cough after taking enalapril, can I continue taking it?", and the standard answer includes the statement "Dry cough may be related to ACEI drugs; a doctor should assess whether to adjust the treatment plan based on blood pressure control, renal function, and blood potassium levels." The evidence path is "Enalapril - ACEI drugs - adverse reactions - dry cough." After local training is complete, client A only uploads encrypted LoRA parameters and anonymized statistics, without uploading the patient's question or follow-up records.

[0063] Step 3: Enhance inference with local FedRAG retrieval.

[0064] When a user enters "I have high blood pressure, and I've been coughing for a long time after taking enalapril. Can I stop taking the medication on my own?", client A first identifies and standardizes entities such as "high blood pressure," "enalapril," and "cough." Then, it retrieves the path "enalapril-ACEI-dry cough adverse reaction" through a graph database and retrieves semantically similar evidence fragments such as "cannot stop taking the medication on my own, and requires a doctor to assess blood pressure, kidney function, and blood potassium" through a vector database. Finally, it writes the evidence fragments into a prompt template to constrain the model to generate an answer.

[0065] Step 4: Generation and Source Tracing.

[0066] The model outputs the following question-and-answer result: "Enalapril belongs to the ACEI class of antihypertensive drugs. Persistent dry cough may be related to this class of drugs, but it is not recommended to stop taking the medication on your own. You should consult a doctor with your recent blood pressure records, kidney function and blood potassium test results so that the doctor can determine whether to adjust to another antihypertensive regimen." The system also outputs the source path: "Disease: Hypertension → Drug: Enalapril → Class: ACEI → Adverse reaction: Dry cough → Treatment: Adjust the regimen after doctor's evaluation." If the local knowledge graph does not find sufficient evidence, the system outputs a message indicating insufficient evidence, without generating a definitive conclusion on medication use.

[0067] Step 5: Personalized feedback optimization.

[0068] If client A experiences a high number of consistency verification failures for answers related to "adverse drug reactions" within a given period, the server, based on anonymous feedback vectors, determines that its knowledge graph lacks sufficient fragments related to adverse drug reactions. It then pushes standardized thesaurus lists and adverse reaction knowledge fragments for drug categories such as ACEIs, ARBs, and calcium channel blockers to client A via an encrypted channel. If the retrieval hit rate falls below a preset threshold, the graph expansion hop count is increased from 1 to 2 hops, and the reordering weight of the "drug-adverse reaction" path is increased. Therefore, this method can continuously improve the fact-anchoring capability in medical question-and-answer scenarios without aggregating raw question-and-answer data.

[0069] Example 2: Application of this method in AI-assisted diagnosis scenarios.

[0070] In this embodiment, the edge computing device of hospital client B is deployed in the emergency auxiliary diagnostic system to provide auxiliary prompts regarding the risk of chest pain patients. The method provided in this embodiment is only intended to provide clinicians with auxiliary reference information on the risk of illness and is not directly used as the basis for disease diagnosis or treatment plans.

[0071] Client B's local data includes 800 anonymized emergency chest pain cases, 2,000 structured test records, 600 ECG summary records, and 400 medication records; the knowledge graph contains medical pathways such as "chest pain - myocardial injury markers - ECG changes - acute coronary syndrome risk warning".

[0072] The present invention provides a federated learning method for reducing hallucinations in large medical models, implemented by a computer system, comprising the following steps:

[0073] Step 1: Construct a knowledge graph for computer system-assisted diagnosis.

[0074] Client B maps the patient's chief complaint of "chest pain for 2 hours," symptoms of "chest tightness and sweating," examination findings of "elevated troponin I," and ECG summary of "ST segment depression" to UMLS standard terminology, and constructs a pathway map of "chest pain - elevated troponin - ECG ST segment changes - risk of acute coronary syndrome." Simultaneously, it uses a history of hypertension, diabetes, and previous coronary artery disease as risk factor nodes, and medication records such as aspirin and nitrates as treatment-related nodes.

[0075] Step 2: Local Fed-LoRA fine-tuning.

[0076] Client B uses samples of "patient summary - risk stratification prompts - evidence path" from locally desensitized cases to train the LoRA adapter. For example, if the patient summary is "58-year-old male, presenting with 2 hours of retrosternal squeezing pain accompanied by sweating, troponin I above the reference range, and ST segment depression on ECG", the standard output does not directly provide a definitive diagnosis, but instead suggests "high risk of acute coronary syndrome, and immediate evaluation by an emergency room or cardiologist is recommended, combining Holter monitoring, troponin retesting, and vital signs assessment".

[0077] Step 3: Enhance inference with local FedRAG retrieval.

[0078] After the emergency room doctor inputs the above patient information, the system first extracts entities such as age, gender, chief complaint, symptoms, examination indicators, and electrocardiogram summary. Then, it retrieves the path "chest pain - troponin - electrocardiogram - ST segment changes - risk of acute coronary syndrome" from the knowledge graph, and retrieves evidence fragments related to "dynamic follow-up electrocardiograms, retesting of myocardial injury markers, and risk assessment combined with vital signs" from the vector database. The model must cite the above retrieved evidence during generation to avoid making unfounded low-risk judgments.

[0079] Step 4: Generate auxiliary diagnostic suggestions.

[0080] The system outputs: "Based on the input information, the patient exhibits risk signals such as chest pain, sweating, elevated troponin I, and ST-segment changes, indicating a high risk of acute coronary syndrome. This system only provides auxiliary suggestions; it is recommended that an emergency room or cardiologist conduct further evaluation immediately, combining continuous electrocardiogram, dynamic changes in troponin I, blood pressure, heart rate, and past medical history." It also outputs the source tracing path: "Symptoms: Chest pain / sweating → Examination: Elevated troponin I → Examination: ST-segment depression → Risk warning: Acute coronary syndrome → Recommendation: Dynamic follow-up and specialist evaluation."

[0081] Step 5: Personalized feedback optimization.

[0082] If client B's LoRA update shows a significant discrepancy with the distribution of general internal medicine cases compared to other institutions due to a high proportion of emergency chest pain cases, the server identifies this data distribution heterogeneity through anonymous feedback vectors. Instead of uploading or retrieving client B's original case texts, the server increases its FedProx regularization strength to reduce aggregation bias. Simultaneously, in the knowledge gap feedback, the server pushes "synonyms for chest pain-related examination indicators," "troponin unit normalization rules," and "ECG summary keyword mapping rules" to client B, and increases the weight of the "examination-disease risk" path in local retrieval. After client B updates, the system prioritizes retrieving evidence paths related to chest pain risk stratification in the next auxiliary diagnostic inference, thereby improving the system's applicability and interpretability in emergency scenarios.

[0083] It should be noted that this method constructs a dual knowledge enhancement mechanism that combines Fed-LoRA parameter fine-tuning with FedRAG retrieval enhancement to generate deep collaboration, which is one of the core differences between this method and existing technologies. This dual guarantee of illusion reduction from the two dimensions of parameter memorization and real-time retrieval solves the core deficiency of existing federated learning schemes, which can only achieve parameter aggregation but cannot achieve effective knowledge alignment and fact anchoring. Specifically, implicit knowledge enhancement uses Fed-LoRA parameter fine-tuning, freezing the backbone parameters of the base model on each client participating in federated training, and only introducing a low-rank adaptation matrix at the attention layer for local training. During training, the client only needs to upload the low-rank adapter parameters to the central server. The server uses a federated averaging algorithm to weighted aggregate the LoRA parameters from multiple clients. This achieves implicit fusion of cross-institutional clinical knowledge without the original medical data leaving the local machine, while significantly reducing the communication bandwidth requirements of federated training and adapting to the computing power and network conditions of medical edge devices. Explicit knowledge enhancement is a FedRAG retrieval enhancement module deployed locally on the client side. It builds a local vector database and standardized retrieval interface based on the knowledge graph. In each generation and inference process of the model, it performs semantic matching and accurate retrieval between the user query and the local knowledge graph in real time, recalls the corresponding authoritative medical knowledge fragments as the generation context, and forces the model to complete the content generation based on real and authoritative clinical knowledge, so as to achieve real-time fact anchoring of the generated content.

[0084] This method employs deep collaboration between implicit and explicit enhancement layers, enabling the model to learn and accumulate cross-institutional knowledge patterns during parameter training and to perform fact-checking and anchoring of generated content through real-time retrieval during inference. This forms a complementary and interconnected closed loop, systematically reducing the illusionary phenomena of large medical models from the two core sources: parameter memory bias and generated fiction. Building upon this, the method utilizes a server-side federated aggregation and feedback optimization mechanism to iteratively optimize the global model. The central server uses a secure aggregation protocol to perform weighted averaging of encrypted LoRA parameters from multiple clients, updating the global model parameters. Simultaneously, differential privacy technology is used to perform gradient pruning and noise injection on uploaded parameters, preventing privacy risks associated with inferring original medical data from parameter updates. Furthermore, based on the task performance metrics of each client, the method accurately identifies error sources. For performance degradation caused by knowledge gaps, corresponding knowledge fragments are pushed through encrypted channels. For aggregation bias caused by heterogeneous data distribution, the FedProx regularization strength is dynamically adjusted to mitigate performance degradation caused by non-independent and identically distributed medical data. Finally, the trained global model is deployed to each client. During the inference phase, after receiving user input, the client model simultaneously retrieves relevant authoritative knowledge from the knowledge graph through the RAG module, generates the final answer by combining the model's parameterized memory, and can provide the authoritative knowledge source and complete tracing path of the answer as needed.

[0085] like Figure 2 As shown in the training experiments, this invention employs a federated learning architecture to ensure that medical data remains locally. Combined with differential privacy and secure aggregation protocols, it provides dual protection for uploaded parameters and resists privacy attacks, overcoming the shortcomings of traditional solutions. Fed-LoRA achieves implicit fusion of cross-institutional knowledge, while FedRAG anchors authoritative knowledge during inference. The two work together to reduce model illusions and improve output reliability. The knowledge graph can flexibly integrate medical knowledge, and the federated framework supports dynamic adjustments by the client. In this embodiment, compared with the original base model, the accuracy of multiple-choice questions was slightly improved, the semantic similarity between answers and responses increased from 57.4% to 71.2%, and the semantic similarity between answers and responses increased by 13.8%; the question-answer relevance increased from 56.48 at the baseline to 70.93, indicating that after fine-tuning, the model can better focus on the user's core questions when handling complex medical instructions, reducing the deviation of irrelevant information; the average Zlib-Loss memory risk ratio decreased from 3989.74 to 1504.07, effectively cutting off direct physical contact with these data; the strict memory retrieval rate (characterizing the probability that the model can continuously and accurately recite 15 to 18 characters) decreased from 64.00% to 10.00%, effectively blocking the word-by-word leakage loop for specific fine-tuned corpora.

Claims

1. A federated learning illusion reduction method for large medical models, executed by a computer system. The computer system adopts a client-server distributed architecture, consisting of multiple medical client edge computing nodes and a central server. All nodes communicate through a network, and each medical institution participating in federated learning deploys at least one edge computing device. Its features are, Includes the following steps: Step 1: The computer system constructs a targeted knowledge graph for medical question answering and assisted diagnosis; Includes the following steps: Step 1.1: The computer system uses the UMLS medical terminology dictionary as a standardized term space to perform medical entity recognition and mapping on the question-answer pairs in PubMed QA, and extracts the entities and their semantic relationships, where the entities include diseases, symptoms, drugs, and examination indicators. Step 1.2: The computer system extracts entities from the structured electronic medical record data in the database and maps them to UMLS standard terminology; Step 1.3: The computer system constructs a four-dimensional knowledge graph of "disease-symptom-drug-examination", and integrates the question-and-answer knowledge in PubMed QA with the clinical instance knowledge in the database to form a knowledge graph covering common and rare diseases; Step 2: The computer system constructs an initial medical language model as the base model, and uses the knowledge graph constructed in Step 1 as the model knowledge base to realize customized updates of the knowledge base; Among them, the base model is a lightweight model pre-trained with medical corpus to adapt to the computing power constraints of medical edge devices; Step 3: The computer system combines Fed-LoRA parameter fine-tuning with FedRAG retrieval enhancement generation to construct a dual knowledge enhancement mechanism; Includes the following steps: Step 3.1: The client uses Fed-LoRA for local parameter fine-tuning; The backbone parameters of the base model are frozen on each client, and a low-rank adaptation matrix is ​​introduced only in the attention layer for training. The client only uploads the low-rank adapter parameters to the server, and the server uses a federated averaging algorithm to weight and aggregate the LoRA parameters from different clients to achieve implicit fusion of cross-institutional knowledge. Step 3.2: Deploy the local search enhancement module on each client; Based on the knowledge graph constructed in step 1, an interface for constructing and retrieving a local vector database is implemented. During the model generation process, user queries are matched and retrieved in real time with the local knowledge graph, and relevant knowledge fragments are recalled as the generation context to achieve fact anchoring of the generated content. Step 3.3: Set up an implicit knowledge enhancement module and an explicit knowledge enhancement module in the computer system. The two modules realize data interaction and collaborative work through the computer's internal bus. The implicit knowledge enhancement module enables the model to learn the knowledge patterns of different institutions through parameter updates, while the explicit knowledge enhancement module forces the model to generate content based on the local authoritative knowledge base during the inference stage. The two work together to reduce illusions from the two dimensions of parameter memory and real-time retrieval. Step 4: On the server side, server-side federated aggregation and feedback optimization are implemented based on the federated learning framework; Includes the following steps: Step 4.1: The server collects the encrypted LoRA parameters uploaded by each client, performs a weighted average using a secure aggregation protocol, and updates the global model parameters; Step 4.2: The server uses differential privacy technology to perform gradient pruning and noise injection on the parameters uploaded by the client to prevent the original data from being inferred from the parameter updates; Step 4.3: Based on the performance metrics of each client in tasks such as medical question answering and assisted diagnosis, the server identifies the source of error. For clients whose performance is degraded due to lack of knowledge, relevant knowledge fragments are pushed through an encrypted channel. For clients whose aggregation bias is caused by heterogeneous data distribution, the FedProx regularization strength is adjusted to alleviate the performance degradation caused by non-independent and identically distributed data. Step 5: Deploy the trained global model to each client; During the reasoning phase, the client receives user input, retrieves relevant knowledge from the knowledge graph, combines it with model parameterized memory to generate the final answer, and provides the authoritative source and tracing path of the answer as needed.

2. The federated learning hallucination reduction method for large medical models as described in claim 1, characterized in that, The method is applied to medical question-and-answer and clinical decision support scenarios, where the clinical decision support scenario provides clinicians with assistance in identifying and prompting risk factors for their conditions.

3. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the federated learning illusion reduction method for large medical models as described in claim 1.