A disease-assisted diagnosis system based on a self-learning large language model intelligent agent
By constructing a disease-aided diagnostic system based on a self-learning large language model and utilizing a closed-loop optimization mechanism to summarize and refine diagnostic experience, the system solves the problems of static knowledge and non-iterative errors in existing diagnostic systems. It achieves continuous learning and self-correction capabilities similar to those of clinicians, thereby improving diagnostic accuracy and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-13
- Publication Date
- 2026-04-03
AI Technical Summary
Existing large language models in medical diagnostic systems suffer from problems such as reliance on high-cost labeled data, difficulty in continuous learning, static knowledge, and inability to iteratively optimize errors, resulting in insufficient diagnostic accuracy and robustness.
A disease-aided diagnosis system based on a self-learning large language model is constructed, including a diagnostic agent unit, a feature optimization agent unit, and a storage agent unit. Diagnostic experience is summarized, refined, and stored through a closed-loop optimization mechanism to achieve autonomous accumulation and incremental updates.
It enables the autonomous accumulation and self-correction of diagnostic experience, significantly improving diagnostic accuracy and generalization performance, and solving the problems of static knowledge and unrepairable errors in existing systems.
Smart Images

Figure CN121506461B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical diagnostic technology, and more specifically, relates to a disease-aided diagnostic system based on a self-learning large language model intelligent agent. Background Technology
[0002] In recent years, artificial intelligence technology has made significant progress, and large language models, as one of its core technologies, have received widespread attention across various industries. Among these, the application of large language models in the medical field is becoming increasingly extensive. To continuously improve the capabilities of large language models in the medical field, various solutions have been developed, including supervised fine-tuning of basic model parameters, constructing external retrieval knowledge bases, and introducing multi-agent diagnostic systems. These methods have achieved certain results in medical application scenarios, providing effective technical approaches to enhance the application capabilities of large language models in medical settings.
[0003] Despite previous research and engineering validation, these approaches still face numerous challenges. Fine-tuning methods heavily rely on labeled data, but high-quality, structured clinical labeled data covering multiple diseases is extremely expensive. Furthermore, once the model is trained, it struggles to continuously absorb new knowledge, lacking a mechanism for continuous learning and experience accumulation. Knowledge retrieval enhancement methods are limited by retrieval accuracy and knowledge coverage. When patient symptoms are rare or ambiguous, the retrieval module easily returns irrelevant or outdated information, introducing noise and reducing diagnostic reliability. While multi-agent diagnosis can improve robustness, the inability of agents to share or iteratively optimize diagnostic experience leads to repeated errors in similar cases, indicating a lack of "learning from experience" capability. More critically, existing systems generally treat "knowledge" as a static, external resource, rather than a dynamic resource that evolves endogenously through the diagnostic process. This makes it difficult for models to evolve their diagnostic capabilities like human doctors do through repeated practice, summarizing common characteristics, identifying individual differences, and eliminating misleading information. Summary of the Invention
[0004] The main objective of this invention is to provide a disease-assisted diagnostic system based on a self-learning large language model intelligent agent, in order to overcome the shortcomings of the prior art.
[0005] This invention provides a disease-assisted diagnosis system based on a self-learning large language model intelligent agent, comprising a diagnostic intelligent agent unit, a feature optimization intelligent agent unit, and a storage intelligent agent unit. The storage intelligent agent unit includes an experience database and a large language model; the experience database contains diagnostic features for various diseases. The diagnostic intelligent agent unit receives patient information, fills the patient information into designated positions of a preset prompt word template to generate standard patient information, selects several diagnostic features corresponding to the standard patient information from the experience database, and selects the diagnostic feature most similar to the category of the standard patient information, generating and outputting corresponding auxiliary diagnostic results and diagnostic criteria. The feature optimization intelligent agent unit, based on the auxiliary diagnostic results and diagnostic criteria output by the diagnostic intelligent agent unit, uses the large language model to perform closed-loop optimization on the experience database, summarizing, filtering redundancies, and iteratively absolving the diagnostic features in the experience database.
[0006] Preferably, the feature optimization agent unit includes a feature summarization agent subunit and a feature refinement agent subunit; the feature summarization agent subunit is used to summarize the common features between different cases of the same disease based on the auxiliary diagnostic results and diagnostic criteria output by the diagnostic agent unit, using the large language model, and correspondingly generate multiple diagnostic feature texts that conform to the disease and transmit them to the feature refinement agent subunit; the feature refinement agent subunit is used to filter out redundant and erroneous feature texts from the multiple diagnostic feature texts to obtain effective feature texts for optimizing the experience database.
[0007] Preferably, the feature summarizing intelligent agent subunit summarizes the common features among different cases of the same disease, and the common features include at least a first preset number of common features related to the disease and a second preset number of rare features related to the disease.
[0008] Preferably, the feature refinement intelligent body subunit filters out redundant feature texts from multiple diagnostic feature texts, specifically including: the feature refinement intelligent body subunit calculates the semantic similarity between each pair of diagnostic feature texts in the multiple diagnostic feature texts; filters out a third preset number of diagnostic feature texts with the highest semantic similarity, or filters out diagnostic feature texts with semantic similarity greater than a set value; and removes the remaining diagnostic feature texts as redundant feature texts.
[0009] Preferably, the semantic similarity is:
[0010] ;
[0011] in, Let A represent the semantic similarity between two diagnostic feature texts; A and B are vectors of the two diagnostic feature texts, respectively. and These are the moduli of the vectors of the two diagnostic feature texts, respectively.
[0012] Preferably, the feature refinement intelligent body subunit filters out erroneous feature texts, specifically including: for each diagnostic feature text retained after filtering out redundant feature texts: the feature refinement intelligent body subunit determines whether the diagnostic accuracy of the corresponding disease increases after removing the diagnostic feature text; if it increases, the diagnostic feature text is an erroneous feature text and is removed; if it decreases, the diagnostic feature text is a correct feature text and is retained.
[0013] Preferably, the number of diagnostic agent units is multiple, and different diagnostic agent units correspond to different large language models. The system also includes a referee agent unit, which is used to determine whether the auxiliary diagnostic results output by each diagnostic agent unit are consistent. If they are consistent, the auxiliary diagnostic result is taken as the final auxiliary diagnostic result.
[0014] Preferably, when the auxiliary diagnostic results output by each diagnostic agent unit are inconsistent, the referee agent unit is further used to analyze the differences between the auxiliary diagnostic results and diagnostic basis output by each diagnostic agent unit, and output the corresponding difference analysis data to the corresponding diagnostic agent unit; the corresponding diagnostic agent unit corrects its output auxiliary diagnostic results and diagnostic basis according to the difference analysis data and then transmits them to the referee agent unit again.
[0015] Preferably, when the auxiliary diagnostic results output by each diagnostic agent unit are still inconsistent after a set number of corrections, the correction is stopped, and the referee agent unit directly outputs the latest auxiliary diagnostic results and diagnostic basis output by each diagnostic agent unit.
[0016] Preferably, the referee intelligent agent unit determines whether the auxiliary diagnostic results output by each diagnostic intelligent agent unit are consistent, specifically including: the referee intelligent agent unit determines whether the auxiliary diagnostic results output by each diagnostic intelligent agent unit are consistent according to a preset disease name mapping rule.
[0017] Compared with existing technologies, the beneficial effects of this invention are as follows: It provides a disease-assisted diagnosis system based on a self-learning large language model intelligent agent, constructing a dynamic experience evolution closed loop consisting of "diagnosis-summarization-refinement-storage". This enables the system to automatically summarize, refine, verify, and optimize diagnostic knowledge from its own diagnostic results, achieving autonomous accumulation of diagnostic experience, redundancy removal, error correction, and incremental updates. This fundamentally solves the defects of existing systems, such as "static knowledge, non-accumulative experience, and uncorrectable errors". It enables the disease-assisted diagnosis system based on a self-learning large language model intelligent agent to have continuous learning and self-correction capabilities similar to those of clinicians, significantly improving diagnostic accuracy and generalization performance. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the structure of a disease-assisted diagnosis system based on a self-learning large language model intelligent agent, provided in an embodiment of the present invention.
[0020] Figure 2 for Figure 1 The diagram shows the working process of the system.
[0021] Figure 3 This is a schematic diagram of the structure of a disease-assisted diagnosis system based on a self-learning large language model intelligent agent, provided as another embodiment of the present invention.
[0022] Figure 4 for Figure 3 The diagram shows the working process of the system. Detailed Implementation
[0023] In view of the shortcomings of the prior art, the inventors of this invention, through long-term research and extensive practice, have proposed the technical solution of this invention. The following will further explain and illustrate this technical solution, its implementation process, and its principles.
[0024] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.
[0025] Furthermore, in the description of this invention, it should be understood that the terms "upper," "lower," "inner," "outer," "horizontal," "vertical," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0026] In the description of this specification, the references to terms such as "an embodiment," "a particular embodiment," or "the embodiment" indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0027] Figure 1 This is a schematic diagram of the structure of a disease-assisted diagnosis system based on a self-learning large language model intelligent agent according to an embodiment of the present invention. Its overall working process is as follows: Figure 2 As shown.
[0028] Please see Figure 1 A disease-assisted diagnosis system based on a self-learning large language model intelligent agent includes a diagnostic intelligent agent unit, a feature optimization intelligent agent unit, and a storage intelligent agent unit. The storage intelligent agent unit contains an experience database and a large language model; the experience database contains diagnostic features for various diseases. The diagnostic intelligent agent unit receives patient information, fills the patient information into designated positions in a preset prompt word template to generate standard patient information, selects several diagnostic features corresponding to the standard patient information from the experience database, and chooses the diagnostic feature most similar to the category of the standard patient information, generating and outputting the corresponding auxiliary diagnostic results and diagnostic basis. The feature optimization intelligent agent unit, based on the auxiliary diagnostic results and diagnostic basis output by the diagnostic intelligent agent unit, uses the large language model to perform closed-loop optimization of the experience database, summarizing, filtering redundancies, and iteratively absolving diagnostic features in the experience database.
[0029] The disease-assisted diagnosis system based on a self-learning large language model intelligent agent provided by this invention constructs a dynamic experience evolution closed loop consisting of "diagnosis-summarization-refinement-storage", enabling the system to automatically summarize, refine, verify and optimize diagnostic knowledge from its own diagnostic results, and realize the autonomous accumulation, redundancy removal and error correction and incremental update of diagnostic experience.
[0030] The disease-assisted diagnosis system based on a self-learning large language model intelligent agent provided by this invention is implemented through three core modules: a diagnostic intelligent agent unit, a feature optimization intelligent agent unit, and a storage intelligent agent unit.
[0031] The diagnostic agent unit acquires patient information, utilizes a correlated large language model, and, based on diagnostic features in an established experience database, obtains and outputs auxiliary diagnostic results and corresponding diagnostic criteria for that patient information. The evaluation metric for the auxiliary diagnostic results is diagnostic accuracy, defined as the percentage of cases diagnosed with a specific disease out of all cases of that disease.
[0032] The experience database is constructed by the feature optimization agent unit during its self-training and iterative process within the closed-loop optimization framework, based on the auxiliary diagnostic results from the diagnostic agent unit. Closed-loop optimization refers to the process of summarizing and refining diagnostic experience using the associated large language model based on the auxiliary diagnostic results and their corresponding diagnostic criteria, and storing this experience in the experience database. For example, initially, the experience database is empty, and the diagnostic agent unit has no experience available for diagnosis in its initial state. During the diagnosis process, a certain number of correctly diagnosed cases are accumulated (the large language model itself has a certain degree of diagnostic accuracy, or some sample cases can be manually provided for subsequent closed-loop optimization). The subsequent closed-loop optimization process uses these correctly diagnosed cases to summarize and refine diagnostic experience and store it in the experience database. In essence, it can be viewed as follows: initially, the experience database is empty; first, some cases are diagnosed to accumulate a portion of the cases; then, some experience is obtained through optimization from these cases; this experience is used to continue diagnosis; and based on the new cases obtained from the diagnoses, the experience database is further optimized. It should be noted that the optimized experience database will differ depending on the type of large language model associated with it. Therefore, independent, dedicated experience databases should be constructed for each diagnostic agent unit associated with different large language models.
[0033] The experience database stores diagnostic features for various diseases. The diagnostic agent can select several diagnostic features from the experience database that correspond to the patient's standard information, and then choose the diagnostic feature most similar to the category of the patient's standard information. For example, if the experience database contains diagnostic features for seven diseases—appendicitis, pancreatitis, cholecystitis, diverticulitis, pneumonia, pulmonary thrombosis, and pericarditis—and the patient is diagnosed with an abdominal disease (i.e., the patient's standard information category is abdominal disease), then the diagnostic agent will retrieve the diagnostic features for appendicitis, pancreatitis, cholecystitis, and diverticulitis from the experience database.
[0034] Patient information serves as the primary information source for the diagnostic agent unit. It can be obtained by reading patient information from existing datasets or through input via an interactive interface. After obtaining the patient information, the diagnostic agent unit populates it into designated positions within a preset prompt word template. This information is then combined with preset diagnostic prompt words to construct a complete input, which is sent to the associated large language model. The diagnostic agent unit utilizes the associated large language model to generate auxiliary diagnostic results and corresponding diagnostic criteria for the patient's case. The preset prompt word template is pre-constructed and possesses a degree of scalability, allowing for expansion of its specific content within a certain scope.
[0035] Preferably, the feature optimization agent unit includes a feature summarization agent subunit and a feature refinement agent subunit. The feature summarization agent subunit, based on the auxiliary diagnostic results and diagnostic criteria output by the diagnostic agent unit, uses a large language model to summarize the common features among different cases of the same disease, generating multiple diagnostic feature texts that match the disease and transmitting them to the feature refinement agent subunit. The feature refinement agent subunit filters out redundant and erroneous feature texts from the multiple diagnostic feature texts, obtaining valid feature texts to optimize the experience database.
[0036] Further preferably, the feature summarizing intelligent agent subunit summarizes the common features among different cases of the same disease, and the common features include at least a first preset number of common features related to the disease and a second preset number of rare features related to the disease.
[0037] Common features are typically the most relevant common clinical manifestations and diagnostic outcomes for the disease; rare features refer to the most relevant specific or unique clinical manifestations and diagnostic outcomes observed in a subset of patients. The first presupposed number is, for example, 3, 4, or 5, and the second presupposed number is, for example, 3, 4, or 5.
[0038] The Feature Summarization Intelligent Subunit, based on the auxiliary diagnostic results and diagnostic criteria generated by the Diagnostic Intelligent Subunit, utilizes the associated large language model to summarize key diagnostic features of the same disease during the diagnostic process, thus identifying commonalities and individual differences among different cases of the same disease. In practice, to ensure the Feature Summarization Intelligent Subunit has sufficient cases for comparison, multiple cases with correct auxiliary diagnostic results for the same disease are selected from the patient data as learning samples. Their corresponding auxiliary diagnostic results and diagnostic criteria are combined with the content of a pre-set summary prompt template and input into the associated large language model. The Feature Summarization Intelligent Subunit searches for commonalities among different cases of the same disease based on the content of the learning samples, providing multiple (e.g., M) diagnostic feature texts that match the disease, which serve as input for the subsequent Feature Refinement Intelligent Subunit. The summary prompt template is a pre-constructed, scalable prompt content that can be expanded within a certain scope.
[0039] Taking appendicitis as an example, the corresponding diagnostic features include: right lower quadrant pain, McBurney's point tenderness, elevated white blood cell count (range 9.9~27.5K / uL), CT imaging showing thickening of the appendix wall, nausea and vomiting, blurred surrounding fat spaces, and enlarged appendix (diameter ranging from 8mm to more than 1cm).
[0040] In a further preferred embodiment, the feature refinement intelligent body subunit filters out redundant feature texts from multiple diagnostic feature texts, specifically including: the feature refinement intelligent body subunit calculates the semantic similarity between each pair of diagnostic feature texts in the multiple diagnostic feature texts; filters out a third preset number of diagnostic feature texts with the highest semantic similarity, or filters out diagnostic feature texts with semantic similarity greater than a set value; and removes the remaining diagnostic feature texts as redundant feature texts.
[0041] Preferably, the semantic similarity is:
[0042] ;
[0043] in, Let A represent the semantic similarity between two diagnostic feature texts; A and B are vectors of the two diagnostic feature texts, respectively. and These are the magnitudes of the vectors of the two diagnostic feature texts, respectively. Besides the method described above that uses cosine similarity to calculate the semantic similarity between two diagnostic feature texts, other methods can also be used to calculate the semantic similarity between them, such as calculating the Euclidean distance between the semantic vectors.
[0044] Preferably, the feature refinement intelligent body subunit filters out erroneous feature texts, specifically including: for each diagnostic feature text retained after filtering out redundant feature texts: the feature refinement intelligent body subunit determines whether the diagnostic accuracy of the corresponding disease increases after removing the diagnostic feature text. If it increases, the diagnostic feature text is an erroneous feature text and is removed; if it decreases, the diagnostic feature text is a correct feature text and is retained.
[0045] The feature refinement intelligent body subunit refines multiple diagnostic feature texts output by the feature summarization intelligent body subunit based on a redundancy screening mechanism and an iterative ablation mechanism.
[0046] To reduce redundancy among the diagnostic feature texts fed back by the feature summarization agent subunit, this invention proposes a redundancy filtering mechanism. For M diagnostic feature texts obtained for the same disease, a text encoder is used to obtain their corresponding representation vectors, and cosine similarity is used to calculate the semantic similarity between different features. A greedy algorithm is then used to filter out the top N diagnostic feature texts with the greatest semantic differences (N < M). The text encoder preferentially uses a text encoder finely tuned for the medical field to improve its understanding of medical information.
[0047] For the N diagnostic feature texts retained after redundancy filtering, a concept-causal intervention scheme is established for iterative ablation. The N retained diagnostic feature texts after redundancy filtering may contain erroneous or false concepts, which could negatively impact the diagnostic agent's auxiliary diagnostic results. A method of sequentially removing N diagnostic feature texts is employed to verify the impact of each removed feature text on the diagnostic agent's diagnostic results. This method is validated on a large-scale real or simulated patient dataset for a specific disease. The validation results are quantified by assessing the change in the diagnostic accuracy of the diagnostic agent on the target disease dataset after the removal of the nth diagnostic feature text. Specifically, if removing the nth diagnostic feature text results in a decrease in the diagnostic accuracy of the diagnostic agent for the target disease, it indicates that the diagnostic feature text played a positive role in the agent's assisted diagnosis and should be retained. Conversely, if removing the nth diagnostic feature text results in an increase in the agent's diagnostic accuracy for the target disease, it indicates that the diagnostic feature text played a negative role in assisted diagnosis and should be removed. Finally, the remaining portion of the N diagnostic feature texts for the target disease after removing them one by one is saved as diagnostic experience in the experience database for the diagnostic agent to access and reference during assisted diagnostic tasks.
[0048] In subsequent implementation, for diagnostic feature texts of different diseases or newly added diagnostic feature texts of existing diseases, multiple iterations can be performed by the feature refinement intelligent agent subunit. The new diagnostic features that have been iterated and ablated are added to the experience database incrementally, thereby updating and improving the content of the experience database so that the diagnostic intelligent agent unit can call the corresponding diagnostic experience during the diagnosis process.
[0049] The experience database stores diagnostic feature text after iterative ablation by the feature summarization and feature refinement subunits. This diagnostic feature text can be stored as a vector or as natural language text. Information in the experience database is stored in key-value pairs, with the disease name as the key and the diagnostic feature as the value. When used, the diagnostic agent unit retrieves the diagnostic features corresponding to one or more diseases through keyword matching and combines them into a specified position within an existing prompt word template.
[0050] The dataset in the experience database, for example, was obtained by processing data using this method based on MIMIC-IV V2.2. It includes patient data for seven common diseases, comprising 4390 privacy-protected real patient cases. Each case covers four different types of medical clinical data: patient complaints, body fluid test reports, physical examination reports, and imaging reports, as well as the doctor's diagnosis. Data usage is as follows: 1314 patient cases were used as training samples for the feature summarization and feature refinement subunits to iteratively acquire diagnostic knowledge for each of the seven diseases; the remaining 3076 patient cases were used for testing. The data scope can be further expanded using private datasets or other public datasets in the future.
[0051] Figure 3 This is a schematic diagram of the structure of a disease-assisted diagnosis system based on a self-learning large language model intelligent agent, provided in another embodiment of the present invention. Its overall working process is as follows: Figure 4 As shown.
[0052] Please see Figure 3 Preferably, the system comprises multiple diagnostic agent units, each corresponding to a different large language model. The system also includes a referee agent unit. The referee agent unit determines whether the auxiliary diagnostic results output by each diagnostic agent unit are consistent. If they are consistent, the auxiliary diagnostic result is used as the final auxiliary diagnostic result.
[0053] Preferably, when the auxiliary diagnostic results output by each diagnostic agent unit are inconsistent, the referee agent unit is also used to analyze the differences between the auxiliary diagnostic results and diagnostic criteria output by each diagnostic agent unit, and output the corresponding difference analysis data to the corresponding diagnostic agent unit. The corresponding diagnostic agent unit corrects its output auxiliary diagnostic results and diagnostic criteria based on the difference analysis data and then transmits them to the referee agent unit again.
[0054] Figure 3 The diagnostic agent unit in the system shown is Figure 1 The diagnostic agent unit in the system shown has the same function, except that... Figure 3 The diagnostic agent unit in the system shown can obtain the auxiliary diagnostic results and diagnostic basis of all diagnostic agents in the previous round, and at the same time receive the difference analysis data from the referee agent unit. The patient information, the auxiliary diagnostic results and diagnostic basis of all diagnostic agents in the previous round, and the difference analysis data are input into the associated large language model to obtain new auxiliary diagnostic results and diagnostic basis, so as to correct the auxiliary diagnostic results of the previous round.
[0055] Preferably, when the auxiliary diagnostic results output by each diagnostic agent unit are still inconsistent after a set number of corrections, the correction is stopped, and the referee agent unit directly outputs the latest auxiliary diagnostic results and diagnostic basis output by each diagnostic agent unit.
[0056] Preferably, the referee intelligent agent unit determines whether the auxiliary diagnostic results output by each diagnostic intelligent agent unit are consistent, specifically including: the referee intelligent agent unit determines whether the auxiliary diagnostic results output by each diagnostic intelligent agent unit are consistent according to the preset disease name mapping rules.
[0057] The same disease may have multiple medical names, and some diseases may even have similar expressions but correspond to different names. Therefore, a pre-defined disease name mapping rule was designed to standardize and unify the auxiliary diagnostic results. Since the referee agent needs to determine whether the auxiliary diagnostic results from different diagnostic agents are consistent, even if the auxiliary diagnostic results are the same, differences in the output disease names may lead to discrepancies. For example, the descriptive term "appendicitis" may differ from the professional term "appendicitis." Directly judging the output of the diagnostic agent risks misclassifying similar expressions as inconsistent, leading to misdiagnosis. Therefore, a validated disease name mapping rule is used to unify the standard for judging result consistency.
[0058] The disease-assisted diagnosis system based on a self-learning large language model intelligent agent provided by this invention can associate multiple different large language models to construct diagnostic experiences for different diseases and store them in an experience database. During the case diagnosis process, a multi-agent diagnosis mechanism can be set up: multiple diagnostic intelligent agent units associate different large language models to simultaneously diagnose the case, each calling upon diagnostic experiences from its corresponding experience database, and discussing and reaching a consensus as the auxiliary diagnostic result.
[0059] Figure 3The system described involves core modules including diagnostic agent units that associate different large language models, an experience database, a referee agent unit, and a medical knowledge base (which internally stores pre-defined disease name mapping rules). In a specific implementation, diagnostic agent units 1, 2, ..., n, which associate different large language models, acquire patient information for the same patient, call upon diagnostic experience from the experience database, and obtain auxiliary diagnostic results and their corresponding diagnostic basis for that patient information. Each diagnostic agent unit inputs its diagnostic result into the referee agent unit. The referee agent unit, in conjunction with the medical knowledge base, determines whether the auxiliary diagnostic results of each diagnostic agent unit are consistent. If consistent, the auxiliary diagnostic results from multiple diagnostic agents are used as the final auxiliary diagnostic result. If inconsistent, the differences between them are analyzed, and the analysis results are fed back to the corresponding diagnostic agent unit, which then provides a new auxiliary diagnostic result, and the referee agent unit's judgment process is repeated. For cases where consensus cannot be reached within a certain number of consultations, a doctor can be introduced to make the final decision to diagnose the patient's case.
[0060] To verify the effectiveness of each agent in this invention, ablation experiments were conducted. By sequentially adding feature summarization and feature refinement agent sub-units, the impact of each agent on the auxiliary diagnostic capability of the diagnostic agent unit was verified. Experimental results show that, compared to directly using clinical diagnostic knowledge provided by experts, the feature summarization agent sub-unit can provide certain background information for the diagnostic agent unit and improve the accuracy of auxiliary diagnosis in similar cases, with an average improvement of approximately 14.5%. However, the information provided has not undergone iterative ablation, which has some negative impact. Adding the feature refinement agent sub-unit, based on redundancy and iterative ablation mechanisms, effectively reduces the information redundancy of diagnostic experience for the same disease in the experience database, and further improves the accuracy of auxiliary diagnosis in similar cases, by approximately 16% compared to the expert knowledge method.
[0061] Furthermore, experiments were conducted to compare the diagnostic experience provided in this invention with clinical knowledge of the same disease summarized by experts. The diagnostic experience proposed in this invention can effectively improve the diagnostic accuracy of the diagnostic agent unit on the same dataset, demonstrating the advantage of this system over existing human experience knowledge. Extensive experiments were conducted on the MIMIC-IVV2.2 dataset after data cleaning. For each disease, accuracy was used as the evaluation metric. Accuracy represents the percentage of correct predictions among all predictions, reflecting the diagnostic ability of the diagnostic agent unit for that disease. In this invention, the method for judging the correctness of the auxiliary diagnostic result is as follows: the diagnostic results given by doctors in the dataset are used as labels, the auxiliary diagnostic results of the diagnostic agent unit are used as predicted values, and the percentage of correctly predicted cases in the diagnostic agent unit relative to the test set of the same disease is calculated. Experimental results show that this invention achieves significant performance improvement in the diagnosis of patient data, demonstrating the advantages of self-learning agents in assisting disease diagnosis.
[0062] The technical solution of the present invention will be further described in detail below with reference to several preferred embodiments and accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Test methods in the following embodiments that do not specify specific conditions are generally performed under conventional conditions.
[0063] This invention proposes a disease-assisted diagnosis system based on a self-learning large language model intelligent agent, and a specific implementation example is as follows.
[0064] For a certain disease, such as Figure 2 As shown, the disease-assisted diagnosis system based on a self-learning large language model agent proposed in this invention firstly uses a diagnostic agent unit to perform auxiliary diagnosis on a subset of diseases in the dataset, obtaining a first auxiliary diagnosis result and diagnostic basis for each case, and calculating the first diagnostic accuracy on the dataset. Cases with correct first auxiliary diagnosis results are selected as the learning set, and the first auxiliary diagnosis results and diagnostic basis are input into a feature summarization agent subunit, which extracts the diagnostic features for the disease. The feature refinement agent subunit uses a text encoder to calculate semantic similarity, selects the top N diagnostic feature texts with low semantic similarity, and ablates the nth diagnostic feature. The ablation process involves inputting the N-1 diagnostic feature texts and patient data into a diagnostic agent unit to obtain a second auxiliary diagnostic result. The second diagnostic accuracy on the dataset is then calculated and compared with the first diagnostic accuracy. K diagnostic experiences (i.e., diagnostic feature texts) that have a positive effect on the diagnosis of the disease and have low semantic similarity are retained and stored in an experience database. A process is presented whereby the diagnostic agent unit, when faced with a new patient case, invokes corresponding diagnostic experiences from the experience database and combines them with patient information for auxiliary diagnosis. It is important to note that the diagnostic agent unit can retrieve diagnostic experiences for multiple diseases, rather than just a single disease, from the experience database.
[0065] More importantly, when implementing multi-intelligence diagnosis, such as... Figure 4 As shown, different diagnostic agent units (which can be associated with different large language models) acquire information about the same patient, call upon diagnostic experience from their respective experience databases (each calling upon diagnostic experience adapted to its own needs), and obtain auxiliary diagnostic results and diagnostic basis for the patient information. This serves as the result of the first round of consultation, and the auxiliary diagnostic results of each diagnostic agent unit are input into the referee agent unit. The referee agent unit, in conjunction with an external medical knowledge base, judges whether the multiple auxiliary diagnostic results are consistent. If they are consistent, the auxiliary diagnostic results of multiple diagnostic agent units are used as the final auxiliary diagnostic result; if they are inconsistent, the differences between them are analyzed and the analysis results are fed back to the corresponding diagnostic agent unit. In the second round of diagnosis, the diagnostic agent unit acquires the patient information, the results of the first round of consultation, and the analysis results of the referee agent unit, and re-provides the auxiliary diagnostic results, then repeats the above-mentioned process of judgment by the referee agent unit. According to the above process, the system performs multiple rounds of diagnosis for a certain case. For cases where consensus cannot be reached within a certain number of consultations, a doctor is introduced to make the final judgment to complete the diagnosis of the patient's case.
[0066] This invention provides a disease-assisted diagnosis system based on a self-learning large language model intelligent agent. It constructs a dynamic experience evolution closed loop consisting of "diagnosis-summarization-refinement-storage," enabling the system to automatically summarize, refine, verify, and optimize diagnostic knowledge from its own diagnostic results, achieving autonomous accumulation, redundancy removal, error correction, and incremental updates of diagnostic experience. Specifically, the diagnostic intelligent agent unit not only outputs diagnostic results but also generates interpretable diagnostic evidence, providing structured input for subsequent knowledge refinement. A knowledge distillation and ablation mechanism is proposed: the feature summarization intelligent agent subunit summarizes the commonalities and differences among similar cases, while the feature refinement intelligent agent subunit automatically selects key features that positively contribute to diagnostic accuracy through semantic redundancy removal and causal ablation experiments. A personalized experience base isolation mechanism is also proposed, configuring independent experience knowledge bases for different large language models to avoid knowledge pollution caused by model heterogeneity and ensure experience adaptability. This design fundamentally solves the shortcomings of existing systems, such as "static knowledge, non-accumulative experience, and uncorrectable errors," enabling disease-assisted diagnosis systems based on self-learning large language model agents to possess continuous learning and self-correction capabilities similar to those of clinicians, significantly improving diagnostic accuracy and generalization performance.
[0067] It should be understood that the above embodiments are merely illustrative of the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent changes or modifications made in accordance with the spirit and essence of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A disease-aided diagnostic system based on a self-learning large language model intelligent agent, characterized in that, It includes a diagnostic agent unit, a feature optimization agent unit, and a storage agent unit; The storage agent unit is equipped with an experience database and a large language model. The experience database contains diagnostic features of various diseases. The diagnostic intelligent agent unit is used to receive patient information, fill the patient information into the designated position of the preset prompt word template to generate patient standard information, select several diagnostic features corresponding to the patient standard information from the experience database, select the diagnostic feature most similar to the category of the patient standard information, generate corresponding auxiliary diagnostic results and diagnostic basis, and output them. The feature optimization agent unit is used to perform closed-loop optimization of the experience database using the large language model based on the auxiliary diagnostic results and diagnostic criteria output by the diagnostic agent unit, so as to summarize, filter redundancies and iteratively ablate the diagnostic features in the experience database. The feature optimization agent unit includes a feature summarization agent subunit and a feature refinement agent subunit: The feature summarization intelligent agent subunit is used to summarize the common features between different cases of the same disease based on the auxiliary diagnostic results and diagnostic basis output by the diagnostic intelligent agent unit, and generate multiple diagnostic feature texts that conform to the disease and transmit them to the feature refinement intelligent agent subunit; the common features include at least a first preset number of common features related to the disease and a second preset number of rare features related to the disease. The feature refinement intelligent agent subunit is used to filter out redundant and erroneous feature texts from multiple diagnostic feature texts to obtain effective feature texts for optimizing the experience database.
2. The disease-assisted diagnosis system based on a self-learning large language model intelligent agent according to claim 1, characterized in that, The feature refinement intelligent agent subunit filters out redundant feature text from multiple diagnostic feature texts, specifically including: The feature refinement intelligent agent subunit calculates the semantic similarity between pairwise diagnostic feature texts in multiple diagnostic feature texts; Select the third preset number of diagnostic feature texts with the highest semantic similarity, or select diagnostic feature texts with semantic similarity greater than a set value; The remaining diagnostic feature texts are treated as redundant feature texts and removed.
3. The disease-assisted diagnosis system based on a self-learning large language model intelligent agent according to claim 2, characterized in that, The semantic similarity is: ; in, Let A represent the semantic similarity between two diagnostic feature texts; A and B are vectors of the two diagnostic feature texts, respectively. and These are the moduli of the vectors of the two diagnostic feature texts, respectively.
4. The disease-assisted diagnosis system based on a self-learning large language model intelligent agent according to claim 1, characterized in that, The feature refinement intelligent agent subunit filters out erroneous feature text, specifically including: For each diagnostic feature text retained after filtering out redundant feature text: the feature refinement intelligent agent subunit determines whether the diagnostic accuracy of the corresponding disease increases after removing the diagnostic feature text. If it increases, the diagnostic feature text is an incorrect feature text and is removed; if it decreases, the diagnostic feature text is a correct feature text and is retained.
5. The disease-assisted diagnosis system based on a self-learning large language model intelligent agent according to any one of claims 1-4, characterized in that, The system comprises multiple diagnostic agent units, each corresponding to a different large language model. The referee agent unit is used to determine whether the auxiliary diagnostic results output by each diagnostic agent unit are consistent. If they are consistent, the auxiliary diagnostic result is taken as the final auxiliary diagnostic result.
6. The disease-assisted diagnosis system based on a self-learning large language model intelligent agent according to claim 5, characterized in that, When the auxiliary diagnostic results output by each diagnostic agent unit are inconsistent, the referee agent unit is also used to analyze the differences between the auxiliary diagnostic results and diagnostic basis output by each diagnostic agent unit, and output the corresponding difference analysis data to the corresponding diagnostic agent unit. The corresponding diagnostic agent unit corrects its output auxiliary diagnostic results and diagnostic basis based on the difference analysis data and then transmits them to the referee agent unit again.
7. The disease-assisted diagnosis system based on a self-learning large language model intelligent agent according to claim 6, characterized in that, When the auxiliary diagnostic results output by each diagnostic agent unit are still inconsistent after a set number of corrections, the correction is stopped, and the referee agent unit directly outputs the latest auxiliary diagnostic results and diagnostic basis output by each diagnostic agent unit.
8. The disease-assisted diagnosis system based on a self-learning large language model intelligent agent according to claim 5, characterized in that, The referee agent unit determines whether the auxiliary diagnostic results output by each diagnostic agent unit are consistent, specifically including: The referee agent unit determines whether the auxiliary diagnostic results output by each diagnostic agent unit are consistent based on the preset disease name mapping rules.
Citation Information
Patent Citations
Medical AI assistant implementation method and system based on data driving and large model
CN118098585A
Intelligent diagnosis and treatment method and system based on knowledge graph and large language model
CN120340806A