Medical question and answer method and system based on knowledge graph and agent

By employing a medical question-answering method based on knowledge graphs and intelligent agents, this paper addresses the problem of poor accuracy in drug recommendation in existing technologies. Through associative reasoning and neighbor path exploration, it improves the accuracy and reliability of disease diagnosis and drug recommendation, reduces the hallucination risk of LLM models, and enhances the quality of generated text.

CN121122680BActive Publication Date: 2026-03-20QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511675908.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-03-20
Estimated Expiration
2045-11-17

AI Technical Summary

Technical Problem

Existing knowledge graph-based medical question answering methods neglect the path relationships within the knowledge graph, resulting in poor accuracy in drug recommendations.

Method used

This approach utilizes a medical question-answering method based on knowledge graphs and intelligent agents, including keyword extraction, entity retrieval, vectorization representation, entity matching, and ranking. It explores related reasoning paths and neighbor paths, employs an LLM model for natural language conversion and diagnostic reasoning, and outputs diagnostic results.

Benefits of technology

It significantly improved the accuracy of the disease diagnosis chain and the reliability of drug recommendations, reduced the hallucination risk of LLM models, improved the quality and rationality of generated text, and enhanced the GPT-4 Ranking index.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121122680B_ABST
    Figure CN121122680B_ABST
Patent Text Reader

Abstract

The application discloses a medical question and answer method and system based on a knowledge graph and an intelligent agent, and relates to the technical field of natural language processing.The application first identifies core entities related to patient description information from a medical knowledge graph, and obtains an optimized entity set; then respectively performs associated reasoning path and neighbor path exploration on the entities in the optimized entity set, and obtains an associated reasoning path set and a second-order neighbor reasoning path set; then converts the paths in the associated reasoning path set and the second-order neighbor reasoning path set into natural language description information based on a prompt statement of natural language conversion, and the natural language description information constitutes a knowledge text set; finally, according to reasoning prompts, performs diagnostic reasoning on the information in the patient description information and the knowledge text set, and outputs a diagnostic result.The method disclosed by the application can achieve significant advantages in terms of factual accuracy, disease diagnosis accuracy and drug recommendation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of natural language processing, and particularly relates to a medical question and answer method and system based on a knowledge graph and an agent. BACKGROUND

[0002] The existing medical question and answer method includes a medical question and answer method based on a knowledge graph. The existing medical question and answer method based on a knowledge graph usually helps a large language model (LLM) to understand the logical association among diseases, symptoms and drugs through a knowledge graph. Although these existing medical question and answer methods based on a knowledge graph have achieved certain effects in improving the knowledge support of medical question and answer, the existing medical question and answer method based on a knowledge graph usually directly converts the structured information of the knowledge graph into a text input model, often neglecting the path relationship in the knowledge graph, resulting in insufficient knowledge utilization and thus poor drug recommendation accuracy. Therefore, the application proposes a medical question and answer method and system based on a knowledge graph and an agent. SUMMARY

[0003] In view of the defects and deficiencies in the prior art, the application proposes a medical question and answer method and system based on a knowledge graph and an agent.

[0004] To achieve the above-mentioned application purposes, the application adopts the following technical solutions:

[0005] A medical question and answer method based on a knowledge graph and an agent, comprising the following steps:

[0006] S1. sequentially performing keyword extraction, entity retrieval, obtaining vectorized representation, entity matching, and entity screening and sorting to identify core entities related to patient description information from a medical knowledge graph and obtain an optimized entity set;

[0007] S2. respectively performing associated reasoning path exploration and neighbor path exploration on the entities in the optimized entity set to obtain an associated reasoning path set and a second-order neighbor reasoning path set;

[0008] S3. writing a natural language conversion prompt for a second LLM model, converting the paths in the associated reasoning path set and the second-order neighbor reasoning path set into natural language description information based on the natural language conversion prompt of the second LLM model, and the natural language description information constitutes a knowledge text set;

[0009] S4. writing a reasoning prompt for a third LLM model, and the third LLM model performs diagnostic reasoning on the information in the patient description information and the knowledge text set according to the reasoning prompt, and outputs a diagnostic result.

[0010] Preferably, step S1 specifically comprises the following steps:

[0011] S1-1, extracting medical-related keywords from the patient description information using a first LLM model to obtain a keyword set;

[0012] S1-2, calling the Neo4j database using the Cypher query language, performing entity retrieval based on the medical knowledge graph to obtain a candidate entity set;

[0013] S1-3, mapping the entities in the candidate entity set and the keywords in the keyword set, and then converting them into vector representations using a pre-trained Word2Vec model;

[0014] S1-4, performing entity matching based on the keyword vector representation and the entity vector representation to obtain a set of semantically similar entities;

[0015] S1-5, the first agent performs entity filtering and sorting on the entities in the set of semantically similar entities based on the patient description information and the first agent prompt of the first agent.

[0016] Preferably, in step S1-3, the pre-trained Word2Vec model is obtained by pre-training the Word2Vec model.

[0017] Preferably, in step S2, the entities in the optimized entity set are subjected to associated reasoning path exploration, including continuous associated reasoning path exploration and non-continuous associated reasoning path exploration.

[0018] Preferably, the continuous associated reasoning path exploration of the entities in the optimized entity set includes the following steps:

[0019] 1) Taking the first entity in the optimized entity set as the starting entity, updating the optimized entity set to obtain set one, the difference between set one and the optimized entity set being that set one lacks the first entity in the optimized entity set; calling the Neo4j database using the Cypher query language to retrieve all entities connected to the first entity from the medical knowledge graph; when the second entity in the optimized entity set is included in all entities connected to the first entity retrieved from the medical knowledge graph, it indicates that there is an associated reasoning path between the first entity and the second entity in the optimized entity set.

[0020] 2) Take the second entity in the optimized entity set as a new starting entity, and update set one to obtain set two, the difference between set two and the optimized entity set is that set two lacks the first and second entities in the optimized entity set; call the Neo4j database by using the Cypher query language, retrieve all entities connected with the second entity from the medical knowledge graph, and when the third entity in the optimized entity set is contained in all entities connected with the second entity retrieved from the medical knowledge graph, it is indicated that there is an associated reasoning path between the second and third entities in the optimized entity set, at this time, there is a continuous associated reasoning path between the first, second and third entities in the optimized entity set;

[0021] 3) Take the third entity in the optimized entity set as a new starting entity, update set two to obtain set three, the difference between set three and the optimized entity set is that set three lacks the first, second and third entities in the optimized entity set; call the Neo4j database by using the Cypher query language, retrieve all entities connected with the third entity from the medical knowledge graph, and when the fourth entity in the optimized entity set is contained in all entities connected with the third entity retrieved from the medical knowledge graph, it is indicated that there is an associated reasoning path between the third and fourth entities in the optimized entity set, at this time, there is a continuous associated reasoning path between the first, second, third and fourth entities in the optimized entity set;

[0022] 4) Take the fourth entity as a new starting entity, update set three to obtain set four, set four is an empty set, at this time, the associated reasoning path exploration process ends, and the continuous associated reasoning path between the first, second, third and fourth entities in the optimized entity set is saved.

[0023] Preferably, the entities in the optimized entity set are subjected to non-continuous associated reasoning path exploration, which specifically includes the following steps:

[0024] a) Take the first entity in the optimized entity set as a starting entity, update the optimized entity set to obtain set five, the difference between set five and the optimized entity set is that set five lacks the first entity in the optimized entity set; call the Neo4j database by using the Cypher query language, retrieve all entities connected with the first entity from the medical knowledge graph; when the second entity in the optimized entity set is contained in all entities connected with the first entity retrieved from the medical knowledge graph, it is indicated that there is an associated reasoning path between the first and second entities;

[0025] b) Using the second entity in the optimized entity set as the new starting entity, and updating set five, we obtain set six. The only difference between set six and the optimized entity set is that set six lacks the first and second entities in the optimized entity set. Using the Cypher query language, we call the Neo4j database to retrieve all entities connected to the second entity from the medical knowledge graph. When the third entity in the optimized entity set is not included among all entities connected to the second entity retrieved from the medical knowledge graph, it indicates that there is no association reasoning path between the second and third entities in the optimized entity set. Since there is an association reasoning path between the first and second entities, we save the association reasoning path between the first and second entities in the optimized entity set.

[0026] c) Using the third entity in the optimized entity set as the new starting entity, update set six to obtain set seven. The only difference between set seven and the optimized entity set is that set seven lacks the first, second, and third entities in the optimized entity set. Using the Cypher query language to call the Neo4j database, retrieve all entities connected to the third entity from the medical knowledge graph. When the fourth entity in the optimized entity set is included among all entities connected to the third entity retrieved from the medical knowledge graph, it indicates that there is an association reasoning path between the third and fourth entities in the optimized entity set.

[0027] d) Using the fourth entity as the new starting entity, update set seven to obtain set eight, which is an empty set. At this point, the process of exploring the associated reasoning path ends, and the associated reasoning path between the third and fourth entities in the optimized entity set is saved.

[0028] Preferably, neighbor path exploration is performed on the entities in the optimized entity set, including first-order neighbor path exploration and second-order neighbor path expansion for the entities in the optimized entity set in sequence.

[0029] Preferably, the first-order neighbor path exploration specifically includes the following steps: sequentially performing first-order neighbor path exploration on entities in the optimized entity set based on primary neighbor relationship comparison and filtering and semantic relevance filtering to obtain a first-order neighbor path set.

[0030] Preferably, the second-order neighbor path is expanded, specifically including the following steps: performing second-order neighbor expansion on the first-order neighbor path set to obtain a second-order neighbor path set; optimizing the path between the entity in the entity set, the first-order neighbor entity in the first-order neighbor path set, and the second-order neighbor entity in the second-order neighbor entity set as a second-order neighbor path, and the second-order neighbor path constitutes the second-order neighbor path set; and verifying and sorting the second-order neighbor path in the second-order neighbor path set by using the second intelligent agent to obtain a second-order neighbor reasoning path set.

[0031] A medical question and answer system based on a knowledge graph and an intelligent agent, the medical question and answer system based on the knowledge graph and the intelligent agent being used to implement a medical question and answer method based on the knowledge graph and the intelligent agent, and the medical question and answer system based on the knowledge graph and the intelligent agent comprising:

[0032] An entity extraction module is used for keyword extraction, entity retrieval, acquisition of vectorized representation, entity matching, and entity screening and sorting to identify core entities related to patient description information from a medical knowledge graph and acquire an optimized entity set.

[0033] A knowledge reasoning path exploration module is used to perform associated reasoning path exploration and neighbor path exploration on the entities in the optimized entity set output by the entity extraction module respectively, and acquire an associated reasoning path set and a second-order neighbor reasoning path set.

[0034] A knowledge conversion module is used to convert the paths in the associated reasoning path set P A and the second-order neighbor reasoning path set P final into natural language description information based on a prompt of natural language conversion, and the natural language description information constitutes a knowledge text set.

[0035] A reasoning module is used to perform diagnostic reasoning on the information in the patient description information and the knowledge text set according to a reasoning prompt, and output a diagnostic result.

[0036] Compared with the prior art, the application has the following beneficial technical effects:

[0037] In the present application, the screening of the first intelligent agent at the entity level and the medical verification of the second intelligent agent at the path level jointly constitute a double guarantee mechanism, which greatly reduces the illusion risk of the LLM model; the associated reasoning path set obtained in the present application effectively improves the accuracy of the diagnosis chain from symptoms to diseases, and the second-order neighbor reasoning path set obtained after clinical value verification provides a reliable basis for drug recommendation, and then through the knowledge conversion module, it is converted into natural language description information, which jointly promotes the synergistic improvement of the three indicators of factual accuracy, disease diagnosis accuracy and drug recommendation accuracy of the generated text. In addition, through testing, the method described in the present application has achieved significant improvement in the GPT-4 Ranking indicator, which is mainly due to the following points: first, the first intelligent agent filters redundant entities according to the "clinical relevance + clinical practicality" double dimension and sorts them according to the clinical value, outputting an optimized entity set to reduce noise from the source; second, through continuous associated reasoning path exploration and non-continuous associated reasoning path exploration, the multi-hop relationship between entities is effectively mined, and the second intelligent agent verifies and sorts the second-order neighbor paths in the second-order neighbor path set P based on the second intelligent agent prompt, the department label D and the patient description information, the second intelligent agent only retains the second-order neighbor paths that meet the clinical value verification rules in the second intelligent agent prompt, and sorts them according to the sorting prompt in the second intelligent agent prompt to obtain a second-order neighbor reasoning path set, effectively guaranteeing the medical compliance of the reasoning logic; third, the second LLM model converts the paths in the associated reasoning path set P A and the second-order neighbor reasoning path set P final into natural language description information based on the prompt of the natural language conversion, and the natural language description information constitutes a knowledge text set, and the third LLM model diagnoses and reasons the information in the patient description information and the knowledge text set according to the reasoning prompt, so that the third LLM model can effectively take into account external medical knowledge (i.e. text information in the knowledge text set) and the Internal Knowledge of the third LLM model (i.e. internal knowledge), ultimately improving the quality and rationality of the generated text, and significantly improving the GPT-4 Ranking indicator. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1 is a schematic diagram of the medical question and answer system based on knowledge graph and intelligent agent of the present application;

[0039] Figure 2 is a structural schematic diagram of the associated reasoning path exploration submodule of the present application;

[0040] Figure 3 is a structural schematic diagram of the neighbor path exploration submodule of the present application. DETAILED DESCRIPTION

[0041] The application will be further described below in conjunction with the accompanying drawings and examples.

[0042] The application provides a medical question and answer system based on a knowledge graph and an agent, as shown in the drawings, comprising an entity extraction module, a knowledge reasoning path exploration module, a knowledge transformation module and a reasoning module; a medical question and answer method based on a knowledge graph and an agent, comprising the following steps: Figure 1

[0043] S1, the entity extraction module comprises a first LLM model and a pre-trained Word2Vec model; the input of the entity extraction module is patient description information and a medical knowledge graph; the entity extraction module is used for keyword extraction, entity retrieval, obtaining vectorized representation, entity matching, and entity screening and sorting, to identify core entities related to the patient description information from the medical knowledge graph and obtain an optimized entity set; comprising the following steps:

[0044] S1-1, the first LLM model is used to extract medical-related keywords from the patient description information to obtain a keyword set; wherein the first LLM model is used to extract medical-related keywords from the patient description information; specifically, the first LLM model receives patient description information, which is usually expressed in daily consultation language, such as "I have been experiencing blurred vision, accompanied by dizziness and nausea"; the first LLM model extracts medical-related keywords k1, k2,..., k n from the patient description information according to the prompt words pre-written by the first LLM model. n The keywords k1, k2,..., k n constitute a keyword set K={k1, k2,...,k}. In this embodiment, the prompt words pre-written by the first LLM model include role setting prompt words and task prompt words; wherein the role setting prompt words are: you are a medical assistant; the task prompt words are to extract medical-related keywords from the following medical problems.

[0045] The medical-related keywords extracted from the patient description information in this embodiment include but are not limited to disease names, symptom descriptions, drug use situations and diagnosis and examination schemes. The first LLM model used in this application is consistent with the GPT-3.5 model structure disclosed in the paper "How Robust is GPT-3.5 to Predecessors? A Comprehensive Study on Language Understanding Tasks", but the functions are not consistent; the first LLM model in this application is used to extract keywords.

[0046] S1-2, call Neo4j database by using Cypher query language, retrieve entities based on medical knowledge graph, and obtain candidate entity set; in this embodiment, all entities e1, e2,..., e n , are retrieved from the medical knowledge graph by calling the Neo4j database by using the Cypher query language. n to form the candidate entity set ;

[0047] S1-3, map the entities in the candidate entity set and the keywords in the keyword set, and then convert them into vector representations by using a pre-trained Word2Vec model; specifically:

[0048] First, map the entities in the candidate entity set and the keywords in the keyword set into the same semantic space to obtain the mapped keywords and the mapped entities, respectively; then, convert the mapped keywords and the mapped entities into vector representations by using a pre-trained Word2Vec model to obtain the corresponding keyword vector representations and entity vector representations, respectively; in this application, converting the mapped keywords and the mapped entities into vector representations can more effectively capture the potential semantic relationships and context association information between medical terms, thereby effectively improving the accuracy of calculating the cosine similarity between the keyword vector representations and the entity vector representations;

[0049] The Word2Vec model selected in this application is consistent with the Word2Vec model structure disclosed in “Efficient Estimation of Word Representations in Vector Space” and has the same function; the pre-trained Word2Vec model in this application is obtained by pre-training the Word2Vec model, which includes the following steps:

[0050] 1) Obtain medical training data:

[0051] 1-1) Select the first 99 diseases and their corresponding symptoms, medical examinations and drug information from the disease database, and then format them according to the arrangement of diseases, symptoms, medical examinations and drug information to obtain formatted data; in this application, the disease database is obtained from the following website: https: / / github.com / Kent0n-Li / ChatDoctor / blob / main / format_dataset.csv);

[0052] For example, the formatted data of a disease and its corresponding symptoms, medical examinations and drug information is as follows:

[0053] Panic disorder,"['Anxiety and nervousness', 'Depression', 'Shortnessof breath', 'Depressive or psychotic symptoms', 'Sharp chest pain', 'Dizziness', 'Insomnia', 'Abnormal involuntary movements', 'Chest tightness','Palpitations', 'Irregular heartbeat', 'Breathing fast']","['Psychotherapy','Mental health counseling', 'Electrocardiogram', 'Depression screen(Depression screening)', 'Toxicology screen', 'Psychological and psychiatric evaluation and therapy']","['Lorazepam', 'Alprazolam (Xanax)', 'Clonazepam','Paroxetine (Paxil)', 'Venlafaxine (Effexor)', 'Mirtazapine', 'Buspirone(Buspar)', 'Fluvoxamine (Luvox)', 'Imipramine', 'Desvenlafaxine (Pristiq)', 'Clomipramine', 'Acamprosate (Campral)']";

[0054] 1-2) The formatted data is converted into natural language descriptive text using the GPT-3.5 model, which serves as the document text; the GPT-3.5 model used in this application has the same structure as the GPT-3.5 model disclosed in the paper "How Robust is GPT-3.5 to Predecessors? A Comprehensive Study on Language Understanding Tasks";

[0055] 1-3) Extract question text from the first 99 dialogues in the GenMedGPT-5k dataset; these 99 extracted question texts involve 99 diseases, which correspond one-to-one with the first 99 diseases selected in the disease database; in this application, the GenMedGPT-5k dataset is a medical domain dataset consisting of 5000 doctor-patient dialogues; the GenMedGPT-5k dataset can be obtained from https: / / github.com / Kent0n-Li / ChatDocto;

[0056] 1-4) Merge the document texts and question texts for the same disease to obtain the merged training data, which is the medical training data.

[0057] 2) The Word2Vec model is pre-trained using medical training data to obtain a pre-trained Word2Vec model. The difference between the method of pre-training the Word2Vec model in this application and the method of training the Word2Vec model disclosed in the paper "MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models" is only that the input for pre-training the Word2Vec model in this application is the medical training data obtained in step 1).

[0058] S1-4. Perform entity matching based on keyword vector representation and entity vector representation to obtain a set of semantically similar entities. This includes the following steps:

[0059] First, calculate the keyword vector representation. and entity vector representation The cosine similarity is used to measure the keyword vector representation. With entity vector representation The semantic relevance between them is determined; then, the cosine similarity value is compared with a preset similarity threshold. The comparison is performed when the cosine similarity value is greater than or equal to a preset similarity threshold. If the entity matches the corresponding keyword, it is considered to be retained; otherwise, it is discarded. All retained entities and their corresponding keywords together constitute a semantically similar entity set. The above steps achieve the purpose of matching keywords with entities based on keyword vector representation and entity vector representation, effectively filtering out entities that match the keywords.

[0060] Among them, keyword vector representation With entity vector representation cosine similarity The calculation method is shown in equation (1):

[0061] (1)

[0062] In equation (1), This indicates the calculation of the dot product. Keyword vector representation The length of the mold, Entity vector representation The length of the mold, Keyword vector representation Magnitude and entity vector representation The product of moduli;

[0063] Keyword vector representation With entity vector representation cosine similarity This represents the keyword vector representation in the vector space. With entity vector representation The cosine of the angle between them, cosine similarity The value range is [-1, 1]; the closer the cosine similarity value is to 1, the stronger the keyword vector representation in the vector space. With entity vector representation The closer the directions are to each other, the more semantically similar the keywords and entities are; the closer the cosine similarity value is to 0, the better the keyword vector representation in the vector space. With entity vector representation The closer the direction is to vertical, the weaker the semantic connection between the keyword and the entity; the closer the cosine similarity value is to -1, the weaker the keyword vector representation in the vector space. With entity vector representation The closer the direction is to the opposite, the more semantically opposed or unrelated the keywords are to the entity.

[0064] S1-5, the entity filtering and ranking submodule takes the semantically similar entity set and patient description information as input. The first agent in the entity filtering and ranking submodule filters and ranks the entities in the semantically similar entity set based on the patient description information and the first agent's prompts, obtaining an optimized entity set. In this application, the setting of step S1-5 can effectively improve the accuracy and efficiency of the reasoning process. Step S1-5 specifically includes the following steps:

[0065] S1-5-1, entity screening, specifically comprising the following steps: writing a first agent prompt for the first agent, the first agent prompt including filtering rules and sorting rules; then, the first agent filters the clinical relevance and clinical utility between the patient description information and the entities in the semantic similar entity set according to the filtering rules set in the first agent prompt, and obtains entities with high clinical relevance and clinical utility;

[0066] The filtering rules include:

[0067] (1) Keep the clinical relevance entities and delete the entities without clinical relevance; wherein the clinical relevance entities refer to entities in the patient description information that are medically related to the condition, and at least include disease-related entities, symptom-related entities, and treatment-related entities, which are kept;

[0068] In this embodiment, the entities medically related to the condition include disease-related entities (such as "hypertension" and "retinopathy" entities), symptom-related entities (such as "weight loss" and "high blood sugar" entities), treatment-related entities (such as "metformin" and "fundus examination" entities), and site-related entities (such as "fundus" and "retina" entities), which are kept;

[0069] Entities without clinical relevance refer to entities in the patient description information that are not medically related to the condition; in this embodiment, entities without medical relevance to the condition (such as "shortness of breath" and "seasonal allergies" entities) are excluded.

[0070] (2) From the retained clinical relevance entities, filter out entities with clinical utility and delete entities without clinical utility; wherein entities with clinical utility refer to entities that can provide reference value for diagnosis or treatment, including disease-related entities, symptom-related entities, and treatment-related entities; wherein:

[0071] Disease-related entities include disease entities, disease entities, and syndrome entities, such as "hypertension" and "retinopathy" entities;

[0072] Symptom-related entities refer to symptoms or signs entities mentioned in the patient description information that are related to health status; such as "weight loss" and "high blood sugar" entities;

[0073] Treatment-related entities include drug entities and medical operation entities; such as "metformin" and "fundus examination" entities;

[0074] S1-5-2, entity ranking, the first intelligent agent ranks the entities with high clinical relevance and clinical utility filtered out according to the ranking rules in the first intelligent agent prompt, and obtains an optimized entity set; wherein the ranking rules are as follows:

[0075] (1) The disease-related entity is the first priority;

[0076] (2) The symptom-related entity is the second priority;

[0077] (3) The treatment-related entity is the third priority.

[0078] S2, the knowledge reasoning path exploration module respectively performs associated reasoning path exploration and neighbor path exploration on the entities in the optimized entity set output by the entity extraction module, and obtains an associated reasoning path set and a second-order neighbor reasoning path set; specifically including the following steps:

[0079] S2-1, the associated reasoning path exploration submodule uses the existing structured relationships in the medical knowledge graph to perform associated reasoning path exploration on the entities in the optimized entity set, and obtains an optimized associated reasoning path set; wherein the associated reasoning path exploration on the entities in the optimized entity set includes continuous associated reasoning path exploration and non-continuous associated reasoning path exploration on the entities in the optimized entity set;

[0080] S2-1-1, the associated reasoning path exploration submodule uses the existing structured relationships in the medical knowledge graph to perform continuous associated reasoning path exploration on the entities in the optimized entity set, specifically including the following contents:

[0081] 1) Take the first entity in the optimized entity set as the starting entity, update the optimized entity set to obtain set one, and the difference between set one and the optimized entity set is only that set one lacks the first entity in the optimized entity set; use Cypher query language to call Neo4j database, retrieve all entities connected with the first entity from the medical knowledge graph; when the second entity in the optimized entity set is contained in all entities connected with the first entity retrieved from the medical knowledge graph, it indicates that there is an associated reasoning path between the first entity and the second entity in the optimized entity set;

[0082] 2) Take the second entity in the optimized entity set as a new starting entity, and update set one to obtain set two, the difference between set two and the optimized entity set is that set two lacks the first and second entities in the optimized entity set; call the Neo4j database by using the Cypher query language, retrieve all entities connected with the second entity from the medical knowledge graph, when the third entity in the optimized entity set is contained in all entities connected with the second entity retrieved from the medical knowledge graph, it indicates that there is an associated reasoning path between the second and third entities in the optimized entity set, at this time, there is a continuous associated reasoning path between the first, second and third entities in the optimized entity set;

[0083] 3) Take the third entity in the optimized entity set as a new starting entity, update set two to obtain set three, the difference between set three and the optimized entity set is that set three lacks the first, second and third entities in the optimized entity set; call the Neo4j database by using the Cypher query language, retrieve all entities connected with the third entity from the medical knowledge graph, when the fourth entity in the optimized entity set is contained in all entities connected with the third entity retrieved from the medical knowledge graph, it indicates that there is an associated reasoning path between the third and fourth entities in the optimized entity set, at this time, there is a continuous associated reasoning path between the first, second, third and fourth entities in the optimized entity set;

[0084] 4) Take the fourth entity as a new starting entity, update set three to obtain set four, set four is an empty set, at this time, the associated reasoning path exploration process ends, and the continuous associated reasoning path between the first, second, third and fourth entities in the optimized entity set is saved;

[0085] Take the optimized entity set {v1, v2, v3, v4} as an example, wherein v1, v2, v3 and v4 in the optimized entity set all represent entities, perform continuous associated reasoning path exploration on the entities in the optimized entity set, as shown in FIG. 1, which specifically includes the following steps: Figure 2

[0086] ​1) Starting with entity v1 in the optimized entity set {v1,v2,v3,v4}, update the optimized entity set to obtain set one {v2,v3,v4}; use the Cypher query language to call the Neo4j database and retrieve all entities connected to entity v1 from the medical knowledge graph; when entity v2 from the optimized entity set is included among all entities connected to v1 retrieved from the medical knowledge graph, it indicates that there is an association reasoning path between entity v1 and entity v2.

[0087] 2) Using entity v2 as the new starting entity, update set one to obtain set two {v3, v4}; use Cypher query language to call the Neo4j database to retrieve all entities connected to entity v2 from the medical knowledge graph. When entity v3 in the optimized entity set is included among all entities connected to v2 retrieved from the medical knowledge graph, it indicates that there is an association reasoning path between entity v2 and entity v3. As can be seen from step 1), there is an association reasoning path between entity v1 and entity v2. Therefore, there is a continuous association reasoning path between entities v1, v2, and v3.

[0088] 3) Using entity v3 as the new starting entity, update set two to obtain set three {v4}; use the Cypher query language to call the Neo4j database and retrieve all entities connected to entity v3 from the medical knowledge graph. When entity v4 from the optimized entity set is included among all entities connected to entity v3 retrieved from the medical knowledge graph, it indicates that there is an association reasoning path between entity v3 and entity v4. In addition, as can be seen from step 2), there is a continuous association reasoning path between entities v1, v2, and v3. Therefore, there is a continuous association reasoning path between entities v1, v2, v3, and v4.

[0089] 4) Using entity v4 as the new starting entity, update set three to obtain set four, which is an empty set. At this point, the process of exploring the related reasoning path ends, and the continuous related reasoning paths between entities v1, v2, v3 and v4 are saved.

[0090] S2-1-2, the submodule for exploring associative reasoning paths utilizes existing structured relationships in the medical knowledge graph to explore non-continuous associative reasoning paths for entities in the optimized entity set. Specifically, it includes the following:

[0091] a) Take the first entity in the optimized entity set as the starting entity, update the optimized entity set to obtain set five, the difference between set five and the optimized entity set is only that set five lacks the first entity in the optimized entity set; call Neo4j database by using Cypher query language, retrieve all entities connected with the first entity from the medical knowledge graph; when the second entity in the optimized entity set is contained in all entities connected with the first entity retrieved from the medical knowledge graph, it indicates that there is an associated reasoning path between the first entity and the second entity;

[0092] b) Take the second entity in the optimized entity set as a new starting entity, and update set five to obtain set six, the difference between set six and the optimized entity set is only that set six lacks the first entity and the second entity in the optimized entity set; call Neo4j database by using Cypher query language, retrieve all entities connected with the second entity from the medical knowledge graph, when the third entity in the optimized entity set is not contained in all entities connected with the second entity retrieved from the medical knowledge graph, it indicates that there is no associated reasoning path between the second entity and the third entity in the optimized entity set, since there is an associated reasoning path between the first entity and the second entity, at this time, the associated reasoning path between the first entity and the second entity in the optimized entity set is saved;

[0093] c) Take the third entity in the optimized entity set as a new starting entity, update set six to obtain set seven, the difference between set seven and the optimized entity set is only that set seven lacks the first entity, the second entity and the third entity in the optimized entity set; call Neo4j database by using Cypher query language, retrieve all entities connected with the third entity from the medical knowledge graph, when the fourth entity in the optimized entity set is contained in all entities connected with the third entity retrieved from the medical knowledge graph, it indicates that there is an associated reasoning path between the third entity and the fourth entity in the optimized entity set;

[0094] d) Take the fourth entity as a new starting entity, update set seven to obtain set eight, set eight is an empty set, at this time, the associated reasoning path exploration process ends, and the associated reasoning path between the third entity and the fourth entity in the optimized entity set is saved.

[0095] Take the optimized entity set {v1, v2, v3, v4} as an example, wherein v1, v2, v3 and v4 in the optimized entity set all represent entities, perform non-continuous associated reasoning path exploration on the entities in the optimized entity set, as shown in Figure 2 , which specifically includes the following steps:

[0096] a) taking the entity v1 in the optimized entity set {v1, v2, v3, v4} as a starting entity, updating the optimized entity set to obtain set five {v2, v3, v4}; calling the Neo4j database by using the Cypher query language to retrieve all entities connected with the entity v1 from the medical knowledge graph; when the entity v2 in the optimized entity set is contained in all entities connected with the entity v1 retrieved from the medical knowledge graph, it indicates that there is an associated reasoning path between the entity v1 and the entity v2;

[0097] b) taking the entity v2 as a new starting entity and updating set five to obtain set six {v3, v4}; calling the Neo4j database by using the Cypher query language to retrieve all entities connected with the entity v2 from the medical knowledge graph; when the entity v3 in the optimized entity set is not contained in all entities connected with the entity v2 retrieved from the medical knowledge graph, it indicates that there is no associated reasoning path between the entity v2 and the entity v3, and there is an associated reasoning path between the entity v1 and the entity v2, at this time, the associated reasoning path between the entity v1 and the entity v2 is saved;

[0098] c) taking the entity v3 as a new starting entity and updating set six to obtain set seven {v4}; calling the Neo4j database by using the Cypher query language to retrieve all entities connected with the entity v3 from the medical knowledge graph; when the entity v4 in the optimized entity set is contained in all entities connected with the entity v3 retrieved from the medical knowledge graph, it indicates that there is an associated reasoning path between the entity v3 and the entity v4;

[0099] d) taking the entity v4 as a new starting entity and updating set seven to obtain set eight, and set eight is an empty set, at this time, the associated reasoning path exploration process ends, and the associated reasoning path between the entity v3 and the entity v4 is saved;

[0100] In the present application, when the entity v4 in the optimized entity set is not contained in all entities connected with the entity v3 retrieved from the medical knowledge graph in step c), it indicates that there is no associated reasoning path between the entity v3 and the entity v4; in this case, there is no associated reasoning path between the entity v2, the entity v3 and the entity v4, and only the continuous associated reasoning path between the entity v1 and the entity v2 exists, and since the associated reasoning path between the entity v1 and the entity v2 has been saved in step b), when the entity v4 in the optimized entity set is not contained in all entities connected with the entity v3 retrieved from the medical knowledge graph in step c), the associated reasoning path exploration process can be ended;

[0101] All the association reasoning paths obtained in steps S2-1-1 and S2-1-2 of the present application jointly constitute an association reasoning path set.

[0102] The present application can automatically discover and save association reasoning paths and corresponding entities in a knowledge graph by a joint exploration mechanism of continuous association reasoning path exploration and non-continuous association reasoning path exploration, and generate an association reasoning path set, which provides high-quality path basis support for subsequent knowledge reasoning.

[0103] In step S2-2, a neighbor reasoning path exploration submodule explores neighbor paths of entities in the optimized entity set based on existing structured relationships in the medical knowledge graph to obtain a second-order neighbor reasoning path set; wherein, the neighbor path exploration of the entities in the optimized entity set comprises first-order neighbor path exploration and second-order neighbor path extension of the entities in the optimized entity set in sequence.

[0104] Taking the optimized entity set as {v1, v2, v3, v4} for example, the working principle of the neighbor reasoning path exploration submodule is as shown in Figure 3 , which comprises the following steps:

[0105] In step S2-2-1, first-order neighbor path exploration of the entities in the optimized entity set is performed in sequence based on primary neighbor relationship comparison screening and semantic correlation screening to obtain a first-order neighbor path set; specifically comprising the following steps:

[0106] In step S2-2-1-1, first-order neighbor path exploration of the entities in the optimized entity set is performed based on primary neighbor relationship comparison screening, specifically comprising the following steps:

[0107] When the entity v1 in the optimized entity set is connected to multiple first-order neighbor entities through multiple relationships; for the case that the entity v1 is connected to multiple first-order neighbor entities through multiple same relationships, only the first-order neighbor path between the entity v1 and the first-order neighbor entity located on the association reasoning path is reserved; for the case that the entity v1 is connected to multiple first-order neighbor entities through other relationships that do not exist, the first-order neighbor path between the entity v1 and the first-order neighbor entity connected by the entity v1 through other relationships is reserved.

[0108] Taking the case that the entity v1 is connected to multiple first-order neighbor entities through relationships r1, r2, r3 and r4 for example.

[0109] When the entity v1 in the optimized entity set establishes a connection with multiple first-order neighbor entities through the relations r1, r2, r3, and r4; for the case that the entity v1 connects multiple first-order neighbor entities through multiple same relations r1, only the first-order neighbor path between the entity v1 and the first-order neighbor entity located on the association reasoning path is reserved; meanwhile, the first-order neighbor path between the entity v1 and the first-order neighbor entity connected by the entity v1 through other relations is also reserved; wherein the other relations refer to relations different from the relation r1, and here the other relations refer to the relations r2, r3, and r4; similarly, for the case that the entity v1 connects multiple first-order neighbor entities through multiple same relations r2 or r3 or r4, the way of reserving the first-order neighbor path is the same as that of reserving the first-order neighbor path in the case that the entity v1 connects multiple first-order neighbor entities through multiple same relations r1;

[0110] Similarly, the way of reserving the first-order neighbor path when other entities (i.e., the entities v2, v3, and v4) in the optimized entity set establish a connection with multiple first-order neighbor entities through the relations r1, r2, r3, and r4 is the same as that of reserving the first-order neighbor path when the entity v1 in the optimized entity set establishes a connection with multiple first-order neighbor entities through the relations r1, r2, r3, and r4;

[0111] Through the setting of step S2-2-1-1, the redundancy of the first-order neighbor entities caused by the relation conflict can be effectively reduced, thereby improving the determinacy and logical consistency of the neighbor path;

[0112] S2-2-1-2, based on the semantic correlation, the first-order neighbor paths reserved in step S2-2-1-1 are screened to obtain a first-order neighbor path set; specifically including the following steps:

[0113] The semantic similarity between the patient description information Q and the first-order neighbor path p reserved in step S2-2-1-1 is calculated by using the BERT model , wherein the first-order neighbor path refers to a path from an entity in the optimized entity set to a directly connected node along an edge (herein, the edge refers to a relation), and then the first-order neighbor path with a semantic similarity not lower than a preset threshold is screened to form a first-order neighbor path set P v ; the construction manner of the set is shown in formula (2):

[0114] (2)

[0115] In formula (2), p represents the first-order neighbor path reserved in step S2-2-1-1, Q represents the patient description information, represents the semantic similarity calculated by the BERT model, and δ is a preset similarity threshold; in this embodiment, δ takes a value of 0.7, and Pv is a first-order neighbor path set;

[0116] The BERT model used in the present application is consistent with the structure disclosed in the paper "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding" and has the same function. The semantic similarity calculation method based on BERT in the present application is consistent with the calculation method disclosed in the paper "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks";

[0117] S2-2-2, the first-order neighbor path set P v is expanded to obtain a second-order neighbor path set P ; the paths between the entities in the entity set, the first-order neighbor entities in the first-order neighbor path set, and the second-order neighbor entities in the second-order neighbor entity set that have connections are optimized as second-order neighbor paths, and these second-order neighbor paths constitute a second-order neighbor path set P ; the second intelligent agent is used to verify and sort the second-order neighbor paths in the second-order neighbor path set P to obtain a second-order neighbor reasoning path set P P final ; specifically including the following steps:

[0118] 1) The first-order neighbor path set P v is expanded to obtain a second-order neighbor path set P ; specifically including the following steps: using the Cypher query language to call the Neo4j database to retrieve all entities directly connected to the first-order neighbor entities in the first-order neighbor path set P v along an edge (here, the edge refers to a relationship) from the medical knowledge graph, to form a second-order neighbor entity set;

[0119] 2) Since the first-order neighbor paths selected by semantic similarity only filter out first-order neighbor paths with a semantic similarity not lower than a preset threshold at the semantic level, and cannot ensure the medical rationality of the selected first-order neighbor paths in the medical knowledge, therefore, in order to further ensure the medical rationality of the second-order neighbor path set obtained by expanding the first-order neighbor paths, the present application specially introduces a second intelligent agent, which acts as a proxy doctor. The second intelligent agent verifies and sorts the second-order neighbor paths in the second-order neighbor path set P ; specifically including the following steps:

[0120] 2-1), Department label identification: Specifically, the GPT-4 model (consistent with the GPT-4 model structure disclosed in the paper “GPT-4 Technical Report”) is used to perform semantic analysis on the patient description information, and the department label D is identified; for example, when the patient description information is “blood sugar rises with blurred vision”, the department label D is “endocrinology department”, and the department label D is used to set the medical department to which the second agent belongs;

[0121] 2-2), Clinical value verification and sorting: Specifically, the second agent prompt is written for the second agent, and the second agent verifies and sorts the second neighbor path in the second neighbor path set based on the second agent prompt, department label D and patient description information, and the second agent only retains the second neighbor path that meets the clinical value verification rule in the second agent prompt, and sorts according to the sorting prompt in the second agent prompt, to obtain the second neighbor reasoning path set P final ;

[0122] P final The acquisition method is as shown in formula (3):

[0123] (3)

[0124] In formula (3), P final represents the second neighbor reasoning path set, D represents the identified department label, and Q represents the patient description information, represents the second neighbor path set, and B represents the second agent prompt.

[0125] The second agent prompt written by the present application for the second agent includes a clinical value verification rule and a sorting prompt;

[0126] In the present application, the clinical value verification rule includes a diagnosis support rule and a treatment plan support rule; specifically, when the second neighbor path in the second neighbor set is verified for clinical value, as long as the second neighbor path in the second neighbor set meets any one of the diagnosis support rule and the treatment plan support rule, the second neighbor path is the second neighbor path that meets the clinical value verification rule, and will be retained by the second agent: wherein,

[0127] The diagnosis support rule is: only retain the path that can directly support the diagnosis of the disease (such as symptoms-diseases related to symptoms, diseases-complications related to diseases); for example, when the patient description information is “blood sugar rises with blurred vision”, the second neighbor path “blood sugar rises → diabetes” is retained; ​

[0128] The treatment plan support rule is: keep the path that can prompt the patient to take medicine, keep the path that can prompt the patient to check; for example: the second-order neighbor path "diabetes -> retinopathy -> fundus examination", "blood sugar rise -> diabetes -> kidney disease -> need dialysis", "diabetes -> insulin injection", "diabetes -> foot infection -> wound treatment" are kept;

[0129] The second intelligent agent sorts and outputs the kept second-order neighbor paths according to the sorting prompt in the second intelligent agent prompt, and the output second-order neighbor paths constitute the second-order neighbor reasoning path set P final The output second-order neighbor paths can assist in diagnosis and decision-making;

[0130] In this application, the sorting prompt is: according to the patient description information, evaluate the contribution of second-order neighbor paths of the same type to diagnosis and decision-making, and the path with higher contribution is sorted in front of the second-order neighbor paths of the same type; wherein, the types of second-order neighbor paths include medication type and examination type;

[0131] For example, when the second intelligent agent sorts the examination type paths "diabetes -> retinopathy -> fundus examination", "blood sugar rise -> diabetes -> kidney disease -> need dialysis" and "diabetes -> foot infection -> wound treatment" in the second-order neighbor path according to the sorting prompt, according to the patient description information (blood sugar rise + blurred vision), it will be determined that the examination type path "diabetes -> retinopathy -> fundus examination" has a higher contribution to diagnosis and decision-making, and is sorted in front.

[0132] For another example, when the second intelligent agent sorts the medication type paths "diabetes -> blood sugar rise -> intensive insulin therapy" and "diabetes -> foot infection -> antibiotic therapy" in the second-order neighbor path according to the sorting prompt, according to the patient description information (blood sugar rise), it will be determined that the medication type path "diabetes -> blood sugar rise -> intensive insulin therapy" has a higher contribution to diagnosis and decision-making, and is sorted in front.

[0133] S3, the knowledge conversion module, as shown in Figure 1 , takes the association reasoning path set P A and the second-order neighbor reasoning path set P final output by the knowledge reasoning path exploration module as input, the paths of the association reasoning path set P A and the second-order neighbor reasoning path set P final show medical entities and their relationships, in this application, the association reasoning path set P A and the second-order neighbor reasoning path set P finalThe paths in the path set P A and the second-order neighbor reasoning path set P final are all in a structured form, such as diabetes -> retinopathy -> fundus examination; the knowledge conversion module of the present application comprises a second LLM model; the present application first writes a natural language conversion prompt for the second LLM model, and then converts the associated reasoning path set P

[0134] into natural language description information based on the natural language conversion prompt of the second LLM model, and the natural language description information constitutes a knowledge text set;

[0135] Task: Some reasoning paths are provided here, which follow the format of "entity -> relationship -> entity", please convert each reasoning path into a description that conforms to the habits of natural language, the expression should be coherent and smooth, avoid literal translation, and refer to the following example format:

[0136] Reasoning path: diabetes -> retinopathy -> fundus examination;

[0137] Natural language: "Diabetes" may cause "retinopathy", and "fundus examination" is needed to rule out related problems.

[0138] S4, reasoning module; in the present application, the reasoning module comprises a third LLM model; the present application writes a reasoning prompt for the third LLM model; the third LLM model performs diagnostic reasoning on the patient description information and the text information in the knowledge text set according to the reasoning prompt, and outputs a diagnostic result; wherein the text information in the knowledge text set is external medical knowledge;

[0139] The reasoning prompt written for the third LLM model in the present application is as follows:

[0140] Task: Specify the LLM model as "AI doctor", and need to base on the patient description information, take the knowledge text set as the external medical knowledge base, and combine the internal knowledge of LLM for reasoning;

[0141] Patient input: {patient description information};

[0142] External medical knowledge: {knowledge text set};

[0143] Output: require the LLM model to output the diagnostic result in a step-by-step reasoning manner, and the diagnostic result covers at least the following three items:

[0144] 1. What diseases may the patient have?

[0145] 2. What examinations need to be performed for diagnosis?

[0146] 3. Which drugs or treatments are recommended for this disease?

[0147] Test:

[0148] To verify the effectiveness of the medical question and answer method described in the present application, the present application tests the existing six medical question and answer methods and the medical question and answer method described in the present application based on test set one and test set two, obtains generated text; then, the generated text is tested using the Bert model (the access address is https: / / github.com / google-research / bert), and the Precision, Recall and F1 Score evaluation results are obtained; the generated text is tested using the GPT-4 model, and the GPT-4 Ranking evaluation result, Total factualness, Disease diagnosis, Drug recommendation, and three evaluation indexes are obtained; the above evaluation results are shown in Tables 1 to 4;

[0149] The evaluation indexes in Tables 1 to 2 include Precision, Recall, F1 Score, GPT-4 Ranking, Total factualness, Disease diagnosis, Drug recommendation, and Average; wherein, Precision is used to measure the proportion of accurate and relevant words in the generated text; Recall is used to measure whether the generated text can cover important information in the reference text; F1 Score is the harmonic mean of Precision and Recall; GPT-4 Ranking is sorted by the GPT-4 model (the GPT-4 model in the present application is consistent with the GPT-4 model disclosed in the paper “GPT-4 Technical Report” in structure and function) according to the reference text, and the quality and rationality of the generated text; Total factualness, Disease diagnosis, and Drug recommendation are three evaluation indexes scored by the GPT-4 model, and Total factualness, Disease diagnosis, and Drug recommendation are three evaluation indexes that show the factual accuracy, disease diagnosis accuracy, and drug recommendation accuracy of the generated text obtained by different methods; Average is the average value of the three evaluation indexes of Total factualness, Disease diagnosis, and Drug recommendation.

[0150] The existing six medical question and answer methods selected in the application include: MindMap method (from MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models), GPT-3.5 method (from Evaluating open question answering evaluation), GPT-4 method (from GPT4Graph: Can large language models understand graph structured data? an empirical evaluation and benchmarking), BM25 Retriever method (from The probabilistic relevance framework: Bm25 and beyond), Embedding Retriever method (from Ontology-based semantic retrieval of documents using word2vec model), and KG Retriever method (from ERNIE 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation).

[0151] In the application, the test set one is obtained by randomly selecting 477 pieces of dialogue data from the dialogue data from the 100th onwards in the GenMedGPT-5k dataset to form the test set one.

[0152] The test set two is obtained by randomly selecting 394 pieces of dialogue data from the dialogue data from the 100th onwards in the CMCQA dataset to form the test set two.

[0153] Before testing, the EMCKG medical knowledge graph and the CMCKG medical knowledge graph are respectively constructed; the EMCKG medical knowledge graph is dedicated to the test task of test set one, and the CMCKG medical knowledge graph is dedicated to the test task of test set two; that is, the method described in the present application and the above six existing medical question and answer methods are input with the data in test set one and the EMCKG medical knowledge graph when testing based on test set one; the method described in the present application and the above six existing medical question and answer methods are input with the data in test set two and the CMCKG medical knowledge graph when testing based on test set two;

[0154] The EMCKG medical knowledge graph is constructed by using the knowledge graph implementation method disclosed in the open source website https: / / github.com / wylwilling / MindMap to construct the EMCKG medical knowledge graph from zero based on the data in the disease database;

[0155] The construction method of the CMCKG medical knowledge graph is to construct a medical knowledge graph based on the medical data of the knowledge graph question and answer system by using the knowledge graph implementation method disclosed in the open source website https: / / github.com / wylwilling / MindMap, to obtain the CMCKG medical knowledge graph; wherein the medical data acquisition website of the knowledge graph question and answer system is: https: / / github.com / liuhuanyong / QASystemOnMedicalKG / blob / master / data / medical.json).

[0156] Table 1 Precision, Recall, F1 Score and GPT-4Ranking evaluation results of each method in the present application based on test set one

[0157]

[0158] As can be seen from Table 1, when tested based on test set one, the evaluation indexes obtained by the MindMap method are better than those of other existing medical question and answer methods, therefore, the evaluation indexes obtained by the MindMap method and the method described in the present application are compared, as follows:

[0159] When tested based on test set one, compared with the MindMap method, the Precision evaluation index of the method described in the present application is improved by (0.7922-0.7596) / 0.7596=4.29%;

[0160] Compared with the MindMap method, the Recall evaluation index of the method described in the application is reduced (0.7979-0.7977) / 0.7979=0.25%;

[0161] Compared with the MindMap method, the F1 Score evaluation index of the method described in the application is improved (0.8115-0.7729) / 0.7729=4.99%;

[0162] Compared with the MindMap method, the GPT-4 Ranking evaluation index of the method described in the application is improved (2.88-2.09) / 2.09=27.43%;

[0163] Obviously, when the method described in the application is tested based on the test set one, the generated text obtained in the Precision, F1 Score and GPT-4 Ranking key evaluation indexes compared with the MindMap method has achieved improvement, especially in the GPT-4 Ranking index, the improvement amplitude reaches 27.43%, which fully embodies the quality and rationality of the generated text obtained by the method described in the application is better.

[0164] Table 2 Precision, Recall, F1 Score and GPT-4 Ranking evaluation results obtained by each method based on test set two

[0165]

[0166] From Table 2, it can be seen that when tested based on the test set two, the evaluation index obtained by the MindMap method is better than that of other existing medical question and answer methods, so the evaluation index obtained by the MindMap method and the method described in the application is compared, as follows:

[0167] Compared with the MindMap method, the Precision evaluation index of the method described in the application is improved (0.9417-0.9392) / 0.9392=0.26%;

[0168] Compared with the MindMap method, the Recall evaluation index of the method described in the application is improved (0.9316-0.9281) / 0.9281=0.37%;

[0169] Compared with the MindMap method, the evaluation index F1 Score of the method described in the application is improved (0.9466-0.9335) / 0.9335=1.40%;

[0170] Compared with the MindMap method, the GPT-4 Ranking evaluation index of the method described in the application is improved by (2.80-2.35) / 2.35=16.07%;

[0171] Obviously, when the method described in the application is tested based on the test set two, the generated text obtained in the test achieves improvement in the key evaluation indexes of Precision, Recall, F1 Score and GPT-4 Ranking compared with the MindMap method, especially in the GPT-4 Ranking index, the improvement amplitude reaches 16.07%, which fully embodies that the quality and rationality of the generated text obtained by the method described in the application are better.

[0172] From Table 1 and Table 2, it can be seen that whether the test is based on the test set one or the test set two, the method described in the application achieves significant improvement in the GPT-4 Ranking index compared with the MindMap method, which is mainly due to the following points: first, the first intelligent agent filters redundant entities by “clinical relevance + clinical practicality” dual dimensions and sorts them by clinical value, outputs an optimized entity set, and reduces noise from the source; second, the second intelligent agent verifies and sorts the second-order neighbor paths in the second-order neighbor path set P based on the second intelligent agent prompt, department label D and patient description information, the second intelligent agent only retains the second-order neighbor paths that meet the clinical value verification rules in the second intelligent agent prompt, and sorts them according to the sorting prompt in the second intelligent agent prompt to obtain a second-order neighbor reasoning path set, which effectively guarantees the medical compliance of the reasoning logic; third, the second LLM model converts the associated reasoning path set P A and the second-order neighbor reasoning path set P final into natural language description information based on the prompt of natural language conversion, and the natural language description information constitutes a knowledge text set, and the third LLM model diagnoses and reasons the information in the patient description information and the knowledge text set according to the reasoning prompt, so that the third LLM model can effectively consider both external medical knowledge (i.e. text information in the knowledge text set) and the Internal Knowledge of the third LLM model (i.e. internal knowledge), finally improve the quality and rationality of the generated text, and significantly improve the GPT-4 Ranking index.

[0173] Table 3: Evaluation results of Total factualness, Disease diagnosis, Drug recommendation and Average of the method of the present application based on test set one relative to the MindMap method, the GPT-3.5 method and the GPT-4 method

[0174]

[0175] Table 4: Evaluation results of Total factualness, Disease diagnosis, Drug recommendation and Average of the method of the present application based on test set one relative to the BM25 Retriever method, the Embedding Retriever method and the KG Retriever method

[0176]

[0177] In Table 3 and Table 4, Win represents the percentage of the method of the present application being better than the existing medical question and answer method, Tie represents the percentage of the method of the present application being tied with the existing medical question and answer method, and Lose represents the percentage of the method of the present application being worse than the existing medical question and answer method.

[0178] From Table 3 and Table 4, it can be seen that the evaluation indicators obtained by the MindMap method relative to other existing medical question and answer methods are better, therefore, the evaluation indicators obtained by the MindMap method and the method of the present application are compared, as follows:

[0179] From the foregoing, it can be seen that Average refers to the average value of the four evaluation indicators Total factualness, Disease diagnosis and Drug recommendation. That is to say:

[0180] The data 55.28% in the rightmost column of the second row in Table 3 represents the average value of the percentage of the method of the present application being better than the existing medical question and answer method in the Total factualness, Disease diagnosis and Drug recommendation indicators, which is referred to as the average win rate in the present application.

[0181] Similarly, the data 22.25% in the rightmost column of the third row in Table 3 represents the average value of the percentage of the method of the present application being tied with the existing medical question and answer method in the Total factualness, Disease diagnosis and Drug recommendation indicators, which is referred to as the average tie rate in the present application.

[0182] Similarly, the data 22.48% on the far right of the fourth row in Table 3 represents the average value of the percentage of the performance of the method described in the present application being inferior to the existing medical question and answer method in the Total factualness, Disease diagnosis, and Drug recommendation indicators, which is referred to as the average failure rate in the present application.

[0183] As can be seen from Table 3,

[0184] 1) The method described in the present application achieves an average winning rate of 55.28% compared to the MindMap method, and from the other data 52.71%, 54.58%, and 58.54% in the first row of Table 3, it can be seen that the method described in the present application has achieved more than half of the performance advantage in multiple dimensions such as factual accuracy, disease diagnosis accuracy, and drug recommendation accuracy, especially in drug recommendation accuracy.

[0185] Obviously, the method described in the present application has better performance than the existing MindMap method in multiple key evaluation indicators, especially in factual accuracy, disease diagnosis accuracy, and drug recommendation accuracy. In the present application, the screening of the first intelligent agent at the entity level and the medical verification of the second intelligent agent at the path level jointly constitute a double protection mechanism, which greatly reduces the illusion risk of the LLM model; the set of associated reasoning paths obtained in the present application effectively improves the accuracy of the diagnosis chain from symptoms to diseases, while the set of second-order neighbor reasoning paths obtained after clinical value verification provides a reliable basis for drug recommendation, and then through the knowledge conversion module to convert into natural language description information, which together contributes to the synergistic improvement of the three indicators of factual accuracy, disease diagnosis accuracy, and drug recommendation accuracy.

Claims

1. A medical question-answering method based on knowledge graphs and intelligent agents, characterized in that: Includes the following steps: S1. Sequentially perform keyword extraction, entity retrieval, vectorization representation acquisition, entity matching, entity filtering, and sorting to identify core entities related to patient description information from the medical knowledge graph and obtain an optimized entity set; S2. Explore the association reasoning path and the neighbor path for the entities in the optimized entity set to obtain the association reasoning path set and the second-order neighbor reasoning path set. In step S2, the entities in the optimized entity set are explored for related reasoning paths, including the exploration of continuous related reasoning paths and non-continuous related reasoning paths. The exploration of continuous association reasoning paths for entities in the optimized entity set includes the following steps: 1) Starting with the first entity in the optimized entity set, update the optimized entity set to obtain set one. The only difference between set one and the optimized entity set is that set one lacks the first entity from the optimized entity set. Use the Cypher query language to call the Neo4j database and retrieve all entities connected to the first entity from the medical knowledge graph. When the second entity from the optimized entity set is included among all entities connected to the first entity retrieved from the medical knowledge graph, it indicates that there is an association reasoning path between the first and second entities in the optimized entity set. 2) Using the second entity in the optimized entity set as the new starting entity, and updating set one, we get set two. The only difference between set two and the optimized entity set is that set two lacks the first and second entities in the optimized entity set. When all entities connected to the second entity are found in the medical knowledge graph, including the third entity in the optimized entity set, it indicates that there is a continuous association reasoning path between the first and third entities in the optimized entity set. 3) Using the third entity in the optimized entity set as the new starting entity, update set two to obtain set three. The only difference between set three and the optimized entity set is that set three lacks the first to third entities in the optimized entity set. When the fourth entity in the optimized entity set is found to be among all entities connected to the third entity in the medical knowledge graph, it indicates that there is a continuous association reasoning path between the first to fourth entities in the optimized entity set. 4) Using the fourth entity as the new starting entity, update set three to obtain set four, which is an empty set. At this point, the process of exploring the related reasoning path ends, and the continuous related reasoning path between the first and fourth entities in the optimized entity set is saved. The exploration of non-continuous association reasoning paths for entities in the optimized entity set includes the following steps: a) Starting with the first entity in the optimized entity set, update the optimized entity set to obtain set five. The only difference between set five and the optimized entity set is that set five lacks the first entity from the optimized entity set. Use the Cypher query language to call the Neo4j database and retrieve all entities connected to the first entity from the medical knowledge graph. When the second entity from the optimized entity set is included among all entities connected to the first entity retrieved from the medical knowledge graph, it indicates that there is an association reasoning path between the first and second entities. b) Using the second entity in the optimized entity set as the new starting entity, and updating set five, we obtain set six. The only difference between set six and the optimized entity set is that set six lacks the first and second entities in the optimized entity set. When the third entity in the optimized entity set is not found among all entities connected to the second entity retrieved from the medical knowledge graph, it indicates that there is no association reasoning path between the second and third entities in the optimized entity set, and the association reasoning path between the first and second entities is saved. c) Using the third entity in the optimized entity set as the new starting entity, update set six to obtain set seven. The only difference between set seven and the optimized entity set is that set seven lacks the first to third entities in the optimized entity set. When all entities connected to the third entity are retrieved from the medical knowledge graph, including the fourth entity in the optimized entity set, it indicates that there is an associated reasoning path between the third and fourth entities in the optimized entity set. d) Using the fourth entity as the new starting entity, update set seven to obtain set eight, which is an empty set. At this point, the process of exploring the related reasoning path ends, and the related reasoning path between the third and fourth entities in the optimized entity set is saved. S3. The prompts based on natural language conversion convert the paths in the set of related reasoning paths and the set of second-order neighbor reasoning paths into natural language description information, and the natural language description information constitutes a set of knowledge text. S4. Based on the reasoning prompts, perform diagnostic reasoning on the patient's description information and the information in the knowledge text set, and output the diagnostic results.

2. The medical question-answering method based on knowledge graphs and intelligent agents according to claim 1, characterized in that: Step S1 specifically includes the following steps: S1-1. Extract medical-related keywords from patient description information to obtain a keyword set; S1-2. Use the Cypher query language to call the Neo4j database, perform entity retrieval based on the medical knowledge graph, and obtain a candidate entity set; S1-3: Map the entities in the candidate entity set to the keywords in the keyword set, and then use the pre-trained Word2Vec model to convert them into vector representations; S1-4. Perform entity matching based on keyword vector representation and entity vector representation to obtain a set of semantically similar entities; S1-5. The first intelligent agent filters and sorts entities in the semantically similar entity set based on its prompts and patient description information.

3. A medical question-answering method based on knowledge graphs and intelligent agents according to claim 2, characterized in that: In steps S1-3, the pre-trained Word2Vec model is obtained by pre-training the Word2Vec model.

4. A medical question-answering method based on knowledge graphs and intelligent agents according to claim 1, characterized in that: The process involves exploring neighbor paths for entities in the optimized entity set, including sequentially exploring first-order neighbor paths and expanding second-order neighbor paths for each entity in the optimized entity set.

5. A medical question-answering method based on knowledge graphs and intelligent agents according to claim 4, characterized in that: The first-order neighbor path exploration includes the following steps: first-order neighbor path exploration is performed on the entities in the optimized entity set based on primary neighbor relationship comparison and semantic relevance filtering to obtain the first-order neighbor path set.

6. A medical question-answering method based on knowledge graphs and intelligent agents according to claim 4, characterized in that: The second-order neighbor path expansion specifically includes the following steps: expanding the first-order neighbor path set with second-order neighbors to obtain a second-order neighbor path set; optimizing the paths connecting entities in the entity set, first-order neighbor entities in the first-order neighbor path set, and second-order neighbor entities in the second-order neighbor entity set to form second-order neighbor paths, and these second-order neighbor paths constitute the second-order neighbor path set; using a second agent to verify and sort the second-order neighbor paths in the second-order neighbor path set to obtain a second-order neighbor inference path set.

7. A medical question-answering system based on knowledge graphs and intelligent agents, characterized in that: The knowledge graph-based and agent-based medical question-answering system is used to implement the knowledge graph-based and agent-based medical question-answering method as described in claim 1. The knowledge graph-based and agent-based medical question-answering system includes: The entity extraction module is used for keyword extraction, entity retrieval, obtaining vectorized representations, entity matching, and entity filtering and sorting to identify core entities related to patient description information from the medical knowledge graph and obtain an optimized entity set. The knowledge reasoning path exploration module is used to explore the associated reasoning path and the neighbor path respectively for the entities in the optimized entity set output by the entity extraction module, and obtain the associated reasoning path set and the second-order neighbor reasoning path set. The knowledge conversion module is used to convert the paths in the set of related reasoning paths and the set of second-order neighbor reasoning paths into natural language description information based on prompts based on natural language conversion. The natural language description information constitutes a set of knowledge text. The reasoning module is used to perform diagnostic reasoning based on the patient description information and information in the knowledge text set, and output the diagnostic results.

Citation Information

Patent Citations

  • Intelligent question-answering system based on medical knowledge graph

    CN111046272A

  • Elderly disease health question and answer method based on multi-agent and knowledge graph

    CN120542573A