Intelligent Construction and Retrieval Method of Medical Knowledge Graph Based on GraphRAG and LLM
By constructing a medical knowledge graph based on GraphRAG and LLM, the problem that traditional medical search methods are difficult to understand complex queries is solved, and efficient and accurate medical information retrieval is achieved, and medical research and clinical practice are supported.
Patent Information
- Application Number
- CN202411766494.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-12-03
AI Technical Summary
Traditional medical search methods are difficult to understand complex query requirements, resulting in insufficient relevance and accuracy of search results, and it is impossible to efficiently screen out relevant and accurate medical information.
Combining GraphRAG and LLM technologies, a medical knowledge graph is constructed, and efficient retrieval of medical information is achieved through data collection, cleaning, segmentation, entity and relationship extraction, entity analysis, community division and abstract generation.
It significantly improves the accuracy and efficiency of medical literature search, users can quickly locate the required medical data, and the system can intelligently identify the intent of query and provide highly relevant information, shorten the search time and improve the accuracy of search.
Smart Images

Figure CN119719383B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for constructing and retrieving a medical knowledge graph, and in particular to an intelligent method for constructing and retrieving a medical knowledge graph based on GraphRAG and LLM, belonging to the technical field of natural language processing. Background Art
[0002] In the contemporary medical field, medical information retrieval systems are important tools for improving medical efficiency and ensuring medical quality. Doctors, researchers, and other medical professionals often need to search for, evaluate, and apply a large amount of medical literature, case reports, and research data in order to make scientific and accurate medical decisions. However, with the continuous development of the medical field, the massive growth and complexity of medical information pose challenges to doctors and researchers, who need to screen out the most relevant and accurate information from a large amount of data within a limited time.
[0003] Traditional medical retrieval methods often rely on keyword matching and encounter some limitations when facing complex query requirements, making it difficult to understand the complex queries of users and resulting in insufficient relevance and accuracy of retrieval results.
[0004] In recent years, with the rapid development of artificial intelligence and natural language processing technologies, large language models (LLMs) have provided new solutions for medical literature retrieval. GraphRAG uses a knowledge graph to organize and represent entities and their relationships in the medical field. By linking entities in medical literature to corresponding entries in the knowledge graph, the content of documents can be more accurately identified and understood, enriching the retrieval context and improving the relevance and precision of retrieval. Summary of the Invention
[0005] The object of the present invention is to provide an intelligent method for constructing and retrieving a medical knowledge graph based on GraphRAG and LLM.
[0006] To achieve the above object, the first technical solution adopted by the present invention is:
[0007] An intelligent method for constructing a medical knowledge graph based on GraphRAG and LLM, comprising the following steps:
[0008] Step 1: Data collection and processing: The data sources include: academic papers, clinical guidelines, drug instructions, health information, and expert interviews; which consists of the following specific steps:
[0009] Step 1-1: Convert the data into text data in a preset format to form more than one record;
[0010] Step 1-2: Data Cleaning: Calculate the hash value of each record, keep only one record with the same hash value, remove the useless information in the retained records, where the useless information includes special symbols, garbled characters, headers, footers, and dates, and convert the retained records into corresponding medical texts;
[0011] Step 2: Text Segmentation: Segment each medical text using full stops, semicolons, question marks, exclamation marks, and line breaks as markers to obtain more than 1 text block; the number of tokens in each text block is less than the preset value of the text block; the number of overlapping tokens between two adjacent text blocks is the preset number of overlapping tokens;
[0012] Step 3: Extract Entities and Relationships: Use the LLM to extract entity-relationship triples, and at the same time merge the same entity-relationship triples from different data sources to build a graph database;
[0013] Step 4: Entity Resolution: Consists of the following specific steps:
[0014] Step 4-1: Obtain the entities and descriptive attributes included in the triples in the graph database, and obtain the text embedding vectors corresponding to the entities;
[0015] Step 4-2: Construct a K-Nearest Neighbor Graph:
[0016] Project all entities and their corresponding text embedding vectors into the graph database;
[0017] Set the predefined number of neighbors topk, where topk is the number of relationships of each entity;
[0018] The similarity between entities is measured using cosine similarity, and its calculation method is:
[0019]
[0020] where n represents the dimension of the vector, i represents the index from 1 to n, and x i represents the value of vector x in the i-th dimension, and y i represents the value of vector y in the i-th dimension;
[0021] Establish connections between two entities that meet the preset number of neighbors topk and the similarity preset conditions; construct a K-Nearest Neighbor Graph;
[0022] Step 5: Identify Weakly Connected Components: Identify at least one weakly connected component in the K-Nearest Neighbor Graph, and there is a relationship between any two entities in the weakly connected component;
[0023] Step 6: Word Distance Filtering: The word distance is:
[0024]
[0025] Among them, is an indicator function. When a i = b j the value is 0, otherwise it is 1; a and b are two entities of the weakly connected component; i and j are the position numbers in entities a and b respectively;
[0026] Reduce the weakly connected component so that the word distances between entities in the weakly connected component are all less than a preset distance threshold;
[0027] Step 7: LLM evaluation: Use LLM to evaluate and decide whether to merge the entities in the weakly connected component, and merge the entities to be merged and the corresponding relationships to obtain an updated knowledge graph;
[0028] Step 8: Generate element summaries: Use LLM to generate element summaries for each entity and relationship in the updated knowledge graph, and store the element summaries as attributes of the entities;
[0029] Step 9: Generate graph communities: Use the Leiden algorithm to perform hierarchical community partitioning on the entities, and establish corresponding community entities for each hierarchical community;
[0030] Step 10: Generate community summaries: Use LLM to generate community summaries based on the community entities and relationships, and store the community summaries as attributes of the community entities; obtain a medical knowledge graph.
[0031] The second technical solution adopted by the present invention is:
[0032] An intelligent retrieval method for a medical knowledge graph constructed by using the technical solution described in Technical Solution 1, including the following steps:
[0033] Step 1: Construct a knowledge graph vector library: Use an Embedding model to convert each entity and attribute of the medical knowledge graph in the knowledge graph library constructed by Technical Solution 1 into vector representations as knowledge graph vectors, and construct a knowledge graph vector database;
[0034] Step 2: Construct a text vector library: Use an Embedding model to represent the text blocks obtained during the construction of the knowledge graph library by Technical Solution 1 as text vectors, and construct a text vector library;
[0035] Step 3: Vectorize user question entities and question vectorization:
[0036] Step 3-1: Use LLM to extract at least one entity entity0 in the user question as the initial root node; use an Embedding model to convert entity entity0 into an entity vector entity0_vector;
[0037] Step 3-2: Use the Embedding model to convert the user question into a question vector query_vector;
[0038] Step 4: Multi-layer entity and relationship retrieval:
[0039] Step 4-1: Determine whether entity entity0 exists in the medical knowledge graph. If it exists, go to Step 4-2; otherwise, go to Step 4-3;
[0040] Step 4-2: Retrieve the first-layer entities entity1 and their relationships relationship1 that are related to entity entity0 to obtain at least one top-level triple (entity0, relationship1, entity1); then retrieve the second-layer entities entity2 and relationships relationship2 that are related to each first-layer entity entity1 to obtain at least one first-layer triple (entity1, relationship2, entity2); output the top-level triples, first-layer triples, the first-layer entity entity1 and its element summary, and the second-layer entity entity2 and its element summary; go to Step 5;
[0041] Step 4-3: Find the entities in the medical knowledge graph vector library that reach or exceed the preset semantic relevance threshold for entity entity0 as similar entities; if there are any, go to Step 4-4; otherwise, go to Step 4-5;
[0042] Step 4-4: Sort the graph communities: Sort according to the number of similar entities in the community from most to least, and select a preset number of communities from most to least, and output the communities and community summaries; go to Step 5;
[0043] Step 4-5: Determine it as a global problem, and output each community and its community summary level by level according to the preset token number limit;
[0044] Step 5: Text retrieval, find the text blocks in the text vector library that have a semantic similarity greater than the preset semantic similarity threshold with the question vector query_vector as similar text blocks, and sort the similar text blocks according to the similarity from high to low; output a preset number of similar text blocks;
[0045] Step 6: LLM summary: Construct a prompt text with the top-level triples, first-layer triples, the first-layer entity entity1 and its element summary, the second-layer entity entity2 and its element summary, similar text blocks, and the user question according to the preset prompt template, input the prompt text into the LLM, and the LLM outputs the LLM retrieval result;
[0046] Step 7: Output the result, and output the LLM retrieval result and similar text blocks as retrieval information.
[0047] Further, it also includes Step 8: User feedback collection: Collect the user's evaluation of the retrieval information.
[0048] Further, in Step 6, the prompt template is "
Instruction
Known Information
Question
[0049] Adopting the above technical solution, the beneficial effects achieved by the present invention are:
[0050] The present invention combines GraphRAG and LLM large model technologies to achieve efficient retrieval and accurate presentation of medical information. By deeply integrating graph-structured data and natural language processing capabilities, the system significantly improves the accuracy and efficiency of medical literature retrieval. Users can quickly locate the required medical materials through simple interactions. The system can intelligently identify the user's query intention and provide highly relevant medical information, greatly shortening the search time and improving the retrieval accuracy. Medical researchers and clinicians can use this system to quickly access the latest research results and clinical guidelines, accelerating the dissemination and application of knowledge, and providing strong support for research and practice in the medical field. Description of the Drawings
[0051] Figure 1 is the flowchart of Embodiment 1 of the present invention;
[0052] Figure 2 is the medical knowledge graph constructed in Embodiment 1 of the present invention;
[0053] Figure 3 is the hierarchical community graph divided by the Leiden algorithm in Embodiment 1 of the present invention;
[0054] Figure 4 is the flowchart of Embodiment 2 of the present invention. Detailed Embodiments
[0055] Embodiment 1:
[0056] A method for intelligent construction of a medical knowledge graph based on GraphRAG and LLM, comprising the following steps:
[0057] Step 1: Data collection and processing: The data sources include: academic papers, clinical guidelines, drug instructions, health information, and expert interviews; it consists of the following specific steps:
[0058] In this embodiment, the latest research literature is obtained from the PubMed database for academic papers; clinical guidelines are downloaded from official institutions such as the National Institute for Health and Care Excellence (NICE) in the UK and the US Food and Drug Administration (FDA); drug instructions are obtained from the official websites of pharmaceutical companies or drug regulatory departments; health information is collected from authoritative medical news websites and professional blogs; expert interviews: record expert opinions and experience sharing.
[0059] Step 1-1: Convert the data into text data in a preset format to form more than one record;
[0060] In this embodiment, for pure text PDF files, the Python library PyMuPDF is used to extract text; for PDF files containing images, OCR technology is used, combined with libraries such as the Python Imaging Library for image recognition and text extraction.
[0061] Step 1-2: Data cleaning: Calculate the hash value of each record, keep only one record with the same hash value, remove the useless information in the retained record, and the useless information includes special symbols, garbled characters, headers, footers, and dates, and convert the retained record into the corresponding medical text;
[0062] In this embodiment, the text is converted to a unified UTF-8 encoding to ensure consistent text encoding.
[0063] Step 2: Split the text: Split each medical text using full stops, semicolons, question marks, exclamation marks, and line breaks as markers to obtain more than 1 text block; the number of tokens in each text block is less than the preset value of the text block; the number of overlapping tokens between two adjacent text blocks is the preset number of overlapping tokens;
[0064] In this embodiment, each text block is cut into a maximum of 600 tokens. To avoid losing the context and reference relationships of scattered specific entities in the document, the number of overlapping tokens between two adjacent text blocks is 50.
[0065] Step 3: Extract entities and relationships: Use the LLM to extract entity relationship triples, and at the same time merge the same entity relationship triples from different data sources to establish a graph database;
[0066] For example, from the relevant information on diabetes, we can extract ("diabetes", "clinical symptoms and signs", "thirst"), indicating that "diabetes" as a disease entity, one of its "clinical symptoms and signs" is "thirst". We extracted approximately 489,337 entities and 657,996 relationships and stored them in the neo4j graph database.
[0067] In this embodiment, the prompt for extracting entities and relationships input to the LLM is:
[0068] "Some text is provided below. According to the text, extract knowledge triples of diseases, drugs, diagnosis and treatment, symptoms, and attributes in the form of (entity, relationship, entity), avoiding using stop words. Only output the triples without any redundant explanations and descriptions.
[0069] ---------------------------------------
[0070] Example:
[0071] Medical text: The typical clinical symptoms of diabetes usually appear a few days to a few weeks before diagnosis, including polyuria, polydipsia, weight loss, fatigue, and blurred vision caused by swelling of the lens due to the osmotic effect of hyperglycemia.
[0072] Triples:
[0073] Diabetes, clinical symptom, polyuria
[0074] Diabetes, clinical symptom, polydipsia
[0075] Diabetes, clinical symptom, weight loss
[0076] Diabetes, clinical symptom, fatigue
[0077] Diabetes, clinical symptom, blurred vision
[0078] Hyperglycemia, causes, lens swelling
[0079] ---------------------------------------
[0080] Text: {text}
[0081] Triples: ”’
[0082] Step 4: Entity resolution: It consists of the following specific steps:
[0083] Step 4-1: Obtain the entities and descriptive attributes included in the triples in the graph database, and use the Embedding model to obtain the text embedding vectors corresponding to the entities;
[0084] Step 4-2: Construct a K-Nearest Neighbor Graph:
[0085] Project all entities and their corresponding text embedding vectors into the graph database;
[0086] Set a predefined number of neighbors topk, where topk is the number of relationships of each entity;
[0087] The similarity between entities is measured using cosine similarity, and its calculation method is:
[0088]
[0089] where n represents the dimension of the vector, i represents the index from 1 to n, and x i represents the value of vector x in the i-th dimension, and y i represents the value of vector y in the i-th dimension;
[0090] Establish connections between two entities that meet the preset number of neighbors topk and the similarity preset condition; construct a K-Nearest Neighbor Graph.
[0091] The purpose of entity resolution is to identify and merge records that refer to the same real-world object from different data sources. This helps ensure that each entity is uniquely and accurately represented in a knowledge graph, thereby preventing duplicate or incorrect associations. Set a similarity threshold for filtering the relationships between entities, where relationships are only established between two entities when the similarity between them is greater than or equal to the similarity threshold; where the topk is set to 10 and the similarity threshold is set to 0.95. In this embodiment, cosine similarity is used as the measurement method for the similarity between entities.
[0092] Step 5: Identify weakly connected components: Identify a local K-Nearest Neighbor Graph in which any two entities are mutually reachable in the K-Nearest Neighbor Graph as a weakly connected component; a weakly connected component is defined as any two vertices in the component being mutually reachable when ignoring the directionality of all edges in the graph; consider the entities in the identified weakly connected component as a set of entities that may be similar.
[0093] Step 6: Word distance filtering: The word distance is:
[0094]
[0095] where, is an indicator function that has a value of 0 when a i = b j and 1 otherwise; a and b are two entities in the weakly connected component; i and j are the position numbers in entities a and b respectively;
[0096] Reduce the weakly connected components so that the word distances between entities in the weakly connected components are all less than the preset distance threshold;
[0097] Relying solely on text embeddings cannot accurately distinguish different entities. To improve the accuracy of entity resolution, a word distance filtering condition is added to further screen candidate entities. The word distance used in this embodiment is the edit distance, and entities with an edit distance greater than 3 are filtered from the set of weakly connected components. Other word distances can also be used.
[0098] Step 7: LLM Evaluation: Use LLM to evaluate and decide whether to merge entities in the weakly connected components, merge the entities to be merged and their corresponding relationships, and obtain an updated knowledge graph;
[0099] Input the prompt for LLM to evaluate:
[0100] "You are a professional knowledge graph analyst. Please evaluate the following weakly connected components to determine whether the entities in them should be merged. Please analyze according to the following criteria:
[0101] 1. Similarity: The degree of similarity between entities in terms of name, attributes, or other features.
[0102] 2. Contextual relevance: Whether these entities can be used interchangeably in a specific context.
[0103] 3. Duplication: Whether there are exactly the same or almost the same entities.
[0104] Please evaluate the following weakly connected components, provide a merge suggestion (merge or not merge), and explain the reason. Please be as detailed as possible in the evaluation to ensure sufficient reasons to support your suggestion.
[0105] Set of weakly connected components: {set_info}"
[0106] Among them, set_info represents the information of the set of weakly connected components.
[0107] Step 8: Generate Element Summaries: Use LLM to generate element summaries for each entity and relationship in the updated knowledge graph, and store the element summaries as attributes of the entities; the element summaries can index this information and entities more effectively for more accurate retrieval, and store the element summaries as entity attributes.
[0108] Step 9: Generate graph communities: Use the Leiden algorithm to partition entities into communities and create corresponding community entities; use the Leiden algorithm to divide them into hierarchical communities. The Leiden algorithm is an efficient community detection algorithm for identifying community structures in complex networks. Store the results of the graph communities as entity attributes, and create an independent entity for each community. We generated 97,446 communities. The communities are hierarchical, with different levels from low to high. The higher the level, the fewer the number of communities at the same level and the more node information they contain.
[0109] Step 10: Generate community summaries: Use the LLM to generate community summaries based on community entities and relationships, and store the community summaries as attributes of the community entities; obtain the medical knowledge graph.
[0110] With the help of the community summaries, the global topic structure and semantics of the dataset can be understood. The prompt for the community summary is: "Generate a natural language summary of the provided information based on the provided nodes and relationships belonging to the same graph community. The summary should be concise and clear, covering all key information while ensuring coherence and readability.
[0111] Output requirements:
[0112] 1. Conciseness: The summary should be concise and clear, avoiding redundancy.
[0113] 2. Completeness: Ensure that all important nodes and relationships are covered.
[0114] 3. Coherence: Use natural and fluent language to make the summary easy to understand.
[0115] 4. Emphasis: Highlight the core concepts and important relationships in the community.
[0116] {community_info}
[0117] Summary: ", where community_info represents the triple information within the community.
[0118] Example 2:
[0119] An intelligent retrieval method using the medical knowledge graph described in Technical Solution 1 includes the following steps:
[0120] Step 1: Construct a knowledge graph vector library: Use the Embedding model to convert each entity and attribute of the medical knowledge graph in the medical knowledge graph library constructed in Example 1 into vector representations as knowledge graph vectors, and construct a knowledge graph vector database;
[0121] In this embodiment, the knowledge graph vectors are stored in the Milvus vector database to construct a knowledge graph vector library. Milvus is an open-source vector database developed by Zilliz and is one of the solutions widely used in approximate nearest neighbor search currently.
[0122] Step 2: Construct a text vector library: Use the Embedding model to represent the text blocks in the medical knowledge graph library constructed in Embodiment 1 as text vectors, and construct a text vector library; in this embodiment, the text vectors are stored in the Milvus vector database.
[0123] Step 3: Vectorize the user question entities and the question: Step 3-1: Use the LLM to extract at least one entity entity0 in the user question as the initial root node; use the Embedding model to convert the entity entity0 into an entity vector entity0_vector;
[0124] Step 3-2: Use the Embedding model to convert the user question into a question vector query_vector;
[0125] Step 4: Multilayer entity and relationship retrieval:
[0126] Step 4-1: Determine whether the entity entity0 exists in the medical knowledge graph. If it exists, go to Step 4-2; otherwise, go to Step 4-3;
[0127] Step 4-2: Retrieve the first-layer entities entity1 and their relationships relationship1 that are related to the entity entity0 to obtain at least one top-level triple (entity0, relationship1, entity1); then retrieve the second-layer entities entity2 and relationships relationship2 that are related to each first-layer entity entity1 to obtain at least one first-layer triple (entity1, relationship2, entity2); output the top-level triples, first-layer triples, the first-layer entity entity1 and its element summary, and the second-layer entity entity2 and its element summary; go to Step 5;
[0128] Step 4-3: Find the entities in the medical knowledge graph vector library that reach or exceed the preset semantic relevance threshold for the entity entity0 as similar entities; in this embodiment, the semantic relevance is calculated using cosine similarity, and the preset semantic relevance threshold is 0.85; if there are any, go to Step 4-4; otherwise, go to Step 4-5;
[0129] Step 4-4: Invert the graph community index according to the number of similar entities in the community: Select a preset number of communities, output the communities and community summaries; go to Step 5;
[0130] Step 4-5: Determine it as a global problem, and output each community and its community summary from high to low according to the preset token number limit;
[0131] Step 5: Text retrieval, search for text blocks in the text vector library whose semantic similarity with the problem vector query_vector is greater than the preset semantic similarity threshold as similar text blocks, and sort the similar text blocks from high to low according to the similarity; Output a preset number of similar text blocks;
[0132] Step 6: LLM summary: Construct a prompt text by using the top-level triple, the first-level triple, the first-level entity entity1 and its element summary, the second-level entity entity2 and its element summary, the similar text blocks and the user question according to the preset prompt template, input the prompt text into the LLM, and output the retrieval result; In this embodiment, the first 10 similar text blocks retrieved from Step 5 and the user question are input into the prompt template, and the prompt is input into the LLM large model, and the result is output after being polished by the large model. In the prompt template, it is restricted that answers unrelated to the question are not allowed to be generated to prevent misleading the user.
[0133] The prompt template for generating the summary:
[0134] "
Instruction
Known Information
Question
[0135] Step 7: Output the result, and output the LLM retrieval result and the similar text blocks as retrieval information.
[0136] Embodiment 3: The difference from Embodiment 2 is that it further includes Step 8: User feedback collection: Collect the evaluation of the user on the retrieval information. It provides a feedback mechanism that allows users to evaluate whether the information provided is helpful to them, and records the user questions, the provided answers, and the user feedback to continuously optimize the retrieval algorithm and answer generation strategy.
[0137] The present invention aims to achieve efficient retrieval and intelligent question-answering functions for medical information by integrating GraphRAG and large model technologies, providing accurate and timely information support for medical professionals and researchers. The present invention constructs a medical knowledge graph to capture the complex relationships between entities such as diseases, drugs, and symptoms, and uses GraphRAG technology to effectively retrieve and utilize this knowledge. In addition, by integrating the powerful generation ability of the LLM, it is able to retrieve the most relevant information fragments from existing medical literature based on understanding the user's query and generate high-quality answers.
[0138] The present invention can not only enhance the understanding and processing ability of medical literature, but also support more complex queries, providing more in-depth analysis and insights. This will greatly improve the information retrieval efficiency of medical professionals, help them better cope with the increasing amount of medical information, and ultimately improve the quality and effectiveness of medical services.
Claims
1. An intelligent construction method of a medical knowledge graph based on the combination of GraphRAG and LLM, characterized in that: It includes the following steps: Step 1: Data collection and processing: The data sources include: academic papers, clinical guidelines, drug instructions, health information, and expert interviews; it consists of the following specific steps: Step 1-1: Convert the data into text data in a preset format to form more than one record; Step 1-2: Data cleaning: Calculate the hash value of each record, keep only one record with the same hash value, remove the useless information in the retained records, and the useless information includes special symbols, garbled characters, headers, footers, and dates, and convert the retained records into corresponding medical texts; Step 2: Split the text: Split each medical text using full stops, semicolons, question marks, exclamation marks, and line breaks as delimiters to obtain more than 1 text block; the number of tokens in each text block is less than the preset value of the text block; the number of overlapping tokens between two adjacent text blocks is the preset number of overlapping tokens; Step 3: Extract entities and relationships: Use the LLM to extract entity-relationship triples, and at the same time merge the same entity-relationship triples from different data sources to build a graph database; Step 4: Entity resolution: It consists of the following specific steps: Step 4-1: Obtain the entities and descriptive attributes included in the triples in the graph database, and obtain the text embedding vectors corresponding to the entities; Step 4-2: Build a K-nearest neighbor graph: Project all entities and their corresponding text embedding vectors into the graph database; Set the predefined number of neighbors topk, and topk is the number of relationships of each entity; The similarity between entities is measured using cosine similarity, and its calculation method is: where n represents the dimension of the vector, i represents the index from 1 to n, and x i represents the value of the vector x in the i-th dimension, and y i represents the value of the vector y in the i-th dimension; Establish a connection between two entities that meet the preset number of neighbors topk and the similarity preset condition; build a K-nearest neighbor graph Step 5: Identify weakly connected components: Identify at least one weakly connected component in the K-nearest neighbor graph, and there is a relationship between any two entities in the weakly connected component; Step 6: Word distance filtering: The word distance is: Among them, is an indicator function, when a i = b j the value is 0, otherwise it is 1; a and b are two entities of the weakly connected component; i and j are the position numbers in entities a and b respectively; Shrink the weakly connected components so that the word distances of the entities in the weakly connected components are all less than the preset distance threshold; Step 7: LLM evaluation: Use the LLM to evaluate and decide whether to merge the entities in the weakly connected components, and merge the entities to be merged and their corresponding relationships to obtain an updated knowledge graph; Step 8: Generate element summaries: Use the LLM to generate element summaries for each entity and relationship in the updated knowledge graph, and store the element summaries as attributes of the entities; Step 9: Generate graph communities: Use the Leiden algorithm to hierarchically partition the entities into communities, and establish corresponding entities for each hierarchical community; Step 10: Generate community summaries: Use the LLM to generate community summaries based on the community entities and relationships, and store the community summaries as attributes of the community entities; obtain the medical knowledge graph.
2. An information retrieval method for a medical knowledge graph constructed according to the method described in claim 1, characterized in that: It includes the following steps: Step 1: Build a knowledge graph vector library: Use the method of claim 1 to build a medical knowledge graph, and use the Embedding model to convert each entity and attribute in the medical knowledge graph into a vector representation as a knowledge graph vector, and build a knowledge graph vector database; Step 2: Construct a text vector library: Use an Embedding model to represent the text blocks obtained during the construction of the knowledge graph library as text vectors, and construct a text vector library; Step 3: Vectorize user question entities and questions: Step 3-1: Use LLM to extract at least one entity entity0 in the user question as the initial root node; Use an Embedding model to convert entity entity0 into an entity vector entity0_vector; Step 3-2: Use an Embedding model to convert the user question into a question vector query_vector; Step 4: Multilayer entity and relationship retrieval: Step 4-1: Determine whether entity entity0 exists in the medical knowledge graph. If it exists, go to Step 4-2; otherwise, go to Step 4-3; Step 4-2: Retrieve the first-layer entities entity1 and their relationships relationship1 that are related to entity entity0, and obtain at least one top-level triple (entity0, relationship1, entity1); Then retrieve the second-layer entities entity2 and relationships relationship2 that are related to each first-layer entity entity1, and obtain at least one first-layer triple (entity1, relationship2, entity2); Output the top-level triples, first-layer triples, first-layer entity entity1 and its element summary, and second-layer entity entity2 and its element summary; Go to Step 5; Step 4-3: Find the entities in the medical knowledge graph vector library that reach or exceed the preset semantic relevance threshold for entity entity0 as similar entities; If there are any, go to Step 4-4; Otherwise, go to Step 4-5; Step 4-4: Sort the graph communities: Sort according to the number of similar entities in the community from most to least, and select a preset number of communities from most to least, and output the communities and community summaries; Go to Step 5; Step 4-5: Determine it as a global problem, and output each community and its community summary level by level according to the preset token number limit; Step 5: Text retrieval, find the text blocks in the text vector library whose semantic similarity with the question vector query_vector is greater than the preset semantic similarity threshold as similar text blocks, and sort the similar text blocks from highest to lowest similarity; Output a preset number of similar text blocks; Step 6: LLM summary: Construct a prompt text from the top-level triples, first-layer triples, first-layer entity entity1 and its element summary, second-layer entity entity2 and its element summary, similar text blocks, and user questions according to the preset prompt template, and input the prompt text into LLM to output the retrieval result; Step 7: Output the result, and output the LLM retrieval result and similar text blocks as retrieval information.
3. The information retrieval method according to claim 2, wherein: It also includes Step 8: User feedback collection: Collect the user's evaluation of the retrieval information.
4. The information retrieval method according to claim 2, wherein: In step 6, the prompt template is: [Instruction] Answer the question professionally based on the known information; if the answer cannot be obtained from it, say "The question cannot be answered based on the known information", and do not allow fabrications in the answer. Provide corresponding explanations for the answer, and the answer should be in Chinese;\n\n[Known Information] {context}\n\n[Question] {question}"; where context is the retrieved similar triples, element summaries of entities, summaries of communities, and similar text blocks, and question is the question input by the user.
Citation Information
Patent Citations
Domain knowledge graph-oriented entity alignment method
CN116578654A
Method and system for realizing entity alignment by applying credibility perception iterative training strategy
CN118364906A