A Retrieval-Augmented Generation Method and Device Based on a Semantic-Enhanced Knowledge Graph
By constructing a semantic enhancement knowledge graph, extracting entity information and relationship information, the problem of inconsistent content generation of large language models is solved, more accurate question-and-answer system answers are achieved, and the semantics and detailed information of the knowledge graph are enhanced.
Patent Information
- Application Number
- CN202510398457.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-01
AI Technical Summary
Large language models have hallucinations when generating content. Traditional retrieval enhancement generation methods cannot capture the relationship information between texts, resulting in low accuracy and redundant answers. Graph RAG method lacks generalization ability and detailed information.
Construct a semantic enhanced knowledge graph, extract entity information, inter-entity relationships and related sentence information, use a large language model to judge whether the sub-graph can answer questions, generate corresponding answers, or extract relevant entity information from a mixed set of related sentences to enhance context information.
It improves the generalization ability and accuracy of the question-and-answer system, provides rich structured relationships and semantic information, and can generate more accurate answers.
Smart Images

Figure CN119917636B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of question - answering systems in human - computer interaction conversations, and particularly to a retrieval - augmented generation method and device based on a semantically enhanced knowledge graph. Background Art
[0002] Large Language Models (LLMs) have demonstrated powerful natural language understanding and human - like text generation capabilities in practical applications, but still face the hallucination problem, that is, the generated content does not conform to common sense or is inconsistent with the source content.
[0003] Retrieval - Augmented Generation (RAG) helps LLMs obtain external knowledge by retrieving relevant information from an external text corpus, thereby improving the accuracy of responses. However, RAG has the following limitations: (1) Ignoring relationship information: Traditional RAG cannot capture the relationship information between texts, resulting in the inability to process knowledge with a graph structure. (2) Redundant information: RAG represents external knowledge using text fragments, which may exceed the input length limit of LLMs or cause information loss.
[0004] To solve these problems, in recent years, some research has proposed using Knowledge Graphs (KGs) as external knowledge sources. Different from traditional RAG, the Graph Retrieval - Augmented Generation (GraphRAG) process retrieves sub - graphs containing relationship information from the knowledge graph as context information. GraphRAG can handle the relationships between texts and improve the reasoning ability, but still has the defects of lack of generalization ability and low answer accuracy. Summary of the Invention
[0005] The purpose of this application is to provide a retrieval - augmented generation method and device based on a semantically enhanced knowledge graph, which can improve the generalization ability and the accuracy of answers.
[0006] To achieve the above - mentioned purpose, this application provides the following solutions:
[0007] In the first aspect, this application provides a retrieval - augmented generation method based on a semantically enhanced knowledge graph, including:
[0008] Obtain the original text document;
[0009] Construct a semantic-enhanced knowledge graph based on the entity information and the relationship information between entities extracted from the original text document. The semantic-enhanced knowledge graph includes an entity set, a relationship set between entities, a relevant sentence set, and a mapping relationship set. The relevant sentence set refers to the set of sentences related to entities and the relationships between entities. The mapping relationship set refers to the set of mapping relationships between the sentences in the original text document and the target objects, where the target objects are the entity nodes and relationship edges in the knowledge graph;
[0010] Obtain a subgraph based on the explicit entity set and the semantic-enhanced knowledge graph;
[0011] Use a large language model to determine whether the entity information and the relationship information between entities in the subgraph can answer the user's question, and obtain a first judgment result;
[0012] If the first judgment result is yes, use the entity information and the relationship information between entities in the subgraph as the first context information, and construct a first prompt message according to the first context information and the user's question;
[0013] Based on the first prompt message, use a large language model to generate an answer corresponding to the user's question;
[0014] If the first judgment result is no, extract relevant entity information from the mixed relevant sentence set, and use the relevant entity information, the entity information and the relationship information between entities in the subgraph as the second context information, and construct a second prompt message according to the second context information and the user's question. The mixed relevant sentence set is the set of sentences related to the mixed entities and the relationships between the mixed entities in the subgraph;
[0015] Based on the second prompt message, use a large language model to generate an answer corresponding to the user's question.
[0016] In a second aspect, the present application provides a retrieval-enhanced generation device based on a semantic-enhanced knowledge graph, including:
[0017] A document acquisition module for acquiring an original text document;
[0018] A semantic-enhanced knowledge graph construction module for constructing a semantic-enhanced knowledge graph according to the entity information and the relationship information between entities extracted from the original text document. The semantic-enhanced knowledge graph includes an entity set, a relationship set between entities, a relevant sentence set, and a mapping relationship set. The relevant sentence set refers to the set of sentences related to entities and the relationships between entities. The mapping relationship set refers to the set of mapping relationships between the sentences in the original text document and the target objects, where the target objects are the entity nodes and relationship edges in the knowledge graph;
[0019] A sub - graph construction module, configured to obtain a sub - graph according to the explicit entity set and the semantic - enhanced knowledge graph;
[0020] A judgment module, configured to use a large - language model to judge whether the entity information and the relationship information between entities in the sub - graph can answer the user's question, and obtain a first judgment result; if the first judgment result is yes, use the entity information and the relationship information between entities in the sub - graph as the first context information, construct a first prompt information according to the first context information and the user's question; based on the first prompt information, use the large - language model to generate an answer corresponding to the user's question; if the first judgment result is no, extract relevant entity information from the mixed relevant sentence set, use the relevant entity information, the entity information and the relationship information between entities in the sub - graph as the second context information, construct a second prompt information according to the second context information and the user's question; based on the second prompt information, use the large - language model to generate an answer corresponding to the user's question, where the mixed relevant sentence set is a sentence set related to the mixed entities and the relationship between mixed entities in the sub - graph.
[0021] According to the specific embodiments provided by the present application, the present application has the following technical effects:
[0022] The present application provides a retrieval - enhanced generation method and device based on a semantic - enhanced knowledge graph. The method includes constructing a semantic - enhanced knowledge graph, performing retrieval based on the retrieval strategy of the semantic - enhanced knowledge graph, and the retrieval result includes the first context information or the second context information. The first context information is determined according to the entity information and the relationship information between entities in the sub - graph, and the second context information is determined according to the entity information, the relationship information between entities in the sub - graph and the relevant entity information. An answer is obtained according to the first context information or the second context information. This method enhances the detailed information of entities and the semantic information of the knowledge graph, and can generate more accurate answers; and because this method extracts the specific information (context information) required by specific entities according to the requirements, rather than extracting the specified entity information during the construction of the knowledge graph, this method has generality. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic flowchart of a retrieval - enhanced generation method based on a semantic - enhanced knowledge graph provided in Embodiment 1 of the present application;
[0024] Figure 2 It is a schematic diagram of the semantic - enhanced knowledge graph in Embodiment 1 of the present application;
[0025] Figure 3 It is a schematic flowchart when applying a specific example of a retrieval - enhanced generation method based on a semantic - enhanced knowledge graph in Embodiment 1 of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] Example 1
[0027] It has been found through research that the Graph RAG method still faces the following challenges:
[0028] (1) Lack of generalization ability: Most existing Graph RAG methods are built based on existing knowledge graphs. However, constructing a high-quality knowledge graph requires profound expertise and a large amount of time investment, making this process costly. In addition, it is necessary to design an architecture suitable for a specific application scenario before constructing the knowledge graph. However, this predefinition also limits the information in the knowledge graph.
[0029] (2) Lack of details: In the actual application of Graph RAG, the knowledge graph usually only contains entities and their relationships, ignoring specific information about the entities or evidence supporting the relationships. Compared with using text corpora as external knowledge sources, although the knowledge graph provides powerful relationship information, it is extremely lacking in terms of details.
[0030] To address the above defects, this example provides a retrieval-augmented generation method based on a semantically enhanced knowledge graph, which can improve the generalization ability and the accuracy of answers. As Figure 1 shown, this method includes the following steps 201 to step 208. Among them:
[0031] Step 201, obtain the original text document.
[0032] Step 202, extract entity information and relationship information between entities from the original text document and construct a semantically enhanced knowledge graph according to the extracted entity information and relationship information between entities , where the semantically enhanced knowledge graph includes an entity set, a relationship set between entities, a relevant sentence set, and a mapping relationship set. The mapping relationship set refers to the set containing the mapping relationship between the sentences in the original text document and the target objects, and the target objects are entity nodes and relationship edges in the knowledge graph. is the entity set, is the entity set is the relationship set between entities in is the set of sentences related to entities and relationships between entities in, that is, the relevant sentence set, is the mapping relationship set between the sentences in the text and the relevant entity nodes and relationship edges on the knowledge graph, that is, the mapping relationship set.
[0033] The process of extracting entity information and relationship information between entities in step 202 (i.e., step 202-1) is as follows:
[0034] Step 202-1-1, for the original text document provided by the user , perform document segmentation respectively to obtain several sets of text blocks .
[0035] Step 202-1-2, use a language model to perform sentence segmentation on the original text document to obtain a set of sentences containing several sentences ;
[0036] Step 202-1-3, use a tokenization model to calculate the number of tokens for each sentence in the set of sentences (i.e., ), and record ;
[0037] Step 202-1-4, divide consecutive sentences whose total number of tokens does not exceed the set text block size into one text block;
[0038] Step 202-1-5, construct a set of text blocks according to all the text blocks .
[0039] Step 202-1-6, extract the entity information and the relationship information between entities from the set of text blocks
[0040] In step 202, construct a semantically enhanced knowledge graph according to the extracted entity information and the relationship information between entities (i.e., step 202-2), specifically including:
[0041] Step 202-2-1, construct a text block semantically enhanced knowledge graph according to the entity information and the relationship information between entities extracted from each text block (i.e., ) ;
[0042] Step 202-2-2, merge all the text block semantically enhanced knowledge graphs to obtain the semantically enhanced knowledge graph of the original text document ; ;
[0043] Step 202-2-1 specifically includes:
[0044] (1) Use a large language model to extract entity information and entity relationship information of a specified entity type from each text block , where the entity information includes entity name, entity type and entity description, and the entity relationship information includes source entity, target entity, relationship description and relationship basis;
[0045] (2)Construct a set of text block entities and a set of relationships between text block entities based on the entity information and entity relationship information of the specified entity type extracted, where the set of relationships between text block entities contains relationships between several entities in the set of text block entities Among them, the set of text block entities is denoted as , and the set of relationships between text block entities is denoted as ;
[0046] (3)Use a large language model to evaluate the information content score of each sentence in each text block, denoted as ;
[0047] (4)When the information content score is not less than the information threshold , and the entities appearing in the sentence belong to the set of text block entities (i.e., ), establish a mapping relationship between the sentence and the entities appearing in the sentence (i.e., map the sentence to );
[0048] (5)Construct a set of mapping relationships between the text content of the text block and the entities according to the mapping relationship between the sentence and the entities appearing in the sentence ;
[0049] (6)When the information content score is not less than the information threshold and the sentence comes from the relationships between entities in the set of relationships between text block entities (i.e., the sentence is the source of the relationship ), establish a mapping relationship between the sentence and the relationships between entities in the set of relationships between text block entities (i.e., map the sentence to the relationship );
[0050] (7)Construct a set of mapping relationships between the text content of the text block and the relationships between entities according to the mapping relationship between the sentence and the relationships between entities in the set of relationships between text block entities ;
[0051] (8)Determine the set of text block mapping relationships and the set of sentences related to the text block ;
[0052] (9) Construct the text block semantic enhancement knowledge graph according to the text block entity set, the relationship set between text block entities, the relevant sentence set of text blocks, and the mapping relationship set of text blocks. .
[0053] Step 203: Perform entity alignment processing on the semantic enhancement knowledge graph obtained in Step 202 to obtain an aligned semantic enhancement knowledge graph, so as to improve the normativity and consistency of the semantic enhancement knowledge graph.
[0054] Step 203 specifically includes:
[0055] Step 203-1: Pre-classify the entities in the entity set according to the entity type to obtain an entity type set containing several different types of entities , where represents the th category;
[0056] Step 203-2: Cluster the entities of the same type in the entity type set ( ) based on semantic similarity to obtain a clustered entity set;
[0057] The specific process of Step 203-2 is as follows:
[0058] Step 203-2-1: Use an embedding model to vectorize the entity names and corresponding entity descriptions of the entities of the same type in the entity type set to obtain each entity vector, that is, convert "<entity name>:<entity description>" into a vector form as the vector representation of;
[0059] Step 203-2-2: Calculate the second similarity according to the entity vectors;
[0060] Step 203-2-3: Add association edges to the entities with the second similarity higher than the second similarity threshold to construct an entity relationship graph , where ;
[0061] Step 203-2-4: Extract all connected subgraphs in the entity relationship graph ;
[0062] Step 203-2-5: Construct a clustered entity set according to the nodes included in the connected subgraph (the nodes included in each connected subgraph are a new entity set).
[0063] Step 203-3: Use a large language model to determine whether the clustered entities in the clustered entity set are the same entity, and obtain a second judgment result;
[0064] Step 203-4: If the second judgment result is yes, then merge the entity nodes in the semantic enhanced knowledge graph that belong to the same entity.
[0065] Step 204: For the user question , use a large language model to extract the entities that appear in the user question, and obtain an explicit entity set containing several explicit entities (i.e., the entities that appear in ).
[0066] Step 205: According to the explicit entity set obtained in Step 204 , perform subgraph expansion on the semantic enhanced knowledge graph , and obtain a mixed entity set containing the explicit entities and implicit entities , , and extract a subgraph from the aligned semantic enhanced knowledge graph according to the mixed entity set , where the implicit entity refers to the entity involved in the user question but not appearing in the user question .
[0067] The subgraph expansion method is specifically as follows:
[0068] Step 205-1: Add the explicit entities in the explicit entity set to the initial entity list ;
[0069] Step 205-2: Use an embedding model to vectorize the user question to obtain a question vector , and use the embedding model to vectorize the text, denoted as ;
[0070] Step 205-3: For each explicit entity in the initial list, traverse the semantic enhanced knowledge graph to determine the neighbor nodes, that is, sequentially take , and traverse all the neighbor nodes on the semantic enhanced knowledge graph ;
[0071] Step 205-4: Use an embedding model to vectorize the neighbor relationship description to obtain a neighbor relationship description vector , where the neighbor relationship description refers to the relationship description between neighbor nodes not in the initial entity list and the corresponding explicit entity ;
[0072] Step 205-5, calculate the first similarity between the problem vector and the neighbor relationship description vector;
[0073] Step 205-6, if the first similarity is greater than the first similarity threshold , then add the corresponding neighbor node to the initial entity list to obtain an updated entity list;
[0074] Step 205-7, use the updated entity list as the new initial entity list, return to Step 205-3, and obtain the final entity list;
[0075] Step 205-8, after the expansion ends, obtain a hybrid entity set including the explicit entity and the implicit entity according to the final entity list, that is .
[0076] Step 206, use the large language model to judge whether the entity information and the entity relationship information in the subgraph can answer the user's question , and obtain a first judgment result.
[0077] Step 207, if the first judgment result is yes, then use the entity information and the entity relationship information in the subgraph as the first context information, construct a first prompt message according to the first context information and the user's question, and based on the first prompt message, use the large language model to generate a final response , that is, generate an answer corresponding to the user's question.
[0078] Step 208, if the first judgment result is no, then prompt the large language model to judge that other relevant entity information is required in addition to the entity information and the entity relationship information in the subgraph, and perform in-depth extraction. Specifically, extract relevant entity information from the hybrid relevant sentence set , use the relevant entity information, the entity information and the entity relationship information in the subgraph as the second context information, construct a second prompt message according to the second context information and the user's question, and based on the second prompt message, use the large language model to generate a final response , that is, generate an answer corresponding to the user's question, where the hybrid relevant sentence set is the sentence set in the subgraph related to the hybrid entity and the hybrid entity relationship.
[0079] Extract relevant entity information from the set of mixed relevant sentences (for the required entity information , retrieve relevant sentences from the set of mixed relevant sentences . Specifically, it includes:
[0080] (1) Use an embedding model to vectorize the text of the required entity information and each sentence in the set of mixed relevant sentences to obtain the required entity information vector and the sentence vector , where the required entity information refers to the entity information in the sub-graph for which the first judgment result is negative;
[0081] (2) Calculate the semantic similarity between the required entity vector and the sentence vector ;
[0082] (3) Retrieve the set of target sentences with a semantic similarity greater than the threshold , that is, construct a set of target sentences according to the sentences corresponding to the semantic similarity greater than the semantic similarity threshold ;
[0083] (4) Use a large language model to extract relevant entity information from the set of target sentences .
[0084] The following takes the retrieval enhancement generation of the question "At what age did so-and-so publish ResNet" as an example to specifically explain the graph enhancement retrieval method provided in this embodiment.
[0085] Before performing retrieval enhancement generation, it is necessary to first construct a semantic enhancement knowledge graph on several documents provided by the user, and its method flow is as Figure 2 shown.
[0086] As Figure 3 shown, this embodiment provides a specific example of a graph enhancement retrieval method, including:
[0087] Step 1, use the "en_core_web_sm" language model to segment the sentences of the document , use the "cl100k_base" model as the tokenization model, and calculate the number of tokens for each sentence. When performing text segmentation, set the text block size to , and divide several consecutive sentences with a total number of tokens not exceeding into a text block to obtain .
[0088] Step 2, for each text block , prompt "gpt-4o-mini" to extract entity information of the specified entity type, including entity name, entity type, and entity description. The specific prompt for entity extraction is: "Given a text document and a list of entity types, identify all entities. For each identified entity, extract the following information:\n- Entity name: the name of the entity in the text\n- Entity type: one of the following types: {entity type}\n- Entity description: a comprehensive description of the entity\n\nText:\n{input text}\nOutput: ". After extracting the entity information, continue to prompt "gpt-4o-min" to extract the relationships between the extracted entities. The prompt for relationship extraction is: "Given a text document and a list of entities, identify all pairs of (source entity, target entity) that are clearly related to each other. For each pair of related entities, extract the following information:\n- Source entity\n- Target entity\n- Relationship description: a description of the relationship between the source entity and the target entity\n- Citation: the sentence in the text that proves this relationship\n\nEntities:\n{entity}\nText:\n{input text}\nOutput: ". Parse the output results of the model to obtain and . Prompt the "gpt-4o-mini" model to evaluate the information content of the sentences in . The evaluation prompt is: "Your task is to score each text according to the information content in the given list. Texts with more detailed information score higher. The score range is from 0 to 10 (including 0 and 10). Return in JSON format\nText:\n{input_text}\nOutput: ". For sentences where the score is greater than or equal to the threshold , if an entity appears , then map it to . If is the support for the relationship between entity and , then map to to obtain the mapping relationships of the sentence set, entity set and the relationship set . From this, construct the semantic-enhanced knowledge graph on . Merge to get the semantic-enhanced knowledge graph on .
[0089] Step 3: For the set of entity nodes , pre-classify according to entity categories, use the "text-embedding-3-small" model as the embedding model, vectorize "<entity name>:<entity description>" as the vector representation of the entity, and calculate the cosine similarity between entities of the same category: for the n-dimensional vector , the cosine similarity calculation formula is as follows:
[0090]
[0091] where, respectively represent the n-dimensional vectors, respectively represent the vectors each dimension of data.
[0092] Add association edges to the entities with similarity higher than the threshold to construct an entity relationship graph , where . Prompt the "gpt-4o-mini" model to judge the same entities in the set of entity nodes of all connected subgraphs in the relationship graph . The judgment prompt is "According to the entity description, group all entities in the given list that refer to the same real-world entity into one group.\nEntity list:\n{entity}\nOutput: ". Merge the same entities on the semantic enhanced knowledge graph .
[0093] Step 4: For the question "At what age did so-and-so publish ResNet", prompt the "gpt-4o-mini" model to extract all explicit entities , and the prompt for extracting entities is "Extract entities from the given text and return them as a JSON list as follows:\nText:\n{input text}\nOutput: ". Extract the entities "so-and-so" and "ResNet".
[0094] Step 5: Perform subgraph expansion based on the set of explicit entities extracted in Step 4 , maintain the entity list , add the entities in the set of explicit entities to the entity list , for the entities in the entity list , traverse their neighbor nodes on the semantic enhanced knowledge graph , use the "text-embedding-3-small" model as the embedding model, vectorize the user question , get and the relationship information between the entity and its neighbors , get , calculate and The cosine similarity. If the similarity is greater than the threshold , then add the neighbor to the entity list . Repeat this process until all entities in the entity list have been visited. Thus, an entity set containing explicit entities and implicit entities is obtained , and according to the entity set , extract a subgraph from the semantic enhanced knowledge graph . .
[0095] Step 6. Prompt "gpt-4o-mini" to determine whether the entity and relationship information provided by the subgraph can answer the user's question . The judgment prompt is "Please determine whether the given knowledge graph can support answering the question. If it can, output 'Yes'; otherwise, output 'No'.\nStandard: The knowledge graph must contain all the information required to answer the question completely.\nKnowledge graph:\n{Knowledge graph}\nQuestion:\n{Question}\nOutput:". If it can answer, then assemble the entity and relationship information as the context and the user's question together into a prompt for "gpt-4o-mini" to finally respond .
[0096] Step 7. Prompt "gpt-4o-mini" to determine what additional entity information is needed besides the entity relationship information provided in Step 6. The judgment prompt is "What specific information of which entities is needed to answer this question?\nOutput:". Analyze the output result of the model. The information of "date of birth" of the entity "so and so" and the information of "publication date" of the entity "ResNet" are needed. Retrieve the sentences related to the specific information from the sentence set related to the entity . Use "text-embedding-3-small" as the embedding model to vectorize the text, and vectorize the required entity information and each sentence in the sentence set respectively. Calculate the cosine similarity between and each sentence in the sentence set , and retrieve the sentences with a similarity greater than the threshold sentence. Prompt "gpt-4o-mini" extracts specific information from the retrieved sentences. The extraction prompt is "Extract the information of the entity according to the provided context, and only output the extracted information.\nEntity: {entity}\nInformation: {attribute}\nContext:\n{context}\nOutput: ". The specific information "date of birth" of the entity "so-and-so" is "so-and-so was born in 1984", and the specific information "publication date" of the entity "ResNet" is "ResNet was published in 2015".
[0097] Step 8, combine the entity and relationship information provided in Step 6 and the specific information of the entity extracted in Step 7 as the context and the user's question and assemble them into a prompt to make the final response of "gpt-4o-mini" 。
[0098] A retrieval-enhanced generation method involved in this embodiment. First, for the original document, use a large language model to extract rich structured relationship information and semantic information, and construct a semantic-enhanced knowledge graph; then, based on the semantic-enhanced knowledge graph, propose an efficient retrieval-enhanced generation method to answer the user's question. The method includes:
[0099] 1) For the original text document, perform text segmentation, use a large language model for information extraction, and construct a semantic-enhanced knowledge graph;
[0100] 2) Based on the semantic-enhanced knowledge graph, propose an efficient retrieval-enhanced generation method, extract explicit entities according to the user's question, retrieve implicit entities on the semantic-enhanced knowledge graph, use a large language model to extract relevant entity information from the semantic-enhanced knowledge graph, and then return the answer by assembling the prompt words and calling the large model.
[0101] Compared with the traditional knowledge graph, the above retrieval-enhanced generation method provided in this embodiment can provide richer structured relationship information and semantic information (construct the relationship between the knowledge graph and the original document, enhance the detailed information of the entity and the semantic information of the knowledge graph), effectively improve the accuracy of question answering; propose a deep extraction scheme, extract the specific information required by a specific entity according to the demand, rather than extracting the specified entity information during the construction of the knowledge graph, making the method of this embodiment universal.
[0102] Embodiment 2
[0103] This embodiment provides a retrieval-enhanced generation device based on a semantic-enhanced knowledge graph, including:
[0104] A document acquisition module for acquiring the original text document;
[0105] The semantic enhancement knowledge graph construction module is used to construct a semantic enhancement knowledge graph according to the entity information and the relationship information between entities extracted from the original text document. Among them, the semantic enhancement knowledge graph includes an entity set, a relationship set between entities, a related sentence set, and a mapping relationship set. The related sentence set refers to the set of sentences related to entities and the relationships between entities. The mapping relationship set refers to the set containing the mapping relationship between the sentences in the original text document and the target objects, and the target objects are the entity nodes and relationship edges in the knowledge graph;
[0106] The sub-graph construction module is used to obtain a sub-graph according to the explicit entity set and the semantic enhancement knowledge graph;
[0107] The judgment module is used to use a large language model to judge whether the entity information and the relationship information between entities in the sub-graph can answer the user's question, and obtain a first judgment result; if the first judgment result is yes, the entity information and the relationship information between entities in the sub-graph are used as the first context information, and a first prompt information is constructed according to the first context information and the user's question; based on the first prompt information, a large language model is used to generate an answer corresponding to the user's question; if the first judgment result is no, relevant entity information is extracted from the mixed relevant sentence set, and the relevant entity information, the entity information and the relationship information between entities in the sub-graph are used as the second context information, and a second prompt information is constructed according to the second context information and the user's question; based on the second prompt information, a large language model is used to generate an answer corresponding to the user's question, where the mixed relevant sentence set is the set of sentences related to the mixed entities and the relationships between the mixed entities in the sub-graph.
[0108] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0109] In this article, specific examples are used to elaborate on the principles and implementation methods of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation methods and application scopes. To sum up, the content of this specification should not be construed as a limitation to this application.
Claims
1. A retrieval-augmented generation method based on a semantically enhanced knowledge graph, characterized in that The retrieval enhanced generation method based on the semantic enhanced knowledge graph includes: Obtain the original text document; Construct a semantic enhanced knowledge graph according to the entity information and the relationship information between entities extracted from the original text document, wherein the semantic enhanced knowledge graph includes an entity set, a relationship set between entities, a related sentence set, and a mapping relationship set. The related sentence set refers to the set of sentences related to entities and the relationships between entities. The mapping relationship set refers to the set containing the mapping relationship between the sentences in the original text document and the target objects, and the target objects are entity nodes and relationship edges in the knowledge graph; Perform entity alignment processing on the semantic enhanced knowledge graph in sequence to obtain the aligned semantic enhanced knowledge graph; Use a large language model to extract the entities appearing in the user's question to obtain an explicit entity set containing several explicit entities; According to the explicit entity set, perform subgraph expansion on the aligned semantic enhanced knowledge graph to obtain a mixed entity set containing the explicit entities and implicit entities, wherein the implicit entities refer to the entities involved in the user's question but not appearing in the user's question; Extract a subgraph from the aligned semantic enhanced knowledge graph according to the mixed entity set; Use a large language model to determine whether the entity information and the relationship information between entities in the subgraph can answer the user's question to obtain a first judgment result; If the first judgment result is yes, use the entity information and the relationship information between entities in the subgraph as the first context information, and construct a first prompt information according to the first context information and the user's question; Based on the first prompt information, use a large language model to generate an answer corresponding to the user's question; If the first judgment result is no, extract relevant entity information from the mixed relevant sentence set, use the relevant entity information and the entity information and the relationship information between entities in the subgraph as the second context information, and construct a second prompt information according to the second context information and the user's question, wherein the mixed relevant sentence set is the set of sentences related to the mixed entities and the relationships between the mixed entities in the subgraph; Based on the second prompt information, use a large language model to generate an answer corresponding to the user's question.
2. The retrieval enhancement generation method based on a semantically enhanced knowledge graph according to claim 1, wherein The process of extracting entity information and relationship information between entities from the original text document specifically includes: Perform document segmentation on the original text document to obtain a text block set containing several text blocks; Extract the entity information and the relationship information between entities from the text block set; The construction process of the text block set specifically includes: Use a language model to perform sentence segmentation on the original text document to obtain a sentence set containing several sentences; Use a word segmentation model to calculate the Token number of each sentence in the sentence set; Divide consecutive sentences with a total Token number not exceeding the set text block size into one text block; Construct a text block set according to all the text blocks.
3. The retrieval enhancement generation method based on a semantically enhanced knowledge graph according to claim 2, characterized in that, Construct a semantic enhanced knowledge graph according to the entity information and the relationship information between entities extracted from the original text document, specifically including: Construct a text block semantic enhancement knowledge graph based on each text block in the set of text blocks; Merge all the text block semantic enhancement knowledge graphs to obtain the semantic enhancement knowledge graph of the original text document; Constructing a text block semantic enhancement knowledge graph according to each text block in the set of text blocks specifically includes: Use a large language model to extract entity information and entity relationship information of a specified entity type from each text block, where the entity information includes entity name, entity type, and entity description, and the entity relationship information includes source entity, target entity, relationship description, and relationship basis; Construct a text block entity set and a text block entity relationship set according to the extracted entity information and entity relationship information of the specified entity type, where the text block entity relationship set contains several entity relationships among the entities in the text block entity set; Use a large language model to evaluate the information quantity score of each sentence in the text block; When the information quantity score is not less than the information quantity threshold and the entities appearing in the sentence belong to the text block entity set, establish a mapping relationship between the sentence and the entities appearing in the sentence; Construct a mapping relationship set between the text content of the text block and the entity according to the mapping relationship between the sentence and the entities appearing in the sentence; When the information quantity score is not less than the information quantity threshold and the sentence comes from the entity relationship in the text block entity relationship set, establish a mapping relationship between the sentence and the entity relationship in the text block entity relationship set; Construct a mapping relationship set between the text content of the text block and the entity relationship according to the mapping relationship between the sentence and the entity relationship in the text block entity relationship set; Determine a text block mapping relationship set and a text block related sentence set according to the mapping relationship set between the text content of the text block and the entity and the mapping relationship set between the text content of the text block and the entity relationship; Construct the text block semantic enhancement knowledge graph according to the text block entity set, the text block entity relationship set, the text block related sentence set, and the text block mapping relationship set; 4. The retrieval enhancement generation method based on a semantically enhanced knowledge graph according to claim 3, wherein Perform entity alignment and entity merging processing on the semantic enhancement knowledge graph in sequence to obtain an aligned semantic enhancement knowledge graph, which specifically includes: Pre-classify entities according to the entity type to obtain an entity type set containing several different types of entities; Cluster entities of the same type in the entity type set based on semantic similarity to obtain a clustered entity set; Use a large language model to judge whether the clustered entities in the clustered entity set are the same entity to obtain a second judgment result; If the second judgment result is yes, merge the entity nodes belonging to the same entity in the semantic enhancement knowledge graph; 5. The retrieval enhancement generation method based on a semantically enhanced knowledge graph according to claim 1, wherein Perform subgraph expansion on the aligned semantic enhancement knowledge graph according to the explicit entity set to obtain a mixed entity set containing the explicit entity and the implicit entity, which specifically includes: Add the explicit entities in the explicit entity set to the initial entity list; Use an embedding model to vectorize the user question to obtain a question vector; For each explicit entity in the initial entity list, traverse the semantic-enhanced knowledge graph to determine neighbor nodes; Use an embedding model to vectorize the neighbor relationship description into a neighbor relationship description vector, where the neighbor relationship description refers to the relationship description between neighbor nodes not in the initial entity list and the corresponding explicit entity; Calculate the first similarity between the problem vector and the neighbor relationship description vector; If the first similarity is greater than the first similarity threshold, add the corresponding neighbor node to the initial entity list to obtain an updated entity list; Use the updated entity list as the new initial entity list, and return to the step "For each explicit entity in the initial entity list, traverse the semantic-enhanced knowledge graph to determine neighbor nodes" to obtain a final entity list; Obtain a hybrid entity set including the explicit entity and the implicit entity according to the final entity list; 6. The retrieval enhancement generation method based on a semantically enhanced knowledge graph according to claim 1, wherein Extract relevant entity information from the hybrid relevant sentence set, specifically including: Use an embedding model to vectorize the required entity information and each sentence in the hybrid relevant sentence set into a required entity information vector and a sentence vector, where the required entity information refers to the entity information in the subgraph for which the first judgment result is negative; Calculate the semantic similarity between the required entity information vector and the sentence vector; Construct a target sentence set according to the sentences corresponding to the semantic similarity greater than the semantic similarity threshold; Use a large language model to extract relevant entity information from the target sentence set; 7. The retrieval enhancement generation method based on a semantically enhanced knowledge graph according to claim 4, wherein Cluster entities of the same type in the entity type set based on semantic similarity to obtain a clustered entity set, specifically including: Use an embedding model to vectorize the entity name and the corresponding entity description of entities of the same type in the entity type set into each entity vector; Calculate the second similarity according to the entity vectors; Add association edges to entities with the second similarity higher than the second similarity threshold to construct an entity relationship graph; Extract all connected subgraphs in the entity relationship graph; Construct a clustered entity set according to the nodes included in the connected subgraph; 8. A retrieval enhanced generation device based on a semantic enhanced knowledge graph, characterized in that The retrieval-enhanced generation device based on the semantic-enhanced knowledge graph includes: A document acquisition module for acquiring an original text document; A semantic-enhanced knowledge graph construction module for constructing a semantic-enhanced knowledge graph according to the entity information and entity relationship information extracted from the original text document, where the semantic-enhanced knowledge graph includes an entity set, an entity relationship set, a relevant sentence set, and a mapping relationship set, the relevant sentence set refers to a sentence set related to entities and entity relationships, and the mapping relationship set refers to a set including the mapping relationship between sentences in the original text document and target objects, and the target objects are entity nodes and relationship edges in the knowledge graph; A semantic-enhanced knowledge graph processing module for sequentially performing entity alignment processing on the semantic-enhanced knowledge graph to obtain an aligned semantic-enhanced knowledge graph; The sub-graph construction module is used to extract the entities appearing in the user's question by using a large language model, and obtain an explicit entity set containing several explicit entities; according to the explicit entity set, perform sub-graph expansion on the aligned semantic enhanced knowledge graph to obtain a mixed entity set containing the explicit entities and implicit entities, where the implicit entity refers to an entity involved in the user's question but not appearing in the user's question; extract a sub-graph from the aligned semantic enhanced knowledge graph according to the mixed entity set. The judgment module is used to use a large language model to judge whether the entity information and the entity relationship information in the sub-graph can answer the user's question, and obtain a first judgment result; if the first judgment result is yes, use the entity information and the entity relationship information in the sub-graph as the first context information, and construct a first prompt information according to the first context information and the user's question; based on the first prompt information, use a large language model to generate an answer corresponding to the user's question; if the first judgment result is no, extract relevant entity information from the mixed relevant sentence set, use the relevant entity information and the entity information and the entity relationship information in the sub-graph as the second context information, and construct a second prompt information according to the second context information and the user's question; based on the second prompt information, use a large language model to generate an answer corresponding to the user's question, where the mixed relevant sentence set is a sentence set in the sub-graph related to the mixed entities and the relationships between the mixed entities.
Citation Information
Patent Citations
Knowledge graph question-answering method for sub-graph retrieval optimization
CN117149974A
Retrieval enhancement generation system and method based on knowledge graph
CN118643134A