Hypergraph structure knowledge representation-based question and answer method for retrieval enhancement generation
By constructing a knowledge representation method of hypergraph structure, the problem of not being able to effectively capture multi-entity relationships in the prior art is solved, and more efficient and accurate knowledge retrieval and answer generation are achieved.
Patent Information
- Application Number
- CN202510318439.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-18
AI Technical Summary
The existing retrieval enhancement generation methods based on text fragments and graph structures cannot effectively capture the complex relationships between entities, resulting in insufficient knowledge representation ability and inefficient retrieval efficiency, and the inability to fully model the widespread multi-entity relationship in the real world, affecting the accuracy of the answers.
By constructing knowledge representation based on hypergraph structure, we use the multivariate relationship of natural language documents to build a knowledge hypergraph, search and fusion of entities and hyperedges, and generate target answers.
It improves the accuracy and reasoning ability of knowledge modeling, improves the retrieval efficiency and the accuracy of answers, and can more comprehensively represent the complex knowledge structure in the real world.
Smart Images

Figure CN120336455A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of question answering, and particularly to a question answering method and device for retrieval-enhanced generation based on hypergraph structure knowledge representation. Background Art
[0002] Retrieval-Augmented Generation (RAG) technology uses large-scale and structured knowledge bases as external information sources and combines language generation models to answer natural language questions, and can be widely applied to knowledge-intensive tasks, such as question answering systems, text generation, and medical consultations.
[0003] Among them, the existing RAG technologies can be divided into two categories: one is the retrieval-enhanced method based on text fragments, and the other is the retrieval-enhanced method based on graph structures. Among them, the RAG method based on text fragments directly takes natural language questions as input, and assists large language models in generating answers by retrieving text fragments in the knowledge base. The whole process depends on dense vector matching and fragment retrieval. Also, the RAG method based on graph structures, such as GraphRAG, represents knowledge as a graph structure, and uses entities and relationships in the graph for retrieval and reasoning to generate more accurate answers. However, the above text fragment-based method cannot effectively capture the complex relationships between entities, resulting in insufficient knowledge representation ability; while the graph structure-based method is limited by the representation of binary relationships and cannot fully model the multi-entity relationships widely existing in the real world, resulting in the incompleteness of knowledge representation and the insufficiency of reasoning ability, thereby reducing the retrieval efficiency and the accuracy of answers. Summary of the Invention
[0004] The present invention provides a question answering method and device for retrieval-enhanced generation based on hypergraph structure knowledge representation, so as to propose a question answering method for retrieval-enhanced generation based on hypergraph structure knowledge representation.
[0005] To this end, the present invention proposes a question answering method for retrieval-enhanced generation based on hypergraph structure knowledge representation, which can construct a knowledge hypergraph through the multi-entity relationships of natural language documents, and generate a target answer based on the retrieved first hypergraph set and second hypergraph set through hypergraph knowledge fusion and generation enhancement, thereby effectively solving the limitations of binary relationship representation, improving the accuracy of knowledge modeling and reasoning ability, and further improving the retrieval efficiency and the accuracy of answers.
[0006] Another object of the present invention is to propose a question answering device for retrieval-enhanced generation based on hypergraph structure knowledge representation.
[0007] To achieve the above object, on the one hand, the present invention proposes a question answering method for retrieval-enhanced generation based on hypergraph structure knowledge representation, and the method includes:
[0008] Construct a knowledge hypergraph based on the multiple relationships in natural language documents;
[0009] Obtain the question data input by the user, and extract the first entity in the question data;
[0010] Based on the first entity, perform entity retrieval from the knowledge hypergraph to obtain a second entity;
[0011] Based on the question data, perform hyperedge retrieval from the knowledge hypergraph to obtain a first hyperedge;
[0012] Perform hypergraph knowledge fusion on the second entity and the first hyperedge to obtain a hypergraph knowledge result;
[0013] Generate a target answer based on the hypergraph knowledge result and the question data input by the user.
[0014] The question-answering method based on retrieval-enhanced generation using hypergraph-structured knowledge representation according to the embodiments of the present invention may further have the following additional technical features:
[0015] In an embodiment of the present invention, the constructing a knowledge hypergraph based on the multiple relationships in natural language documents includes:
[0016] Extract multiple multiple relationships in the natural language document through an LLM-driven extraction method, where each multiple relationship consists of a second hyperedge and multiple third entities;
[0017] Construct a knowledge hypergraph based on the multiple multiple relationships;
[0018] Store the knowledge hypergraph in a database through vector representation and a bipartite graph structure.
[0019] In an embodiment of the present invention, the third entities in the knowledge hypergraph include entity names, types, explanations, and first confidence scores; the performing entity retrieval from the knowledge hypergraph based on the first entity to obtain a second entity includes:
[0020] Perform vector representation on the first entity to obtain a corresponding first vector;
[0021] Obtain the second vector and the first confidence score corresponding to the third entity in the knowledge hypergraph;
[0022] Calculate a first similarity value between the first vector and the second vector;
[0023] Based on the first similarity value and the first confidence score, determine a first retrieval score for the third entity;
[0024] Determine the entity in the third entity whose first retrieval score meets the first retrieval condition as the second entity.
[0025] In one embodiment of the present invention, the second hyperedge in the knowledge hypergraph includes a natural language description and a second confidence score; the retrieving the first hyperedge from the knowledge hypergraph based on the question data includes:
[0026] Perform vector representation on the question data to obtain a corresponding third vector;
[0027] Obtain the fourth vector and the second confidence score corresponding to the second hyperedge in the knowledge hypergraph;
[0028] Calculate a second similarity value between the third vector and the fourth vector;
[0029] Determine the second retrieval score of the second hyperedge based on the second similarity value and the second confidence score;
[0030] Determine the hyperedge whose second retrieval score in the second hyperedge meets the second retrieval condition as the first hyperedge.
[0031] In one embodiment of the present invention, the fusing the second entity and the first hyperedge for hypergraph knowledge to obtain a hypergraph knowledge result includes:
[0032] Perform hyperedge expansion on the second entity through the knowledge hypergraph to obtain a corresponding first hypergraph set;
[0033] Perform entity expansion on the first hyperedge through the knowledge hypergraph to obtain a corresponding second hypergraph set;
[0034] Perform hypergraph knowledge fusion on the first hypergraph set and the second hypergraph set to obtain a hypergraph knowledge result.
[0035] In one embodiment of the present invention, the generating a target answer based on the hypergraph knowledge result and the question data input by the user includes:
[0036] Determine the retrieval result of the question data input by the user;
[0037] Combine the hypergraph knowledge result and the retrieval result through a hybrid RAG fusion mechanism to generate a target knowledge input;
[0038] Input the target knowledge input and the question data input by the user into a large language model to generate a target answer.
[0039] To achieve the above object, on the other hand, the present invention proposes a question-answering device for retrieval-enhanced generation based on hypergraph-structured knowledge representation, and the device includes:
[0040] A construction module for constructing a knowledge hypergraph based on the multiple relationships of natural language documents;
[0041] An extraction module for obtaining the problem data input by the user and extracting the first entity in the problem data;
[0042] A first retrieval module for performing entity retrieval from the knowledge hypergraph based on the first entity to obtain a second entity;
[0043] A second retrieval module for performing hyperedge retrieval from the knowledge hypergraph based on the problem data to obtain a first hyperedge;
[0044] A fusion module for performing hypergraph knowledge fusion on the second entity and the first hyperedge to obtain a hypergraph knowledge result;
[0045] A generation module for generating a target answer based on the hypergraph knowledge result and the problem data input by the user.
[0046] Another object of the present invention is to propose an electronic device, comprising:
[0047] At least one processor; and
[0048] A memory communicatively connected to the at least one processor; wherein,
[0049] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of the foregoing aspects.
[0050] Another object of the present invention is to propose a computer storage medium, wherein the computer storage medium stores computer-executable instructions; the computer-executable instructions, when executed by a processor, cause the computer to execute the method according to any one of the foregoing aspects.
[0051] The question-answering method and device for retrieval-enhanced generation based on hypergraph structure knowledge representation according to the embodiments of the present invention construct a knowledge hypergraph based on the multiple relationships of natural language documents; obtain the question data input by the user, and extract the first entity in the question data; perform entity retrieval from the knowledge hypergraph based on the first entity to obtain a second entity; perform hyperedge retrieval from the knowledge hypergraph based on the question data to obtain a first hyperedge; perform hypergraph knowledge fusion on the second entity and the first hyperedge to obtain a hypergraph knowledge result; and generate a target answer based on the hypergraph knowledge result and the question data input by the user. Thus, the present invention can construct a knowledge hypergraph through the multiple relationships of natural language documents, and through hypergraph knowledge fusion and generation enhancement, generate a target answer based on the retrieved first hypergraph set and second hypergraph set, thereby effectively solving the limitations of binary relationship representation, improving the accuracy of knowledge modeling and reasoning ability, and further improving the retrieval efficiency and the accuracy of the answer.
[0052] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0054] Figure 1 is a flowchart of the question-answering method for retrieval-enhanced generation based on hypergraph structure knowledge representation according to the embodiments of the present invention;
[0055] Figure 2 is a structural diagram of the question-answering device for retrieval-enhanced generation based on hypergraph structure knowledge representation according to the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0057] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0058] In the prior art, the RAG method based on text fragments (such as standard RAG) divides the text in the knowledge base into fragments of a fixed length and retrieves the fragments related to the question through dense vector matching. Subsequently, the relevant fragments are used as context inputs for a large language model to assist in generating answers. However, the above method cannot effectively capture the complex relationships between entities, resulting in insufficient knowledge representation ability and low retrieval efficiency. For example, when dealing with multi-entity relationships, standard RAG may not accurately understand the semantic associations between entities, thus affecting the quality of the generated answers. Also, the RAG method based on graph structure (such as GraphRAG) improves the retrieval accuracy by representing knowledge as a graph structure, where the nodes in the graph represent entities and the edges represent the relationships between entities. However, the above method GraphRAG is limited by the representation of binary relationships and cannot fully model the widespread multi-entity (n-ary) relationships in the real world. For example, in the medical field, when describing the fact that "male hypertensive patients with serum creatinine levels between 115 - 133 μmol / L are diagnosed with mild serum creatinine elevation", GraphRAG needs to decompose it into multiple isolated binary relationships, resulting in incomplete knowledge representation and insufficient reasoning ability.
[0059] To solve the above technical problems in the related art, the present invention proposes a question-answering method for retrieval-enhanced generation based on hypergraph-structured knowledge representation to improve the performance of the question-answering system through hypergraph-structured knowledge representation.
[0060] The question-answering method and device for retrieval-enhanced generation based on hypergraph-structured knowledge representation proposed according to an embodiment of the present invention will be described below with reference to the accompanying drawings.
[0061] Figure 1 is a flowchart of the question-answering method for retrieval-enhanced generation based on hypergraph-structured knowledge representation according to an embodiment of the present invention.
[0062] As Figure 1 shown, the method includes:
[0063] S1, constructing a knowledge hypergraph based on the multi-entity relationships in natural language documents;
[0064] In an embodiment of the present invention, a knowledge hypergraph is defined as a structured knowledge representation form that can represent complex multi-entity relationships. Among them, the knowledge hypergraph can be composed of an entity set V, a hyperedge set E H and a relationship set F between hyperedges and entities. Specifically, the entity set V can contain all entities participating in knowledge representation, which can be specific concepts, objects, or attributes; the hyperedge set E H is the core component of the knowledge hypergraph, and each hyperedge e H ∈EH It can connect two or more entities, thus being able to represent complex multi-entity relationships; the relationship set F between the hyperedge and the entities consists of the hyperedge and the set of entities it connects, denoted as Fn = (e H , V eH ), where e H is the natural language description of the hyperedge, and V eH is the set of entities associated with this hyperedge. For example, a hyperedge can be represented as e H = (hypertensive patients, male, serum creatinine level between 115–133 μmol / L, mild elevation of serum creatinine). This representation can fully capture the complex relationships between multiple entities, not just simple binary relationships. Thus, through the above knowledge hypergraph structure, the complex knowledge structure in the real world can be represented more comprehensively, providing stronger support for knowledge-intensive tasks.
[0065] Among them, in an embodiment of the present invention, the method for constructing a knowledge hypergraph based on the above-mentioned multi-relationship in natural language documents may include the following steps:
[0066] S11, Extract multiple multi-relationships in the natural language document through an LLM-driven extraction method, where each multi-relationship consists of a second hyperedge and multiple third entities;
[0067] S12, Construct a knowledge hypergraph based on the multiple multi-relationships;
[0068] S13, Store the knowledge hypergraph in the database through vector representation and bipartite graph structure.
[0069] Specifically, in an embodiment of the present invention, the above-mentioned LLM (Large Language Model)-driven extraction method segments the knowledge fragments in the natural language document into hyperedges and identifies the entities therein, thereby extracting multiple multi-relationships in the natural language document. Among them, the above extraction process can be formalized as:
[0070]
[0071] Among them, P ext is the designed extraction prompt, π is the LLM, d is the input text, f n = (e H , V eH ), e H is the second hyperedge, representing the relationship between multiple third entities, and V eH is the set of entities associated with the hyperedge e H , and F n are both sets including multiple multi-relationships.
[0072] Also, in one embodiment of the present invention, the above third entity v = (vname, vtype, vexplain, vscore) includes an entity name vname, a type vtype, an explanation vexplain, and a first confidence score vscore, where the first confidence score represents the credibility of the third entity extracted from the hyperedge text; the second hyperedge e H = (etext, escore) includes a natural language description etext and a second confidence score escore, where the second confidence score represents the association strength between the second hyperedge and the original text.
[0073] Furthermore, in one embodiment of the present invention, after extracting multiple multi-relations in a natural language document through the above steps, a knowledge hypergraph can be constructed based on the multiple multi-relations. Also, in one embodiment of the present invention, after constructing the knowledge hypergraph, the knowledge hypergraph can be stored in a database through vector representation and a bipartite graph structure.
[0074] Specifically, in one embodiment of the present invention, the above knowledge hypergraph can be stored in a database using a bipartite graph structure for subsequent efficient storage and querying of the knowledge hypergraph. Among them, the bipartite graph structure G B = (V B , E B ) is constructed from the knowledge hypergraph G H = (V, E H ) through a transformation function:
[0075] V B = (V, E H )
[0076] E V = {(e H , v)|e H ∈ e H , v ∈ V H}
[0077] Among them, V B is the node set of the bipartite graph, including the third entity and the second hyperedge; E B is the edge set, representing the connection relationship between the second hyperedge and the third entity. The bipartite graph storage structure not only retains the integrity of the hypergraph but also utilizes the efficient querying ability of a general graph database. Also, the bipartite graph structure allows incremental addition of new hypergraph information in the bipartite graph database, thus supporting the dynamic update of the knowledge base, facilitating the expansion of the hypergraph knowledge base, while retaining the integrity of the hypergraph structure and query efficiency.
[0078] Moreover, in an embodiment of the present invention, the second hyperedge and the third entity in the knowledge hypergraph can be embedded into the same vector space to support efficient semantic retrieval. In an embodiment of the present invention, an embedding model z can be used to convert the second hyperedge and the third entity into vector representations:
[0079]
[0080] where is the vector set of the second hyperedge, E V ={E V |v∈V}, and the vector representations of each second hyperedge e H and the third entity v can be and h v =z(v) respectively. Through the above vector representations, the knowledge in the knowledge hypergraph can be efficiently semantically matched with the user's question, thus supporting the subsequent retrieval and generation processes.
[0081] S2. Obtain the question data input by the user, and extract the first entity from the question data;
[0082] In an embodiment of the present invention, the question data input by the user can be obtained through the front-end page.
[0083] Moreover, in an embodiment of the present invention, after obtaining the question data input by the user, it can be completed by designing a specific prompt and using a large language model (LLM). Specifically, the entity extraction process can be formalized as:
[0084] V q ~π(V|p q_ext ,q)
[0085] where V q is the entity set extracted from the question data q input by the user, p q_ext is the prompt designed for entity extraction, and π is the large language model for extraction.
[0086] S3. Based on the first entity, perform entity retrieval from the knowledge hypergraph to obtain the second entity;
[0087] where, in an embodiment of the present invention, after obtaining the first entity through the above steps, the second entity can be obtained by performing entity retrieval from the knowledge hypergraph based on the first entity.
[0088] In an embodiment of the present invention, the method for performing entity retrieval from the knowledge hypergraph based on the first entity to obtain the second entity may include the following steps:
[0089] S31. Perform vector representation on the first entity to obtain the corresponding first vector;
[0090] S32. Obtain the second vector and the first confidence score corresponding to the third entity in the knowledge hypergraph;
[0091] S33. Calculate the first similarity value between the first vector and the second vector;
[0092] S34. Determine the first retrieval score of the third entity based on the first similarity value and the first confidence score;
[0093] S35. Determine the entities in the third entity whose first retrieval scores meet the first retrieval condition as the second entity.
[0094] Among them, in an embodiment of the present invention, the above-mentioned first entity V q is vectorized to obtain the corresponding first vector as and obtain the second vector h corresponding to the third entity in the knowledge hypergraph V and the first confidence score v score .
[0095] In addition, in an embodiment of the present invention, the first similarity value between the first vector and the second vector h v can be calculated through the similarity function sim The similarity function sim can be a cosine similarity function.
[0096] Furthermore, in an embodiment of the present invention, after obtaining the first similarity value and the first confidence score v score through the above steps, the first retrieval score of the third entity can be determined through .
[0097] In addition, in an embodiment of the present invention, the above-mentioned first retrieval condition may be that the first retrieval score is greater than the first retrieval threshold τ V for the fourth entity, and the entities of the first threshold k v with the highest first retrieval scores among the fourth entities sorted in descending order according to the first retrieval score. Based on this, in an embodiment of the present invention, the above-mentioned second entity where ⊙ represents element-wise multiplication.
[0098] In an embodiment of the present invention, entity retrieval through the above steps not only considers semantic similarity but also combines the confidence of entities, thereby improving the accuracy and reliability of retrieval.
[0099] S4. Based on the question data, perform hyperedge retrieval from the knowledge hypergraph to obtain the first hyperedge;
[0100] Among them, in an embodiment of the present invention, after obtaining the problem data through the above steps, based on the problem data, a first hyperedge can be retrieved from the knowledge hypergraph.
[0101] In an embodiment of the present invention, the method for retrieving the first hyperedge from the knowledge hypergraph based on the problem data may include the following steps:
[0102] S41, perform vector representation on the problem data to obtain a corresponding third vector;
[0103] S42, obtain the fourth vector and the second confidence score corresponding to the second hyperedge in the knowledge hypergraph;
[0104] S43, calculate the second similarity value between the third vector and the fourth vector;
[0105] S44, determine the second retrieval score of the second hyperedge based on the second similarity value and the second confidence score;
[0106] S45, determine the hyperedges in the second hyperedge whose second retrieval scores meet the second retrieval condition as the first hyperedge.
[0107] Among them, in an embodiment of the present invention, the above-mentioned vector representation of the problem data q to obtain the corresponding third vector is h q = z(q), and obtain the fourth vector h corresponding to the second hyperedge in the knowledge hypergraph eH and the second confidence score
[0108] Moreover, in an embodiment of the present invention, the second similarity value between the third vector h q and the fourth vector can be calculated through the similarity function sim The similarity function sim can be a cosine similarity function.
[0109] Furthermore, in an embodiment of the present invention, after obtaining the second similarity value sim(g q , h eH ) and the second confidence score , the second retrieval score of the second hyperedge can be determined through Moreover, in an embodiment of the present invention, the above-mentioned second retrieval condition may be that the second retrieval score is greater than the second retrieval threshold τ
[0110] of the third hyperedge, and the top second threshold k H hyperedges obtained by sorting the third hyperedges in descending order according to the second retrieval score. Based on this, in an embodiment of the present invention, the above-mentioned first hyperedge H
[0111] In one embodiment of the present invention, through the above steps for hyperedge retrieval, by expanding the retrieval scope, not only hyperedges directly related to the user's question are retrieved, but also the knowledge coverage is further expanded through the entities connected by the hyperedges, thereby improving the accuracy of the answer.
[0112] S5. Perform hypergraph knowledge fusion on the second entity and the first hyperedge to obtain a hypergraph knowledge result;
[0113] In one embodiment of the present invention, after obtaining the second entity and the first hyperedge through the above steps, the second entity and the first hyperedge can be subjected to hypergraph knowledge fusion to obtain a hypergraph knowledge result.
[0114] Among them, in one embodiment of the present invention, the method of performing hypergraph knowledge fusion on the second entity and the first hyperedge to obtain a hypergraph knowledge result may include the following steps:
[0115] S51. Expand the hyperedges of the second entity through the knowledge hypergraph to obtain a corresponding first hypergraph set;
[0116] S52. Expand the entities of the first hyperedge through the knowledge hypergraph to obtain a corresponding second hypergraph set;
[0117] S53. Perform hypergraph knowledge fusion on the first hypergraph set and the second hypergraph set to obtain a hypergraph knowledge result.
[0118] Among them, in one embodiment of the present invention, the second entity can be retrieved through the knowledge hypergraph, the hyperedges connecting the second entity in the knowledge hypergraph are retrieved, and all the retrieved hyperedges are determined as the corresponding first hypergraph set
[0119] And, in one embodiment of the present invention, the first hyperedge can be retrieved through the knowledge hypergraph, the entities connected to the first hyperedge in the knowledge hypergraph are retrieved, and all the retrieved entities are determined as the corresponding second hypergraph set
[0120] Furthermore, in one embodiment of the present invention, the first hypergraph set expanded from the second entity and the second hypergraph set expanded from the first hyperedge are subjected to hypergraph knowledge fusion to form a complete retrieved n-ary relation fact set, obtaining a hypergraph knowledge result, thereby ensuring that the knowledge input into the large prediction model subsequently is as complete as possible.
[0121] S6. Generate a target answer based on the hypergraph knowledge result and the question data input by the user.
[0122] In one embodiment of the present invention, after the first hypergraph set and the second hypergraph set complete hypergraph knowledge fusion through the above steps, a target answer can be generated based on the obtained hypergraph knowledge result and the problem data input by the user.
[0123] Specifically, in one embodiment of the present invention, the method for generating a target answer based on the hypergraph knowledge result and the problem data input by the user may include the following steps:
[0124] S61, determining the retrieval result of the problem data input by the user;
[0125] S62, combining the hypergraph knowledge result and the retrieval result through a hybrid RAG fusion mechanism to generate a target knowledge input;
[0126] S63, inputting the target knowledge input and the problem data input by the user into a large language model to generate a target answer.
[0127] Among them, in one embodiment of the present invention, the retrieval result of the problem data input by the user can be obtained through a traditional RAG method. For example, by calculating the vector similarity between the problem data q and the segment d in the knowledge base, and determining the segments with a similarity higher than the threshold τC as the retrieval result of the problem data input by the user.
[0128] And, in one embodiment of the present invention, the generation enhancement stage can generate a target knowledge input by combining the hypergraph knowledge result and the retrieval result. Specifically, in one embodiment of the present invention, a hybrid RAG fusion mechanism can be adopted to combine the hypergraph knowledge result and the segment-based retrieval result to generate a target knowledge input, thereby not only making full use of the complex relationship modeling ability of the hypergraph, but also ensuring the diversity and reliability of the generation process through hybrid knowledge input, providing strong support for knowledge-intensive applications.
[0129] Furthermore, in one embodiment of the present invention, the hypergraph-guided generation mechanism combines the structured knowledge in the knowledge hypergraph with the traditional segment-based retrieval result through two stages of hypergraph knowledge fusion and generation enhancement, not only making full use of the complex relationship modeling ability of the hypergraph, but also ensuring the diversity and reliability of the generation process through hybrid knowledge input, providing strong support for knowledge-intensive applications, thereby improving the factuality and accuracy of the generated answer.
[0130] Based on the above description, an example of a question-answering method for retrieval-enhanced generation based on hypergraph structure knowledge representation is given. For example, assume that after obtaining the question "How to diagnose male hypertensive patients when their serum creatinine level is between 115 - 133 μmol / L?" entered by the user, the first entities of this question, such as "male hypertensive patients", "serum creatinine level", and "diagnosis", are extracted. Then, entity retrieval and hyperedge retrieval are performed in the knowledge hypergraph to obtain the second entity and the first hyperedge. For example, "the relationship between male hypertensive patients and serum creatinine level between 115 - 133 μmol / L and the corresponding diagnosis results". Then, the retrieved first hyperedge and the second entity are subjected to hypergraph knowledge fusion to generate a hypergraph knowledge result, and the hypergraph knowledge result, the retrieval result, and the question entered by the user are input into a large language model to generate the final answer.
[0131] In the question-answering method for retrieval-enhanced generation based on hypergraph structure knowledge representation proposed by the present invention, a knowledge hypergraph is constructed based on the multiple relationships of natural language documents; question data input by the user is obtained, and the first entity in the question data is extracted; based on the first entity, entity retrieval is performed in the knowledge hypergraph to obtain the second entity; based on the question data, hyperedge retrieval is performed in the knowledge hypergraph to obtain the first hyperedge; the second entity and the first hyperedge are subjected to hypergraph knowledge fusion to obtain a hypergraph knowledge result; based on the hypergraph knowledge result and the question data input by the user, a target answer is generated. Thus, the present invention can construct a knowledge hypergraph through the multiple relationships of natural language documents, and through hypergraph knowledge fusion and generation enhancement, generate a target answer based on the retrieved first hypergraph set and the second hypergraph set, thereby effectively solving the limitations of binary relationship representation, improving the accuracy of knowledge modeling and reasoning ability, and further improving the retrieval efficiency and the accuracy of the answer.
[0132] To implement the above embodiments, as Figure 2 shown, in this embodiment, a question-answering device 10 for retrieval-enhanced generation based on hypergraph structure knowledge representation is further provided. The device includes a construction module 301, an extraction module 302, a first retrieval module 303, a second retrieval module 304, a fusion module 305, and a generation module 306;
[0133] The construction module 301 is used to construct a knowledge hypergraph based on the multiple relationships of natural language documents;
[0134] The extraction module 302 is used to obtain the question data input by the user and extract the first entity in the question data;
[0135] The first retrieval module 303 is used to perform entity retrieval in the knowledge hypergraph based on the first entity to obtain the second entity;
[0136] The second retrieval module 304 is used to perform hyperedge retrieval in the knowledge hypergraph based on the question data to obtain the first hyperedge;
[0137] A fusion module 305, configured to perform hypergraph knowledge fusion on a second entity and a first hyperedge to obtain a hypergraph knowledge result;
[0138] A generation module 306, configured to generate a target answer based on the hypergraph knowledge result and the question data input by the user.
[0139] In an embodiment of the present invention, the above-mentioned construction module 301 is specifically configured to:
[0140] Extract a plurality of multi - element relationships from a natural language document by an LLM - driven extraction method, where each multi - element relationship consists of a second hyperedge and a plurality of third entities;
[0141] Construct a knowledge hypergraph based on the plurality of multi - element relationships;
[0142] Store the knowledge hypergraph in a database through vector representation and a bipartite graph structure.
[0143] In an embodiment of the present invention, the third entities in the above - mentioned knowledge hypergraph include entity names, types, explanations, and first confidence scores; the above - mentioned first retrieval module 303 is specifically configured to:
[0144] Perform vector representation on the first entity to obtain a corresponding first vector;
[0145] Obtain the second vector and the first confidence score corresponding to the third entity in the knowledge hypergraph;
[0146] Calculate a first similarity value between the first vector and the second vector;
[0147] Determine a first retrieval score of the third entity based on the first similarity value and the first confidence score;
[0148] Determine the entities in the third entity whose first retrieval scores meet the first retrieval condition as the second entity.
[0149] Further, in an embodiment of the present invention, the second hyperedges in the above - mentioned knowledge hypergraph include natural language descriptions and second confidence scores; the above - mentioned second retrieval module 304 is specifically configured to:
[0150] Perform vector representation on the question data to obtain a corresponding third vector;
[0151] Obtain the fourth vector and the second confidence score corresponding to the second hyperedge in the knowledge hypergraph;
[0152] Calculate a second similarity value between the third vector and the fourth vector;
[0153] Determine a second retrieval score of the second hyperedge based on the second similarity value and the second confidence score;
[0154] Determine the hyperedges in the second hyperedge whose second retrieval scores meet the second retrieval condition as the first hyperedges.
[0155] Furthermore, in an embodiment of the present invention, the above-mentioned fusion module 305 is specifically configured to:
[0156] Perform hyperedge expansion on the second entity through the knowledge hypergraph to obtain the corresponding first hypergraph set;
[0157] Perform entity expansion on the first hyperedge through the knowledge hypergraph to obtain the corresponding second hypergraph set;
[0158] Perform hypergraph knowledge fusion on the first hypergraph set and the second hypergraph set to obtain the hypergraph knowledge result.
[0159] Furthermore, in an embodiment of the present invention, the above-mentioned generation module 306 is specifically configured to:
[0160] Determine the retrieval result of the problem data input by the user;
[0161] Combine the hypergraph knowledge result and the retrieval result through a hybrid RAG fusion mechanism to generate the target knowledge input;
[0162] Input the target knowledge input and the problem data input by the user into the large language model to generate the target answer.
[0163] In the question and answer device for retrieval-enhanced generation based on hypergraph structure knowledge representation proposed by the present invention, a knowledge hypergraph is constructed based on the multiple relationships of natural language documents; the problem data input by the user is obtained, and the first entity in the problem data is extracted; based on the first entity, entity retrieval is performed from the knowledge hypergraph to obtain the second entity; based on the problem data, hyperedge retrieval is performed from the knowledge hypergraph to obtain the first hyperedge; the second entity and the first hyperedge are subjected to hypergraph knowledge fusion to obtain the hypergraph knowledge result; based on the hypergraph knowledge result and the problem data input by the user, the target answer is generated. Thus, the present invention can construct a knowledge hypergraph through the multiple relationships of natural language documents, and through hypergraph knowledge fusion and generation enhancement, generate the target answer based on the retrieved first hypergraph set and second hypergraph set, thereby effectively solving the limitations of binary relationship representation, improving the accuracy of knowledge modeling and reasoning ability, and further improving the retrieval efficiency and the accuracy of the answer.
[0164] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0165] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
Claims
1. A question answering method for retrieval-augmented generation based on hypergraph structure knowledge representation, characterized in that, The method includes: Constructing a knowledge hypergraph based on the multiple relationships in natural language documents; Obtaining the problem data input by the user and extracting the first entity in the problem data; Performing entity retrieval from the knowledge hypergraph based on the first entity to obtain a second entity; Performing hyperedge retrieval from the knowledge hypergraph based on the problem data to obtain a first hyperedge; Performing hypergraph knowledge fusion on the second entity and the first hyperedge to obtain a hypergraph knowledge result; Generating a target answer based on the hypergraph knowledge result and the problem data input by the user.
2. The method according to claim 1, wherein The constructing a knowledge hypergraph based on the multiple relationships in natural language documents includes: Extracting multiple multiple relationships in the natural language document through an LLM-driven extraction method, where each multiple relationship consists of a second hyperedge and multiple third entities; Constructing a knowledge hypergraph based on the multiple multiple relationships; Storing the knowledge hypergraph in a database through vector representation and a bipartite graph structure.
3. The method according to claim 2, characterized in that, The third entities in the knowledge hypergraph include entity names, types, explanations, and first confidence scores; the performing entity retrieval from the knowledge hypergraph based on the first entity to obtain a second entity includes: Performing vector representation on the first entity to obtain a corresponding first vector; Obtaining the second vector and the first confidence score corresponding to the third entity in the knowledge hypergraph; Calculating a first similarity value between the first vector and the second vector; Determining a first retrieval score of the third entity based on the first similarity value and the first confidence score; Determining the entity whose first retrieval score in the third entity meets the first retrieval condition as the second entity.
4. The method according to claim 2, wherein The second hyperedges in the knowledge hypergraph include natural language descriptions and second confidence scores; the performing hyperedge retrieval from the knowledge hypergraph based on the problem data to obtain a first hyperedge includes: Performing vector representation on the problem data to obtain a corresponding third vector; Obtaining the fourth vector and the second confidence score corresponding to the second hyperedge in the knowledge hypergraph; Calculating a second similarity value between the third vector and the fourth vector; Determining a second retrieval score of the second hyperedge based on the second similarity value and the second confidence score; Determining the hyperedge whose second retrieval score in the second hyperedge meets the second retrieval condition as the first hyperedge.
5. The method according to claim 1, characterized in that, The performing hypergraph knowledge fusion on the second entity and the first hyperedge to obtain a hypergraph knowledge result includes: Performing hyperedge expansion on the second entity through the knowledge hypergraph to obtain a corresponding first hypergraph set; Performing entity expansion on the first hyperedge through the knowledge hypergraph to obtain a corresponding second hypergraph set; Performing hypergraph knowledge fusion on the first hypergraph set and the second hypergraph set to obtain a hypergraph knowledge result.
6. The method according to claim 1, wherein The generating a target answer based on the hypergraph knowledge result and the problem data input by the user includes: Determining the retrieval result of the problem data input by the user; Combining the hypergraph knowledge result and the retrieval result through a hybrid RAG fusion mechanism to generate a target knowledge input; Inputting the target knowledge input and the problem data input by the user into a large language model to generate a target answer.
7. A question answering device for retrieval-enhanced generation based on hypergraph structure knowledge representation, characterized in that, The device includes: A construction module, configured to construct a knowledge hypergraph based on the multiple relationships of a natural language document; An extraction module, configured to obtain the problem data input by a user and extract a first entity from the problem data; A first retrieval module, configured to perform entity retrieval from the knowledge hypergraph based on the first entity to obtain a second entity; A second retrieval module, configured to perform hyperedge retrieval from the knowledge hypergraph based on the problem data to obtain a first hyperedge; A fusion module, configured to perform hypergraph knowledge fusion on the second entity and the first hyperedge to obtain a hypergraph knowledge result; A generation module, configured to generate a target answer based on the hypergraph knowledge result and the problem data input by the user.
8. The device according to claim 6, characterized in that, The construction module is specifically configured to: Extract multiple multiple relationships in the natural language document by an extraction method driven by an LLM, where each multiple relationship consists of a second hyperedge and multiple third entities; Construct a knowledge hypergraph based on the multiple multiple relationships; Store the knowledge hypergraph in a database through vector representation and a bipartite graph structure.
9. A computer storage medium, wherein, The computer storage medium stores computer-executable instructions; after the computer-executable instructions are executed by a processor, the method according to any one of claims 1-6 can be implemented.
10. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method according to any one of claims 1-6 is implemented.
Citation Information
Cited By
Large model question and answer data generation method and device, medium and electronic equipment
CN120910112A
Cross-document question and answer method and system based on sparse hypergraph
CN121412280A
A cross-document question answering method and system based on sparse hypergraphs
CN121412280B
Elevator operation and maintenance retrieval enhanced question and answer method, device and equipment and storage medium
CN121412345A
Elevator operation and maintenance retrieval enhanced question and answer method, device, equipment and storage medium
CN121412345B