Question answering method and device based on retrieval enhancement generation, medium and equipment

By using the GRAFT-RAG framework, combined with structured graph retrieval and unstructured sparse-dense hybrid retrieval, the multi-hop question-answering capability of large language models is optimized, solving the problem of insufficient information integration in existing methods and improving the accuracy and consistency of generated answers.

CN121388104APending Publication Date: 2026-01-23SHENYANG AEROSPACE UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511504599.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing retrieval enhancement generation methods have shortcomings in multi-hop reasoning and information integration retrieval recall, especially in the ineffective integration of structured knowledge utilization, sparse and dense retrieval, and retrieval context redundancy, which affects the effectiveness and efficiency of the generation model.

Method used

The GRAFT-RAG framework is adopted, which integrates structured graph retrieval and unstructured sparse-dense hybrid retrieval. Multi-hop neighborhood information is aggregated through graph convolutional networks, and the maximum spanning tree algorithm and self-verifying inference module are introduced to optimize the information retrieval and generation process.

Benefits of technology

It significantly improves the reasoning accuracy and information relevance of large language models in multi-hop question answering tasks, reduces information redundancy and the risk of generating illusions, and improves the logical reliability and factual consistency of generated answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121388104A_ABST
    Figure CN121388104A_ABST
Patent Text Reader

Abstract

The invention discloses a question answering method and device based on retrieval enhancement generation, a medium and equipment, and relates to the technical field of computers. According to the method, related candidate sub-graphs are matched in an existing structured knowledge graph according to the query problem of a user, and the candidate sub-graphs are further judged to be insufficient to deal with the query problem through logical reasoning; according to the method, sparse keyword vectors of query questions based on surface vocabularies and dense question vectors based on context deep dependency are further extracted; matching the query question with a sparse semantic vector of each text block of the unstructured text based on surface vocabularies and a dense semantic vector of each text block based on context deep dependency, which are acquired in advance, so as to determine the text block related to the query question from the unstructured text; the candidate sub-graphs are further converted into graph structures to supplement the candidate sub-graphs, answers corresponding to the query questions are generated based on the graph structures, and the performance of questions and answers in multi-hop reasoning and information integration retrieval recall is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and particularly relates to a question and answer method, device, medium and equipment based on retrieval enhancement generation. BACKGROUND

[0002] At present, large language models (LLMs) have made remarkable progress in the field of natural language processing, showing excellent language understanding and content generation capabilities, and are widely used in question and answer, text generation, code assisted writing and other tasks. However, the essence of LLMs is still a static parameterized model, which has the key challenges of knowledge update lag and hallucination of generated content. To alleviate these problems, the method of retrieval-augmented generation (RAG) is proposed and gradually becomes the mainstream solution. RAG introduces an external knowledge retrieval mechanism before generating content, providing LLMs with the latest or specific domain information related to the query, effectively improving the accuracy and traceability of the generated results, and has become an important technical paradigm for enhancing the knowledge ability of LLMs.

[0003] The typical RAG system adopts a "retrieval-generation" two-stage architecture: first, the retriever retrieves relevant fragments from a large-scale document set, and then inputs the retrieved content to the LLM for answer generation. However, this method still has many limitations: on the one hand, the retrieval strategy based on keywords or semantic similarity can introduce redundant information and ignore the semantic and logical connections between retrieval fragments, limiting the reasoning potential of LLMs; on the other hand, the accuracy limitations of the retriever itself can lead to the omission of important information, which is particularly prominent when dealing with complex tasks such as multi-hop reasoning and multi-document integration.

[0004] To improve the reasoning performance and factual consistency of RAG, researchers have begun to introduce structured information such as knowledge graphs (KG) into RAG systems, using the relationships between entities to assist in the organization and connection of knowledge, and enhancing the logical coherence of generated content. In addition, some research has explored generation-driven RAG methods, such as the "generate-read" paradigm, trying to use the internal knowledge of LLMs to perform preliminary document construction and selection, thereby simplifying the traditional "retrieval-generation" process. Overall, as a bridge connecting language models and external knowledge, RAG is still in the stage of continuous evolution, and there is a broad research space in multi-hop reasoning, information integration and logical consistency guarantee.

[0005] From the perspective of data processing, the current RAG framework can be roughly divided into two categories: RAG based on unstructured text and RAG based on structured graph data. The former converts local text knowledge into high-dimensional vector representation, calculates the semantic similarity between the input question and the knowledge unit, and selects the top K most relevant fragments as the input of the LLM. The latter converts unstructured text into a structured knowledge graph in the form of triples (h, r, t), significantly improving the consistency of facts and the controllability of information in the question answering task, and reducing the illusion risk in the generation process.

[0006] However, the retrieval augmented generation method based on unstructured text ignores the potential logical structure relationship between the retrieved fragments, and the reasoning accuracy is insufficient, which limits its performance in multi-hop reasoning tasks. The retrieval augmented generation method based on structured graph data has insufficient retrieval recall due to the limited knowledge coverage of the knowledge graph. In summary, the existing retrieval augmented generation method still has deficiencies in multi-hop reasoning and information integration retrieval recall. SUMMARY

[0007] Therefore, it is necessary to provide a question answering method, device, medium and equipment based on retrieval augmented generation to solve the above technical problems.

[0008] The application adopts the following technical solutions: The application provides a question answering method based on retrieval augmented generation, comprising: obtaining a query question of a user, a pre-constructed structured knowledge graph, and an unstructured embedding vector; the unstructured embedding vector comprises a sparse semantic vector of each text block in the unstructured text based on surface vocabulary and a dense semantic vector of each text block based on context deep dependency; retrieving at least part of the graph structure information matching the query question from the pre-constructed structured knowledge graph, obtaining a candidate subgraph corresponding to the query question, and converting the candidate subgraph into a triple set; When the matching of the triple set and the query question is determined to be less than a preset threshold through logical reasoning, extracting a sparse keyword vector of the query question based on surface vocabulary and a dense problem vector based on context deep dependency; determining the mixed retrieval score of each text block according to the matching of the sparse semantic vector and the sparse keyword vector and the matching of the dense semantic vector and the dense problem vector; sorting each text block in descending order according to the mixed retrieval score; converting a preset number of text blocks into a graph structure to expand the candidate subgraph in sequence; generating an answer corresponding to the query question according to the expanded candidate subgraph and feeding back to the user.

[0009] Optionally, the matching of the triple set and the query question is determined through logical reasoning, specifically comprising: Based on a large language model, the matching degree between the set of triples and the query question is determined through logical reasoning.

[0010] Optionally, determining the hybrid retrieval score for each text block based on the matching between the sparse semantic vector and the sparse keyword vector, and the matching between the dense semantic vector and the dense question vector, specifically includes: The matching between sparse semantic vectors and sparse keyword vectors is determined by the following formula: ; The matching between dense semantic vectors and dense question vectors is determined by the following formula: ; The mixed retrieval score for each text block is determined by the following formula, based on the matching between sparse semantic vectors and sparse keyword vectors, and the matching between dense semantic vectors and dense question vectors: ; in, It is a sparse keyword vector. For text blocks c sparse semantic vectors, For text blocks c The matching value between the sparse semantic vector and the sparse keyword vector. For dense problem vectors, For text blocks c Dense semantic vectors For text blocks c The matching value between the dense semantic vector and the dense question vector. As weight, For text blocks c The combined search score.

[0011] Optionally, generating the answer to the query question based on the expanded candidate subgraph specifically includes: Based on the similarity between the triples corresponding to the entity pairs with paths in the expanded candidate subgraph and the query question, the weights of the edges between each entity pair in the expanded candidate subgraph are determined, resulting in a weighted candidate subgraph. Based on the weighted subgraph, the maximum spanning tree algorithm is used to obtain the spanning tree with the largest sum of edge weights. This tree serves as the set of entity paths most relevant to the query question, resulting in a compressed candidate subgraph. The compressed candidate subgraph is converted into a set of triples, and the answer to the query question is generated based on the set of triples.

[0012] Optionally, generating the answer to the query question based on the set of triples specifically includes: The query question is decomposed into multiple sub-questions, and for each sub-question, an answer corresponding to the sub-question is generated according to a triple set; When it is determined by the large language model that the answer corresponding to the sub-question is unreasonable, the answer corresponding to the sub-question is generated again according to the triple set; The answer corresponding to the query question is generated according to the answers corresponding to the sub-questions.

[0013] Optionally, the structured knowledge graph is constructed, and specifically includes: A document database is acquired, knowledge triples are extracted from the document database by a large language model, and the knowledge triples are preprocessed including data cleaning and deduplication processing; the knowledge triples include a head entity, a tail entity, a relationship between the head entity and the tail entity, and a corresponding text block; The preprocessed knowledge triples are converted into a knowledge graph according to the relationship between the head entity and the tail entity.

[0014] Optionally, the unstructured embedding vector is constructed, and specifically includes: Unstructured text is acquired, sparse semantic vectors of each text block in the unstructured text based on surface vocabulary are extracted by a BGE-Lexical model, and dense semantic vectors of each text block in the unstructured text based on deep context dependency are extracted by an mxbai-embed-large model.

[0015] The application provides a question and answer device based on retrieval enhancement generation, comprising: An acquisition module is configured to acquire a query question of a user, a pre-constructed structured knowledge graph, and an unstructured embedding vector; the unstructured embedding vector includes sparse semantic vectors of each text block in unstructured text based on surface vocabulary and dense semantic vectors of each text block based on deep context dependency; A structured retrieval module is configured to retrieve at least part of graph structure information matching the query question from the pre-constructed structured knowledge graph, obtain a candidate sub-graph corresponding to the query question, and convert the candidate sub-graph into a triple set; A question feature extraction module is configured to extract sparse keyword vectors of the query question based on surface vocabulary and dense question vectors based on deep context dependency when it is determined by logical reasoning that the matching of the triple set and the query question is less than a preset threshold; An unstructured retrieval module is configured to determine a hybrid retrieval score of each text block according to the matching of the sparse semantic vectors and the sparse keyword vectors and the matching of the dense semantic vectors and the dense question vectors, sort each text block in descending order according to the hybrid retrieval score, and convert a preset number of text blocks into a graph structure to expand the candidate sub-graph; The generating module is configured to generate an answer corresponding to the query question according to the extended candidate subgraph and feed back to the user.

[0016] The application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the above-mentioned question and answer method based on retrieval enhancement generation.

[0017] The application provides a computer device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor realizes the above-mentioned question and answer method based on retrieval enhancement generation when executing the program.

[0018] The above-mentioned at least one technical scheme adopted by the application can achieve the following beneficial effects: The application firstly matches a candidate subgraph in the existing structured knowledge graph according to the query question of the user, and further judges whether the candidate subgraph is sufficient to generate an answer to the query question through logical reasoning, when the candidate subgraph is insufficient to cope with the query question, the application further extracts a sparse key word vector based on surface vocabulary and a dense problem vector based on context deep dependence of the query question, and matches the sparse key word vector and the dense problem vector with a previously obtained unstructured embedding vector, the unstructured embedding vector comprises a sparse semantic vector based on surface vocabulary and a dense semantic vector based on context deep dependence of each text block in unstructured text, respectively, so as to accurately determine a text block related to the query question from the unstructured text, improve the reasoning accuracy, and further convert the text block into a graph structure to supplement the candidate subgraph, generate an answer corresponding to the query question based on this, and improve the performance of question and answer in multi-hop reasoning and information integration retrieval recall. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings, which are included to provide a further understanding of the application, constitute a part of this application, and the illustrative embodiments of the application and their description serve to explain the application, and do not limit the application in any way. In the drawings:

[0020] Figure 1 A question and answer method based on retrieval enhancement generation provided by the application is shown in the flowchart; Figure 2 A GRAFT-RAG framework system flowchart provided by the application is shown in the flowchart; Figure 3 A question and answer device based on retrieval enhancement generation provided by the application is shown in the flowchart; Figure 4 A computer device for realizing the question and answer method based on retrieval enhancement generation provided by the application is shown in the flowchart. DETAILED DESCRIPTION

[0021] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described below in connection with specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0022] Currently, the proposed tree structure and graph structure can effectively capture and express the logical and hierarchical relationship between knowledge. In recent years, there have been multiple studies trying to introduce tree structure and graph structure into RAG framework to improve the logical coherence and structured organization ability of retrieving information. In terms of tree structure, RAPTOR constructs a multi-level tree structure from bottom to top through recursive embedding, clustering and summarization; MemWalker regards a large language model as an intelligent agent that explores interactively in a summary tree. Similarly, SiReRAG proposes a double-tree structure of constructing similarity tree and relevance tree to capture similar and relevant knowledge fragments at the same time. However, these tree structure methods are usually limited to the logical organization within a single document, and it is difficult to fuse knowledge across documents.

[0023] In terms of graph structure, GNN-RAG first uses graph neural network (GNN) to reason on subgraphs of knowledge graph, effectively improving the retrieval effect of candidate answers. GraphRAG establishes a hierarchical graph structure index through knowledge graph construction and recursive summarization, effectively supporting query-focused text summarization and reasoning tasks. However, existing graph structure RAG methods usually require high computational overhead and complex graph structure construction process, and mostly return structured triples rather than directly available natural text, increasing the burden of the generation phase. In contrast, HopRAG provides a more lightweight solution to construct inter-document association graphs through pseudo queries to efficiently support multi-hop reasoning, but its accuracy in accurate knowledge representation and complex reasoning tasks still needs to be improved.

[0024] Although existing RAG methods have made significant progress, there are still obvious shortcomings in the following aspects: (1) Insufficient use of structured knowledge: Traditional RAG methods are mostly limited to unstructured text indexing, lack explicit modeling of knowledge structure and logical relationships, and are difficult to effectively support complex logical reasoning and multi-hop reasoning tasks. (2) Splitting of sparse and dense retrieval: Although sparse retrieval can efficiently capture keyword information, dense retrieval is better at capturing deep semantic information, but most existing methods do not effectively integrate the advantages of these two retrieval modes, which can lead to poor recall or inaccurate semantics. (3) Redundancy problem of retrieval context: Although existing structured RAG methods provide better knowledge organization, they generally have problems of subgraph redundancy and context overload in knowledge graphs, affecting the effectiveness and efficiency of the generation model.

[0025] To overcome the above problems, the present application proposes a novel generative question answering framework GRAFT-RAG, which integrates structured retrieval of knowledge graphs and sparse-dense hybrid retrieval of unstructured data to significantly improve the reasoning ability of LLMs in complex multi-hop problems. In the indexing phase, the local documents are first divided into semantic blocks by paragraphs and sentences, and subgraphs are constructed based on the smallest unit of semantic blocks to capture potential fact-level association relationships between blocks. In the retrieval phase, word-level matching and semantic matching signals are integrated to construct query-related subgraphs, and multi-hop neighborhood information is aggregated through a graph convolutional network (GCN). In the post-processing phase after retrieval, a context organization strategy for triple-path triples is further proposed to effectively filter redundant and ambiguous information, and only the information subgraph most closely related to the query is retained and input into the LLM along with the question to generate answers.

[0026] In summary, the main content of the present application is as follows: 1. GRAFT-RAG framework is proposed: integrating structured graph retrieval and unstructured sparse-dense hybrid retrieval, effectively enhancing the reasoning accuracy and information relevance of LLMs in multi-hop question answering tasks.

[0027] 2. Design of graph-enhanced dual-channel retrieval mechanism: construct subgraphs based on knowledge graphs and perform structured information verification, and cooperate with sparse-dense hybrid retrieval to improve the coverage and accuracy of information retrieval, especially suitable for complex information dispersion scenarios.

[0028] 3. Introducing context pruning and self-verification reasoning module: a subgraph compression strategy guided by maximum spanning tree (MST) and a self-verification reasoning mechanism (SVRM) are proposed, which effectively reduces information redundancy and generation illusion, and improves the factual consistency and logical reliability of the generated answers.

[0029] The technical solutions provided by the embodiments of the present application are described in detail below with reference to the drawings.

[0030] Figure 1 The present application is a question and answer method based on retrieval enhancement generation, which specifically includes the following steps: S101: Obtain the user's query question, the structured knowledge graph constructed in advance according to the document database, and the unstructured embedding vector; the unstructured embedding vector includes sparse semantic vectors of each text block in the corresponding document database based on surface vocabulary and dense semantic vectors of each text block based on context deep dependency.

[0031] S102: Retrieve at least part of the graph structure information matching the query question from the pre-constructed structured knowledge graph, obtain the candidate subgraph corresponding to the query question, and convert the candidate subgraph into a triple set.

[0032] S103: When the matching of the triple set and the query question is determined to be less than a preset threshold value through logical reasoning, extract the sparse keyword vector of the query question based on surface vocabulary and the dense problem vector based on context deep dependency.

[0033] S104: Determine the mixed retrieval score of each text block according to the matching of the sparse semantic vector and the sparse keyword vector and the matching of the dense semantic vector and the dense problem vector; sort each text block in descending order according to the mixed retrieval score; convert a preset number of text blocks into graph structure to expand the candidate subgraph in sequence.

[0034] S105: Generate an answer corresponding to the query question according to the expanded candidate subgraph and feed back to the user.

[0035] For convenience of description, the following will be described only with the server as the execution subject. The server mentioned in the present application can be a server arranged in a business platform, or a device such as a desktop computer, a notebook computer, etc. capable of executing the scheme of the present application.

[0036] Figure 2For a GRAFT-RAG framework system flowchart in the present application, GRAFT-RAG is committed to alleviating the multi-hop reasoning difficulty, information redundancy and illusion risk problems occurring in the traditional RAG framework. The following details the framework process, including the data preparation stage, the dual-channel hybrid retrieval stage, the context pruning and generation stage.

[0037] Figure 2 The overall workflow of the GRAFT-RAG framework is shown, and the system consists of three main stages: (1) Data preparation stage: simultaneously construct a structured knowledge graph (through triple extraction and graph database construction) and a sparse / dense hybrid index of unstructured text; (2) Dual-channel hybrid retrieval stage: preferentially perform structured graph retrieval, and judge whether it is sufficient to answer the question through the context adequacy verification module, if not, start sparse-dense hybrid retrieval to expand the context; (3) Context pruning and generation stage: first compress information redundancy through the maximum spanning tree guided graph pruning strategy, and then generate by the reasoning model for multiple rounds of reasoning, and the generated result is evaluated by the self-verification module to form the final reliable answer. The whole system integrates structured reasoning ability, unstructured coverage and verification mechanism, aiming to improve the accuracy and logical consistency of multi-hop question answering. The following details each stage.

[0038] 1. Data preparation stage.

[0039] In order to effectively utilize local knowledge resources to support subsequent hybrid retrieval and generation processes, in one or more embodiments of the present application, the data preprocessing stage covers parallel processing of knowledge graph construction (structured data) and vector embedding index (unstructured data).

[0040] Structured knowledge graph construction: Traditional knowledge graph-based question answering systems often rely on artificially constructed knowledge bases (e.g., Wikidata), but such methods have the disadvantages of insufficient knowledge coverage and high cost of manual maintenance.

[0041] Based on this, in order to construct a more flexible and efficient domain-specific knowledge graph, in one or more embodiments of the present application, the server can obtain a document database, extract knowledge triples from the document database through a large language model, and preprocess the knowledge triples including data cleaning and deduplication; the knowledge triples include head entity, tail entity, relationship between head entity and tail entity, and corresponding text block; so as to convert the preprocessed knowledge triples into a knowledge graph according to the relationship between the head entity and the tail entity.

[0042] For example, a pre-trained large language model (LLM) can be used for automatic triple extraction. Specifically, based on the local document database, the document database is denoted as wherein d i represents the i-th document. A set of dedicated triple extraction prompts can be designed, and a large language model (e.g., Qwen, LLaMA, etc.) can be used to extract high-quality knowledge triples, which can be formulated as: i G = {( h , r , t , c ) | c ∈ d i , h , t ∈ entity set, r ∈ relation set}. Here, (h, r, t) represents the head entity, the relationship, and the tail entity, c respectively, and the corresponding specific text chunk (chunk). Subsequently, all extracted triples are converted into node-edge structures, a knowledge graph database is constructed, and a graph database (e.g., Neo4j-community) is used for storage, so as to facilitate subsequent structured graph retrieval.

[0043] During the construction of the structured knowledge graph, necessary data cleaning and deduplication processing can be performed to ensure the quality of the triples. Further, a graph neural network (GNN) can be introduced to pretrain the preliminary structured knowledge graph to capture deep semantic relationships between entity nodes:

[0044] wherein represents the entity vector representation after GNN pretraining, which can be used for indexing of the graph database to improve the efficiency and accuracy of the subsequent graph retrieval stage.

[0045] Vectorization embedding and index construction of unstructured data: In addition to the structured graph, it is also necessary to effectively manage and use the unstructured raw document data. Therefore, a hybrid index based on sparse-dense vectors can be further constructed to support efficient semantic and keyword joint retrieval.

[0046] Specifically, in one or more embodiments of the present application, the server can obtain unstructured text, extract sparse semantic vectors of each text chunk in the unstructured text based on surface vocabulary through the BGE-Lexical model, and extract dense semantic vectors of each text chunk in the unstructured text based on context deep dependency through the mxbai-embed-large model.

[0047] For example, for a set of unstructured raw documents, each text chunk c ​Two types of models are used to generate sparse and dense vector representations respectively: Sparse Vector Generation (Lexical Sparse Embedding): The BGE-Lexical model can be used to generate sparse semantic vectors for text blocks c Sparse representation vectors are extracted: to obtain their lexical-based sparse semantic expressions. In the formula, is the sparse semantic vector of the text block c .

[0048] Dense Vector Generation (Semantic Dense Embedding): The mxbai-embed-large model can be used to generate dense semantic vectors: to capture deep semantic information, in the formula, is the dense semantic vector of the text block c .

[0049] To efficiently perform subsequent retrieval, further, the sparse vectors are stored in an inverted index-based retrieval engine (such as ElasticSearch) to support fast keyword-based text retrieval; at the same time, the dense vectors are stored in a vector index engine based on approximate nearest neighbor (ANN) algorithm (such as FAISS) to support semantic-level similarity matching. The two types of index methods complement each other, taking into account the accuracy, semantic coverage, and response efficiency of the retrieval, achieving high-quality candidate knowledge acquisition.

[0050] 2. Dual-channel hybrid retrieval phase.

[0051] This phase effectively addresses the shortcomings of traditional RAG methods in information richness and reasoning accuracy through the synergistic action of structured graph retrieval and unstructured hybrid retrieval. Next, the specific implementation and theoretical details of these two sub-phases are introduced in detail.

[0052] Structured Graph Retrieval and Preliminary Verification: In practical applications, the server can obtain the user's query question q Based on the query question, the server can first perform structured graph retrieval to quickly obtain knowledge information directly related to the query question from the structured knowledge graph.

[0053] In one or more embodiments of the present invention, the query language (Cypher) provided by the graph database (Neo4j) can be used to perform efficient queries and obtain preliminary candidate sub-graphs The query here involves the matching of entity nodes and relations to retrieve the graph structure information that is explicitly related to the user's query question. After obtaining the candidate subgraph, the invention introduces a "structure context sufficiency estimation module" based on a large language model to determine whether the current candidate subgraph can effectively support the answer to the query question.

[0054] For example, the candidate subgraph is first converted into a set of triples in natural language form Then a lightweight language model LLM CSE is used to determine the sufficiency of the matching of the query question and the set of triples, and the output is a logical binary decision (Yes or No):

[0055] Specifically, specific prompt words can be designed, and the matching of the set of triples and the query question is determined by logical reasoning based on a large language model, that is, the large language model is guided to determine whether the existing structured information is sufficient to support the answer to the query question in a simple logical reasoning manner. If the CES module returns Yes, it means that the structured information of the current knowledge graph is sufficient to support a high-quality answer, and the candidate subgraph is directly input to the generation model for the next stage of answer generation: If the CES module returns No, it means that the structured information is insufficient, and the server can start the next stage of unstructured hybrid retrieval to expand more relevant context information: .

[0056] Sparse-dense hybrid unstructured retrieval: When the structured graph information is insufficient to effectively answer the question, the server can start the sparse-dense hybrid retrieval module (HySD-GR) to further supplementary retrieval on the text blocks of unstructured text to enrich the context of the candidate subgraph. The HySD-GR module makes full use of semantic retrieval and keyword matching to improve the retrieval accuracy and recall rate of unstructured data.

[0057] Sparse signal calculation: In order to fully capture the explicit lexical matching information between the query question and the unstructured text, in one or more embodiments of the invention, the server can use the BG model (BGE-M3) to generate the sparse keyword vector corresponding to the query question, and for a given query question q and the document block cThe matching score between the sparse semantic vector and the sparse keyword vector is determined by the dot product method using the following formula: In the formula, It is a sparse keyword vector. For text blocks c The matching value between the sparse semantic vector and the sparse keyword vector. Here Vector sum All vectors were generated by the BGE model, and corresponding instructions have been given in the data preprocessing.

[0058] Dense Signal Calculation: Simultaneously, to capture the potential semantic relevance between the query question and unstructured text, in one or more embodiments of this invention, the server can use the mxbai-embed-large model to generate a dense semantic embedding (i.e., a dense question vector) for the query question, and determine the matching degree between the dense semantic vector and the dense question vector based on cosine similarity using the following formula: [Formula omitted for brevity] ; In the formula, For dense problem vectors, For text blocks c The matching value between the dense semantic vector and the dense question vector.

[0059] Sparse-Dense Signal Fusion: Furthermore, the sparse matching score and the dense matching score can be linearly combined using the following formula to determine the mixed retrieval score for each text block: In the formula, As weight, For text blocks c The combined search score. Parameters This can be optimized through experiments, aiming to balance keyword matching and semantic relevance matching, thereby improving the overall retrieval quality of unstructured text.

[0060] Subsequently, based on the hybrid retrieval score, the text blocks in the unstructured text can be sorted and the top-scoring blocks can be selected. k One text block: In the formula, The top score in mixed search indicates the highest score. k A text block.

[0061] Transformation and encoding of unstructured data into graph structures: Further, the selected set of text blocks can be... Transform into graph structure (first get the graph nodes in the text block, and then construct the edges between the nodes through semantic correlation), and expand the candidate subgraph through the graph structure: To further capture and strengthen the high-order neighborhood information of the nodes in the expanded candidate subgraph, a GNN model can also be used for context graph encoding: The finally encoded context subgraph Will be sent to the next stage (context pruning and answer generation stage) for subsequent processing.

[0062] 3. Context pruning and generation stage (MST-guided compression and generation).

[0063] After structured graph retrieval and unstructured mixed retrieval, GRAFT-RAG obtains a semantic subgraph constructed by fusing multiple sources of information (structured information and unstructured information) To alleviate the context window pressure of large language models (LLM) and strengthen the accuracy and controllability of the generation stage, further, in one or more embodiments of the present application, a context pruning mechanism and self-verification reasoning process can also be introduced to further filter and verify the information, thereby generating a reliable final answer. The context pruning mechanism pseudo code is shown in Table 1:

[0064] Table 1 Context pruning mechanism pseudo code Wherein, ConnectedComponents(G) represents all connected subgraph components obtained.

[0065] Maximum spanning tree guided subgraph pruning (MST-guided context pruning): To alleviate the problem of a large number of redundant entities and relationships in the encoded context subgraph G encoded In one or more embodiments of the present application, a maximum spanning tree (Maximum Spanning Tree, MST) strategy can be introduced to structurally prune the information from a graph theory perspective. Specifically, the weight of the edge between each entity pair in the expanded candidate subgraph can be determined according to the similarity of the triples corresponding to the entity pairs and the query question, and a weighted candidate subgraph is obtained. Then, the maximum spanning tree algorithm is used to obtain the spanning tree with the maximum sum of edge weights from the weighted subgraph, as the most relevant entity path set related to the query question, to obtain the compressed candidate subgraph. Then, the compressed candidate subgraph is converted into a triple set, and the answer corresponding to the query question is generated according to the triple set.

[0066] For example, first the query question can be decomposed into a set of sub-questions q The semantic relevance of the path between any entity pair is defined as the weight of the edge in the graph That is, the semantic similarity between the query and the path is taken as the edge weight.

[0067] On this basis, the maximum spanning tree algorithm is performed on the weighted graph G encoded From which the set of entity paths most relevant to the query is extracted, denoted as Thus obtaining the compressed candidate sub-graph.

[0068] Finally, the compressed candidate sub-graph is converted into a structured triple set form for use in the generation phase, the process can be represented as .

[0069] This process not only guarantees semantic coverage, but also effectively compresses redundant information, significantly improving the quality of input in the generation phase and the accuracy of model answers, achieving both form and spirit and increasing efficiency.

[0070] Self-verified reasoning generation (Self-Verified Reasoning Mechanism, SVRM) based on double models: To avoid the spread of errors and hallucination output, the invention GRAFT-RAG introduces a double language model structure to simulate the human "thinking-checking" mechanism. In one or more embodiments of the invention, the server can decompose the query question into multiple sub-questions, and for each sub-question, generate an answer corresponding to the sub-question according to the triple set; when the answer corresponding to the sub-question is judged to be unreasonable by the large language model, the answer corresponding to the sub-question is generated again according to the triple set; finally, the answer corresponding to the query question is generated according to the answers corresponding to each sub-question. Specifically, it can include:

[0071] 1) LLM gen : Perform sub-graph-based reasoning generation; 2) LLM ver : Responsible for self-reflection and consistency verification of the current reasoning result.

[0072] (1) Bottom-up multi-round reasoning: The original question is decomposed into multiple independent solvable sub-questions by the language model, and a reasoning task structure graph Mq reflecting the dependency relationship of the sub-questions is constructed. The graph organizes all sub-question nodes and their dependency relationship edges in the form of a directed acyclic graph, and performs multi-round sub-question solving from bottom to top, and finally synthesizes the answer to the original question. The construction of this structure graph combines natural language analysis, dependency reasoning rules, and language model prompt control to automatically generate. Starting from the leaf nodes, each sub-question is solved step by step from bottom to top by the following formulaq t :

[0073] wherein, is to generate a large language model LLM gen According to the sub-problems q t , a set of triples and previously solved sub-problems answers solved sub-problems q t current answers. .

[0074] (2) Answer verification and rethinking mechanism: current answers by a verification large language model LLM ver to determine its rationality: If is judged to be invalid, the rethinking mechanism is triggered to regenerate the answer: (3) Final answer generation: finally, the root node q 0= q answer is composed of all sub-problem answers: To avoid generating illusions, an explicit prompt can be further set to guide the model to return "I don't know" when there is not enough basis, thereby improving the credibility and robustness of the answer.

[0075] Based on the question and answer method based on retrieval enhancement generation shown in Figure 1 , the application first matches the relevant candidate sub-graph in the existing structured knowledge graph according to the user's query problem, and further judges whether the candidate sub-graph is sufficient to generate an answer to the query problem through logical reasoning. When the candidate sub-graph is not sufficient to cope with the query problem, the application further extracts the sparse keyword vector based on the surface vocabulary and the dense problem vector based on the context deep dependence of the query problem, and matches it with the previously obtained unstructured embedding vector. The unstructured embedding vector includes sparse semantic vectors of each text block in unstructured text based on surface vocabulary and dense semantic vectors of each text block based on context deep dependence, respectively, so as to determine the text block related to the query problem from the unstructured text, and further convert it into a graph structure to supplement the candidate sub-graph, and generate the answer corresponding to the query problem based on this, which improves the performance of question and answer in multi-hop reasoning and information integration retrieval recall.

[0076] In the application based on retrieval enhancement generation of the question and answer method, the query question of the user, the structured knowledge graph constructed in advance according to the document database and the unstructured embedding vector can be obtained without being executed according to the order of each step shown in the above embodiment, and the execution order of each step can be determined according to the needs, and the application does not limit this. Figure 1 The execution order of each step can be determined according to the needs, and the application does not limit this.

[0077] The above is the question and answer method based on retrieval enhancement generation provided by one or more embodiments of the application. Based on the same idea, the application also provides a corresponding question and answer device based on retrieval enhancement generation, as shown in the above embodiment. Figure 3

[0078] Figure 3 The question and answer device based on retrieval enhancement generation provided by the application is shown in the above embodiment. The acquisition module 201 is configured to acquire a query question of a user, a structured knowledge graph constructed in advance according to a document database and an unstructured embedding vector. The unstructured embedding vector includes a sparse semantic vector of each text block in the corresponding document database based on surface vocabulary and a dense semantic vector of each text block based on context deep dependence. The structured retrieval module 202 is configured to retrieve at least part of the graph structure information matched with the query question from the structured knowledge graph constructed in advance, to obtain a candidate subgraph corresponding to the query question, and to convert the candidate subgraph into a triple set. The question feature extraction module 203 is configured to extract a sparse keyword vector of the query question based on surface vocabulary and a dense question vector based on context deep dependence when the matching of the triple set and the query question is determined to be less than a preset threshold through logical reasoning. The unstructured retrieval module 204 is configured to determine the mixed retrieval score of each text block according to the matching of the sparse semantic vector and the sparse keyword vector and the matching of the dense semantic vector and the dense question vector, to sort each text block in descending order according to the mixed retrieval score, and to convert a preset number of text blocks into a graph structure to expand the candidate subgraph. The generation module 205 is configured to generate an answer corresponding to the query question according to the expanded candidate subgraph and to feed back to the user.

[0079] The specific limitations of the question and answer device based on retrieval enhancement generation can be referred to the limitations of the question and answer method based on retrieval enhancement generation in the above, which will not be repeated here. Each module in the above question and answer device based on retrieval enhancement generation can be realized by software, hardware and their combination. Each module can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so as to call and execute the operations of each module by the processor.

[0080] ​The application further provides a computer readable storage medium, which stores a computer program. Figure 1 The application provides a question and answer method based on retrieval enhancement generation.

[0081] The application further provides Figure 4 The application further provides a computer device as shown in the structural schematic diagram of the computer device. Figure 4 As shown in the structural schematic diagram of the computer device, the computer device comprises a processor, an internal bus, a network interface, a memory and a non-volatile memory, and can further comprise other hardware required by a business. Figure 1 The application provides a question and answer method based on retrieval enhancement generation.

[0082] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware, and the computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiment methods can be included. In the embodiments of the present application, any reference to the memory, storage, database or other medium can include at least one of the non-volatile and volatile memories. The non-volatile memory can include a read-only memory (ROM), a magnetic tape, a floppy disk, a flash memory or an optical memory. The volatile memory can include a random access memory (RAM) or an external cache memory. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0083] The technical features of the above embodiments can be combined in any way, and to make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the range disclosed by the present application.

Claims

1. A method for generating a question and answer based on search enhancement, characterized in that, The method comprises the following steps: acquiring a query question of a user, a pre-constructed structured knowledge graph, and an unstructured embedding vector; the unstructured embedding vector comprises sparse semantic vectors of each text block in unstructured text based on surface vocabulary and dense semantic vectors of each text block based on context deep dependency; retrieving at least part of graph structure information matching the query question from the pre-constructed structured knowledge graph to obtain a candidate subgraph corresponding to the query question, and converting the candidate subgraph into a triple set; when the matching of the triple set and the query question is determined to be less than a preset threshold through logical reasoning, extracting a sparse keyword vector of the query question based on surface vocabulary and a dense question vector based on context deep dependency; determining a hybrid retrieval score of each text block according to the matching of the sparse semantic vectors and the sparse keyword vector and the matching of the dense semantic vectors and the dense question vector, and performing descending order sorting on each text block according to the hybrid retrieval score; and converting a preset number of text blocks into a graph structure to expand the candidate subgraph; generating an answer corresponding to the query question according to the expanded candidate subgraph and feeding back to the user.

2. The method for generating question and answer based on retrieval enhancement according to claim 1, wherein, The matching of the triple set and the query question through logical reasoning specifically comprises: determining the matching of the triple set and the query question through logical reasoning based on a large language model.

3. The method for generating question and answer based on retrieval enhancement of claim 1, wherein, The determination of the hybrid retrieval score of each text block according to the matching of the sparse semantic vectors and the sparse keyword vector and the matching of the dense semantic vectors and the dense question vector specifically comprises: the matching of the sparse semantic vectors and the sparse keyword vector is determined by the following formula: ; the matching of the dense semantic vectors and the dense question vector is determined by the following formula: ; A mixed search score of each text block is determined from the matching of the sparse semantic vector and the sparse keyword vector and the matching of the dense semantic vector and the dense problem vector by the following equation: ; in, It is a sparse keyword vector. For text blocks c sparse semantic vectors, For text blocks c The matching value between the sparse semantic vector and the sparse keyword vector. For dense problem vectors, For text blocks c Dense semantic vectors, For text blocks c The matching value between the dense semantic vector and the dense question vector. As weight, For text blocks c The combined search score.

4. The question-answering method based on retrieval enhancement as described in claim 1, characterized in that, The generation of the answer corresponding to the query question according to the expanded candidate subgraph specifically comprises: determining the weight of the edge between each entity pair in the expanded candidate subgraph according to the similarity of the triple corresponding to the entity pair and the query question, to obtain a weighted candidate subgraph; obtaining a spanning tree with the maximum sum of edge weights as the most relevant entity path set related to the query question through a maximum spanning tree algorithm based on the weighted subgraph, to obtain a compressed candidate subgraph; converting the compressed candidate subgraph into a triple set, and generating the answer corresponding to the query question according to the triple set.

5. The method for generating question and answer based on retrieval enhancement according to claim 4, wherein, The generation of the answer corresponding to the query question according to the triple set specifically comprises: dissolving the query question into multiple sub-questions, and generating the answer corresponding to each sub-question according to the triple set; when the answer corresponding to the sub-question is determined to be unreasonable through a large language model, generating the answer corresponding to the sub-question again according to the triple set; generating the answer corresponding to the query question according to the answers corresponding to each sub-question.

6. The method for generating question and answer based on retrieval enhancement of claim 1, wherein, The construction of the structured knowledge graph specifically comprises: acquiring a document database, extracting knowledge triples from the document database through a large language model, and performing preprocessing including data cleaning and deduplication processing on the knowledge triples; the knowledge triples comprise a head entity, a tail entity, a relationship between the head entity and the tail entity, and a corresponding text block; converting the preprocessed knowledge triples into a knowledge graph according to the relationship between the head entity and the tail entity.

7. The method of claim 1, wherein the query is generated based on a search enhancement. The unstructured embedding vector is constructed, and specifically includes: ​ The unstructured text is acquired, sparse semantic vectors of each text block in the unstructured text based on surface vocabulary are extracted through a BGE-Lexical model, and dense semantic vectors of each text block in the unstructured text based on deep context dependency are extracted through an mxbai-embed-large model.

8. A question answering apparatus based on retrieval augmentation generated, characterized by, The method comprises: An acquisition module is configured to acquire a query question of a user, a structured knowledge graph constructed in advance according to a document database, and an unstructured embedding vector; The unstructured embedding vector includes sparse semantic vectors of each text block in the corresponding document database based on surface vocabulary and dense semantic vectors of each text block based on deep context dependency; A structured retrieval module is configured to retrieve at least part of graph structure information matching the query question from the structured knowledge graph constructed in advance, to obtain a candidate subgraph corresponding to the query question, and to convert the candidate subgraph into a triple set; A question feature extraction module is configured to extract a sparse keyword vector of the query question based on surface vocabulary and a dense question vector based on deep context dependency when the matching of the triple set and the query question is determined to be less than a preset threshold through logical reasoning; An unstructured retrieval module is configured to determine a hybrid retrieval score of each text block according to the matching of the sparse semantic vectors and the sparse keyword vector and the matching of the dense semantic vectors and the dense question vector, to sort each text block in descending order according to the hybrid retrieval score, and to convert a preset number of text blocks into a graph structure to expand the candidate subgraph; A generation module is configured to generate an answer corresponding to the query question according to the expanded candidate subgraph and to feed back the answer to the user.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is executed by the processor to implement the method in any one of claims 1-7.

10. A computer device, comprising: The computer program is stored in the memory and can be run on the processor, and the processor implements the method in any one of claims 1-7 when executing the program. The computer program is stored in the memory and can be run on the processor, and the processor implements the method in any one of claims 1-7 when executing the program.

Citation Information

Cited By

  • Retrieval enhancement generation method and system based on multi-dimensional reordering

    CN121681787A

  • Deep semantic matching-based knowledge base content accurate generation method and system

    CN121979905A

  • A Search Enhancement Generation Method and Electronic Device Based on Directed Acyclic Graphs

    CN122414269A