Self-adaptive retrieval enhancement generation method based on topic filtering and iterative reasoning
By introducing topic filtering and iterative inference technology into the RAG model, using the BERTopic model and vector searcher, the problem of inefficiency in RAG model when dealing with complex multi-hop queries is solved, multi-document access and query difficulty evaluation is realized, and the response speed of the question-and-answer system and the accuracy of the answers are improved.
Patent Information
- Application Number
- CN202510048758.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing RAG models are inefficient when handling complex multi-hop queries, cannot effectively access the entire document text, and lack consideration for the difficulty of query answers.
Adaptive retrieval enhancement generation method based on topic filtering and iterative inference is adopted, and topics related to input queries are predicted using pre-trained BERTopic model, documents with high similarity are extracted through vector searchers, and long-term memory-dependent networks are used for iterative inference, and answers are finally generated and evaluated.
It improves the efficiency of handling complex multi-hop queries, realizes multi-document access, evaluates and adjusts query difficulty, and improves the response speed of the Q&A system and the accuracy of the answers.
Smart Images

Figure CN119961442A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to an adaptive retrieval enhancement generation method based on topic filtering and iterative reasoning. Background Art
[0002] LLM stands for Large Language Model, and RAG stands for Retrieval Enhanced Generation. Although existing RAG models have achieved certain success, they face efficiency issues when processing complex multi-hop queries. In addition, current RAG models also have certain defects: First, most RAG models use a single-step approach to retrieve documents, which means that they cannot access the entire document text in a single process, limiting the model's retrieval capabilities. Finally, existing methods do not take into account the difficulty of answering queries, and it is difficult to automatically adjust to match the complexity of the question. Summary of the invention
[0003] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide an adaptive retrieval enhancement generation method based on topic filtering and iterative reasoning to improve the efficiency of processing complex multi-hop queries in question-answering systems.
[0004] To achieve the above object, the present invention provides the following solutions:
[0005] An adaptive retrieval enhancement generation method based on topic filtering and iterative reasoning, comprising:
[0006] Predicting several topics related to the input query using a pre-trained BERTopic model; the BERTopic model is fine-tuned based on a multi-hop dataset; the multi-hop dataset is constructed and annotated based on domain experts; the fine-tuning includes: gradient descent and cross-validation techniques;
[0007] Calculating the similarity between the documents in the external knowledge base and the input query based on all the topics;
[0008] Using a vector retriever to sequentially sort the similarities, and sequentially extracting a preset number of documents corresponding to the similarities to obtain retrieved documents;
[0009] Iterating the topic and the retrieval document using a long-term dependency memory network based on chain thinking reasoning technology until a maximum number of iterations is reached or the generated answer passes a quality check, thereby obtaining a number of reasoning results;
[0010] Integrate and polish the reasoning results to obtain the target generated answer;
[0011] The target generated answer is evaluated using a preset LLM model according to a preset evaluation standard, and key information in the target generated answer is cross-referenced with the search document to obtain an evaluation result; the key information includes: name, date and event;
[0012] When the evaluation result is lower than a preset threshold, the target user is prompted to reformulate the input query.
[0013] Preferably, the target generated answer format includes: text, table and chart.
[0014] Preferably, the evaluation result ranges from 0 to 10.
[0015] The present invention discloses the following technical effects:
[0016] The present invention provides an adaptive retrieval enhancement generation method based on topic filtering and iterative reasoning. By adopting a BERTopic model that adaptively adjusts a multi-hop data set, the efficiency problem of a traditional RAG model in processing complex multi-hop queries is solved, and a fast response to multi-hop queries is achieved. By using a vector retriever for document extraction, the defect that the traditional RAG model uses a single-step method to retrieve documents, resulting in an inability to access the entire document text, is solved, and multi-document access in the process is achieved. By setting an evaluation feedback process, the problem that the existing RAG model lacks consideration of the difficulty of query answering is solved, and the evaluation of the difficulty of query answering and the adjustment feedback of target user queries are achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0018] Figure 1 A schematic diagram of the adaptive retrieval enhancement generation process provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] The purpose of the present invention is to provide an adaptive retrieval enhancement generation method based on topic filtering and iterative reasoning to improve the efficiency of processing complex multi-hop queries in a question-answering system.
[0021] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] Figure 1 A schematic diagram of the adaptive search enhancement generation process provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown, the present invention provides an adaptive retrieval enhancement generation method based on topic filtering and iterative reasoning, comprising:
[0023] Step 100: Use the pre-trained BERTopic model to predict several topics related to the input query; the BERTopic model is fine-tuned based on the multi-hop dataset; the multi-hop dataset is constructed and annotated based on domain experts; the fine-tuning includes: gradient descent method and cross-validation technique;
[0024] Step 200: Calculate the similarity between the documents in the external knowledge base and the input query based on all topics;
[0025] Step 300: using a vector retriever to sequentially sort the similarities, and sequentially extracting a preset number of documents corresponding to the similarities to obtain retrieved documents;
[0026] Step 400: Iterate the topics and retrieved documents using the long-term dependency memory network based on the chain thinking reasoning technology until the maximum number of iterations is reached or the generated answer passes the quality check, and obtains a number of reasoning results;
[0027] Step 500: Integrate and polish the reasoning results to obtain the target generated answer;
[0028] Step 600: Evaluate the target generated answer using the preset LLM model according to the preset evaluation criteria, and cross-reference the key information in the target generated answer with the search document to obtain the evaluation result; the key information includes: name, date and event;
[0029] Step 700: When the evaluation result is lower than a preset threshold, prompt the target user to reformulate the input query.
[0030] Optionally, the target generated answer format includes: text, table and chart.
[0031] Specifically, the evaluation results range from 0 to 10.
[0032] Furthermore, the BERTopic model is fine-tuned based on a multi-hop question-answering dataset constructed and annotated by domain experts. The dataset contains diverse question and answer pairs from public resources (such as Wikipedia, academic papers, etc.). The gradient descent method is used to adjust the model parameters during the fine-tuning process, and cross-validation technology is used to ensure the generalization ability of the model.
[0033] Preferably, the traditional RAG model lacks consideration of topics when retrieving documents, resulting in inaccurate retrieval results and high computational costs. This embodiment introduces a topic assignment module to dynamically assign one or more topics to the input query, thereby narrowing the retrieval scope and improving retrieval efficiency and accuracy. The BERTopic model is used as the topic assignment model. The BERTopic model is based on the transformer model and can effectively discover topics in documents. It effectively captures the meaning of context words by clustering the embedding vectors of documents and applying a class-based model to create a consistent topic representation.
[0034] Furthermore, for an input query, the BERTopic model predicts the most relevant topics. Topics can represent the domain of the query and filter the document database to improve retrieval accuracy and reduce computational complexity.
[0035] Preferably, for each multi-hop dataset, the pre-trained BERTopic model is fine-tuned using gradient descent, while cross-validation techniques are used to ensure model generalization, and the model is applied during the inference and data ingestion phases. The fine-tuning process adapts the model to the unique characteristics of each dataset, thereby improving topic consistency and retrieval accuracy.
[0036] Furthermore, a vector retriever, such as Chroma, is used to retrieve relevant documents from an external knowledge base based on the assigned topic and query. Using a vector retriever, documents can be efficiently sorted based on their similarity to the query and the most relevant documents can be returned. A vector retriever can sort documents based on the similarity of topics and return the most relevant documents to the assigned topic.
[0037] Optionally, you can enrich the original query based on the topic and query using natural language processing techniques such as synonym replacement, phrase expansion, or query expansion based on semantic similarity to retrieve more relevant documents. For example, you can use WordNet or other vocabulary resources for synonym replacement, or analyze the keywords in the query and combine them with external knowledge bases to generate additional relevant terms to improve the relevance and coverage of the search results.
[0038] Specifically, using LLM, such as GPT-4, iterative reasoning is performed based on the retrieved documents and contextual information to gradually build a comprehensive understanding of the query. Chain of Thought (CoT) Reasoning: Using CoT reasoning technology, LLM can generate intermediate reasoning steps, such as problem decomposition, hypothesis formulation and verification, etc., to gradually build a comprehensive understanding of the query. Reasoning Context: The retrieved documents and intermediate reasoning steps are passed to the LLM as reasoning context so that the LLM can perform further reasoning. Iterative process: The document retrieval and reasoning process is performed iteratively until the maximum number of iterations is reached or the generated answer passes the quality check.
[0039] Furthermore, based on the reasoning results, the final answer is generated using LLM. Answer integration: The intermediate reasoning steps and the final answer are integrated into a coherent answer, and polished using natural language generation technology. Answer format: The answer can be formatted into different forms, such as text, tables, charts, etc., according to different application scenarios.
[0040] Going further, using the answer evaluation module, the quality and relevance of the generated answers are evaluated, and the query is reformulated as needed. Use the LLM model to evaluate the relevance and value of the answers to the user's query. Use an LLM model, such as GPT-4, to evaluate the generated answers. The LLM model can score the answers based on preset evaluation criteria, such as whether the answer contains all the necessary information, whether it is relevant to the query, etc. Scoring range: The scoring range is usually 0 to 10 points, and the higher the score, the more relevant and valuable the answer. Threshold setting: A threshold can be set. If the answer scores below this threshold, the answer is considered irrelevant or of low value and the query needs to be reformulated. Use the LLM model to evaluate the factual accuracy of the answer and cross-reference it with the retrieved documents to reduce the risk of hallucinations.
[0041] Specifically, the LLM model is used to evaluate the generated answers. The LLM model verifies whether the answer is consistent with the facts based on the retrieved documents. Cross-reference: Cross-reference the key information in the answer with the retrieved documents, such as matching names, dates, events, etc. with the document content. Scoring range: The scoring range is usually 0 to 10 points, and the higher the score, the more accurate the answer. Threshold setting: A threshold can be set. If the answer score is lower than the threshold, it is considered that the answer is an illusion and the query needs to be reformulated.
[0042] Further, this embodiment includes the following steps: 1) The use of a topic assignment model, which accepts an input query, analyzes it through an algorithm (such as clustering, classification or regression, etc.), and assigns one or more topics to it. 2) The role of the document retrieval module is to find relevant documents from an external knowledge base based on the user query and the topic obtained in the previous step. Traditional retrieval algorithms can be used here, such as retrieval based on inverted indexes, retrieval based on topic models, etc. 3) The iterative reasoning model uses a machine learning model (such as a deep neural network) to simulate the human reasoning process. The model gradually understands and constructs the meaning of the query through multiple iterations and inferences. 4) Iterative reasoning is performed based on the query and topic and the retrieved documents. At this stage, a long-term dependency memory network will be used for reasoning to understand the meaning and relationship of the query in the document and related context. Through multiple iterations, the model can gradually refine and abstract the core idea of the query. 5) Based on the results of iterative reasoning, the final answer is generated. In addition, an answer evaluation module can also be included to evaluate the quality and relevance of the generated answer. This evaluation may be achieved by measuring the semantic similarity with the target sentence. If the generated answer does not meet the user's needs, appropriate feedback or prompts can be returned to the user to narrow the scope or try again.
[0043] The beneficial effects of the present invention are as follows:
[0044] The present invention improves the efficiency of processing complex multi-hop queries by adopting a BERTopic model that adaptively adjusts multi-hop data sets; realizes multi-document access in the process by using a vector retriever for document extraction; and realizes the evaluation of query answer difficulty and adjustment feedback of target user queries by setting an evaluation feedback process.
[0045] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0046] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of the present invention. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. An adaptive retrieval enhancement generation method based on topic filtering and iterative reasoning, characterized in that: include: Use the pre-trained BERTopic model to predict several topics related to the input query; The BERTopic model will be fine-tuned based on the multi-hop dataset; The multi-hop dataset is constructed and annotated based on domain experts; the fine-tuning includes: gradient descent method and cross-validation technology; Calculating the similarity between the documents in the external knowledge base and the input query based on all the topics; Using a vector retriever to sequentially sort the similarities, and sequentially extracting a preset number of documents corresponding to the similarities to obtain retrieved documents; Iterating the topic and the retrieval document using a long-term dependency memory network based on chain thinking reasoning technology until a maximum number of iterations is reached or the generated answer passes a quality check, thereby obtaining a number of reasoning results; Integrate and polish the reasoning results to obtain the target generated answer; The target generated answer is evaluated using a preset LLM model according to a preset evaluation standard, and key information in the target generated answer is cross-referenced with the search document to obtain an evaluation result; the key information includes: name, date and event; When the evaluation result is lower than a preset threshold, the target user is prompted to reformulate the input query.
2. The adaptive retrieval enhancement generation method based on topic filtering and iterative reasoning according to claim 1 is characterized in that: The target generated answer formats include: text, table and chart.
3. The adaptive retrieval enhancement generation method based on topic filtering and iterative reasoning according to claim 1 is characterized in that: The evaluation results range from 0 to 10.
Citation Information
Patent Citations
Dynamic adaptation question answering system and method based on hierarchical structure and retrieval enhancement
CN118193714A
Knowledge intensive question reasoning and generating method based on LLM
CN118798367A
Document question and answer method, system and equipment based on RAG and medium
CN119003725A
Open domain question and answer method and device, equipment and storage medium
CN119066183A
Neural network-based semantic information retrieval
WO2021237082A1