Complex reasoning method based on retrieval enhanced verification and improvement
Patent Information
- Application Number
- CN202411844363.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has logical errors, factual illusions or inconsistencies when handling complex queries, and the single-chain reasoning process can easily lead to error accumulation and it is difficult to get the optimal solution; at the same time, directly inputting external knowledge to the model may lead to knowledge conflicts, affecting the accuracy of generated answers.
A complex inference method based on retrieval enhancement verification and improvement is adopted, combined with Monte Carlo tree search algorithm and retrieval enhancement technology, a comprehensive exploration of the solution space through the tree search algorithm is used, and the retrieved information is used as external guidance for verification and correction, avoiding knowledge conflicts, and improving the depth and accuracy of reasoning.
Through the use of tree search algorithm, we can explore and solve space more comprehensively, reduce errors in multi-step inference, improve the consistency and accuracy of large models in the inference process, and enable them to better deal with complex problems.
Smart Images

Figure CN119990302A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a complex reasoning method based on retrieval enhancement verification and improvement. Background Art
[0002] To answer users' complex queries, existing solutions usually combine thought chain reasoning technology and retrieval enhancement technology to complete knowledge-intensive complex reasoning. Specifically, this type of method first iteratively decomposes the original complex question and obtains multiple sub-questions. For each sub-question, the model can make necessary rewrites or rely on internal knowledge to generate partial responses to enrich the contextual information for retrieval. Then, for each sub-question, documents are retrieved from the foreign corpus as input to the model, allowing the model to solve the sub-questions. Finally, the model summarizes the answers to each sub-question and obtains the predicted output. The latest research attempts to expand the number of steps in the reasoning process in the above method, as well as the number of documents in the retrieval process, to explore the relationship between reasoning latency and solution performance.
[0003] There are two main limitations of existing technologies when processing complex queries. First, the decomposition process relies solely on the internal knowledge of the model, which may cause logical errors, factual illusions or inconsistencies in a certain sub-problem. Furthermore, as the problem becomes more and more complex, relying on a single-chain reasoning process will amplify the above-mentioned error accumulation, which may easily lead to the failure to obtain the optimal solution in the end. It is necessary to extend it to a tree-structured reasoning process to support the exploration of more complex problem solving. On the other hand, when solving each sub-problem, although external knowledge can be introduced to make up for the knowledge limitations of the model through retrieval enhancement generation, existing retrieval enhancement generation methods usually directly input external knowledge into the model. The model needs to process internal and external knowledge at the same time, which is prone to knowledge conflicts and leads to the generation of wrong answers. Although it has been explored that increasing the reasoning delay can help the performance of retrieval enhancement generation to a certain extent, there has not yet been an effective solution to the above two problems.
[0004] The information disclosed in this background technology section is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as acknowledging or suggesting in any form that the information constitutes the prior art already known to those skilled in the art. Summary of the invention
[0005] In view of the problems existing in the prior art, the object of the present invention is to provide a complex reasoning method based on retrieval enhancement verification and improvement.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] A complex reasoning method based on retrieval enhancement verification and improvement, the complex reasoning method comprising the following steps:
[0008] S1, the root node is a multi-hop problem with input, and the initialization root node is an original problem with output, that is, the problem to be solved;
[0009] S2. Start a search simulation: use the Monte Carlo tree search algorithm to perform a search simulation;
[0010] S3, select a candidate node from the tree for exploration;
[0011] S4, expand the child nodes of the selected candidate node;
[0012] S5. Calculate rewards for the expanded child nodes;
[0013] S6, reward is sent back, the node value of the entire tree is updated, and a simulation ends;
[0014] S7. Repeat the above process until the maximum number of simulations is reached or the end node is obtained.
[0015] Further, step S3 is specifically as follows:
[0016] Select a node with the highest value from the child nodes obtained in the previous simulation, and then select the next best child node along the tree layer until the leaf node is reached, which is the final state representing the final answer; in order to better balance exploration and utilization, the UCT algorithm is used to calculate the value of the node based on the number of visits and expected rewards of each node. The calculation formula is as follows:
[0017] ;
[0018] in, Indicates the current node, Indicates the parent node of the current node. Represents node rewards, represents the exploration coefficient, Indicates the number of visits.
[0019] Further, step S4 is specifically as follows:
[0020] After selecting a node, the search tree is expanded by repeatedly sampling multiple child nodes. Since each node includes subquery and answer attributes, the expansion process includes two steps: subquery generation and answer derivation.
[0021] The subquery generation is specifically as follows: firstly, context information is constructed by concatenating state information from the root node to the currently selected node, and then the strategy model is instructed to sample the next subquery according to the context information;
[0022] The answer derivation is specifically as follows: after generating the subqueries, further instructing the strategy model to generate answers to explore the intrinsic knowledge of the LLM; for each subquery, using the historical context and the subquery as input to prompt the strategy model to generate candidate answers by utilizing the intrinsic knowledge encoded in its parameters, and in this process, retrieval knowledge from the outside is not considered to avoid knowledge conflicts;
[0023] After the expansion operation is completed, each parent node obtains multiple child nodes, each of which contains a subquery and its answer.
[0024] Further, step S5 is specifically as follows:
[0025] Node rewards involve two types of reward scores, including answer-aware rewards and query-aware rewards;
[0026] The answer-aware reward is specifically as follows: given a node subquery, retrieve the top K documents from the external corpus; based on the retrieved documents, use the reward model to assign an answer-aware reward to the currently generated answer; specifically, for the knowledge consistency between the generated answer and the retrieved document, there are three cases and different reward values, as shown below:
[0027] ;
[0028] in, represents the answer-aware reward, Indicates The answer information stored in the layer nodes, Represents the retrieved document collection;
[0029] In the second case, when the generated answer conflicts with the document, a medium score is assigned to the answer, and the new potential answer from external knowledge is used to correct the original answer, supporting the policy model to continue reasoning from the current node; however, if the generated answer cannot be verified by external knowledge, the lowest score will be assigned to avoid the policy model from exploring the solution space that may be risky;
[0030] The query-aware reward is specifically: using the historical context information from the root node to the current node to measure the rationality of the subquery; if the subquery evaluated by the reward model is logically inconsistent with the historical plan, the score is 0; otherwise, the score is 1;
[0031] The node reward is obtained based on the product of the above two rewards.
[0032] Further, step S6 is specifically as follows:
[0033] After obtaining the newly expanded node reward, the reward is back-propagated to update the value from the root node to the current node; for each node in the path, its visit count and node value will be updated according to the following formula:
[0034] ;
[0035] in, Indicates the current node. Indicates the number of visits that have not been updated. Indicates the updated number of visits. Indicates the node rewards that have not been updated. Indicates the updated node rewards, Represents the child node reward.
[0036] By adopting the above technical solution, the present invention has the following beneficial effects:
[0037] Existing methods mainly rely on a single-chain multi-step reasoning process, which limits the exploration space of the model during reasoning. The model mainly relies on the model's intrinsic knowledge to perform multi-step decomposition of the problem, and lacks corresponding verification and correction strategies, which may lead to insufficient confidence in the solution, such as outdated, erroneous and other problems. On the other hand, although the direct use of retrieval-enhanced generation can introduce external knowledge to make up for the knowledge limitations of the model, it may cause conflicts between internal and external knowledge. Compared with these methods, the solution proposed in this patent combines retrieval-enhanced verification and tree search algorithms to solve the limitations of sequential reasoning structures and knowledge conflicts, and promote the performance of large models in complex reasoning tasks. Through the tree search algorithm to fully explore the solution space, the retrieved information is used as external guidance for verification and correction, and integrated into the reasoning process, thereby avoiding knowledge conflicts and improving the depth and accuracy of reasoning, reducing model errors in multi-step reasoning, and improving the consistency and accuracy of large models in the reasoning process, so that they can better handle complex problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0039] Figure 1 A flow chart of the complex reasoning method based on retrieval-enhanced verification and improvement provided by the present invention.
[0040] Figure 2 This is an overall architecture diagram of the search reasoning technology combined with retrieval enhancement verification and improvement provided by the present invention.
[0041] Figure 3 A schematic diagram of an actual case provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0042] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0043] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.
[0044] Combination Figure 1 As shown, the present invention proposes a complex reasoning method based on retrieval enhancement verification and improvement, and the complex reasoning method includes the following steps:
[0045] Step 1: A multi-hop problem with the root node as input
[0046] Initializing the root node is an original problem of output, that is, the problem that needs to be solved.
[0047] Step 2: Start a search simulation
[0048] Use the Monte Carlo tree search algorithm to perform search simulations.
[0049] Step 3: Select a candidate node from the tree to explore
[0050] Select a node with the highest value from the child nodes obtained in the previous simulation, and then select the next best child node along the tree layer until you reach the leaf node, which is the final state representing the final answer. In order to better balance exploration and utilization, the UCT algorithm is used to calculate the value of the node based on the number of visits and expected rewards of each node, as shown below:
[0051] ;
[0052] in, Indicates the current node, Indicates the parent node of the current node. Represents node rewards, represents the exploration coefficient, Indicates the number of visits.
[0053] Step 4: Expand the child nodes of the selected candidate nodes
[0054] After selecting a node, the search tree is expanded by repeatedly sampling multiple child nodes. Since each node includes subquery and answer attributes, the expansion process includes two steps: subquery generation and answer derivation.
[0055] Subquery generation: First, context information is built by concatenating the state information from the root node to the currently selected node, and then the policy model is instructed to sample the next subquery based on the context information.
[0056] Answer derivation: After generating the subqueries, the application further instructs the strategy model to generate answers to explore the intrinsic knowledge of the LLM. In particular, for each subquery, the application takes the historical context and the subquery as input to prompt the strategy model to generate candidate answers by leveraging the intrinsic knowledge encoded in its parameters. In this process, the application does not consider the retrieval knowledge from the outside to avoid knowledge conflicts.
[0057] After the expansion operation is completed, each parent node obtains multiple child nodes, each of which contains a subquery and its answer.
[0058] Step 5: Calculate rewards for the expanded child nodes
[0059] Node rewards involve two types of reward scores, including answer-aware rewards and query-aware rewards.
[0060] Answer-aware reward: Given a subquery of a node, retrieve the top K documents from the external corpus. Based on the retrieved documents, use the reward model to assign an answer-aware reward to the currently generated answer. Specifically, there are three cases and different reward values for the knowledge consistency between the generated answer and the retrieved document, as shown below:
[0061] ;
[0062] in, represents the answer-aware reward, Indicates The answer information stored in the layer nodes, Represents a collection of retrieved documents.
[0063] In the second case (i.e., when the generated answer conflicts with the document), this application assigns a medium score to the answer and uses the new potential answer from external knowledge to correct the original answer, supporting the policy model to continue reasoning from the current node. However, if the generated answer cannot be verified by external knowledge, this application will assign the lowest score to avoid the policy model from exploring the solution space that may be risky.
[0064] Query-aware reward: Use the historical context information from the root node to the current node to measure the rationality of the subquery. If the subquery evaluated by the reward model is logically inconsistent with the historical plan, the score is 0; otherwise, the score is 1.
[0065] The node reward is obtained based on the product of the above two rewards.
[0066] Step 6: Reward feedback, update the node value of the entire tree, and end a simulation
[0067] After obtaining the newly expanded node reward, the reward is back-propagated to update the value from the root node to the current node. For each node in the path, its visit count and node value will be updated according to the following formula:
[0068] ;
[0069] in, Indicates the current node. Indicates the number of visits that have not been updated. Indicates the updated number of visits. Indicates the node rewards that have not been updated. Indicates the updated node rewards, Represents the child node reward.
[0070] Step 7: Repeat the above process until the maximum number of simulations is reached or the end node is obtained.
[0071] Retrieval-augmented generation (RAG) has become an indispensable technology to solve the inherent knowledge limitations of large models, effectively combining the required information with reliable sources. However, existing technical solutions mainly rely on chain-like problem decomposition and directly use RAG to provide supplementary knowledge, while ignoring the in-depth research of RAG in enhancing the complex reasoning ability of LLM. This patent fully explores the optimal combination strategy of internal knowledge and external knowledge of large models in the complex reasoning process through the verification and modification of Monte Carlo tree search algorithm and retrieval enhancement.
[0072] Combination Figure 2 As shown, it mainly includes two technical steps. First, this application proposes a tree-based sub-query multi-step decomposition process, which expands the traditional chain reasoning into a reasoning process on the tree, and enhances the exploration and utilization of the large language model during reasoning through the Monte Carlo tree. This process mainly relies on the intrinsic knowledge of the LLM. Furthermore, in order to enhance the confidence of the reasoning process, this application designs verification and modification based on retrieval enhancement, using external knowledge to help determine whether the model correctly uses internal knowledge to answer, and when necessary, uses external knowledge to make corrections to support the subsequent reasoning process of the model.
[0073] Specifically, under this framework, the model first selects a node from the tree to explore, then generates the next subquery and answer to obtain a new child node, and calculates the reward for the expanded node. Finally, the reward is back-propagated to update the value of the parent node on the tree. This process will be iterated until the task is completed.
[0074] Training and Inference:
[0075] Using knowledge distillation techniques, we transfer the capabilities of advanced LLMs with more parameters into relatively smaller models. This consists of two stages: data synthesis and instruction fine-tuning.
[0076] In the data synthesis stage, data from the training set are mixed to maintain diversity. First, a contextual learning guidance strategy model is used to generate a solution in the chain reasoning (CoT) format and decompose it into multiple sub-steps, each of which contains the input question, the accumulated reasoning path, and the sub-query specific to the current step. To further ensure diversity, this application randomly selects a step from each sample for subsequent instruction data creation. Then, this application uses a more advanced LLM combined with a retrieval system to evaluate the sub-queries and their answers for each step, and filters out outputs that fail to meet the format requirements. Finally, this application compiles a dataset containing intermediate steps and their queries and answer rewards.
[0077] During the instruction fine-tuning phase, this application uses synthetic samples to fine-tune a smaller LLM to enhance its capabilities in reward modeling.
[0078] After training the reward model, search for an optimal reasoning path on the inference tree according to the above process.
[0079] During the entire search process, LLM initializes the input problem as the root node and performs multiple simulations to eventually reach the terminating leaf node. This process can be represented as a tree.
[0080] like Figure 3As shown in the figure, given a complex question (i.e., Input Question), in the first simulation, the tree search algorithm will expand from the root node, and the sampling strategy will allow the policy model to generate multiple sub-questions of the first step, and then each sub-question will be verified and modified through retrieval enhancement (for example, the second sub-question of the first step is successfully verified by external documents). In the second simulation, the tree search algorithm selects the second query of the first step (i.e., Who is the director of "Life Hits"?) for expansion according to the UCT formula, and LLM expands multiple sub-nodes through multiple sampling. Similarly, the model corrects the generated answer (i.e., "1998") of the sub-query "When was Christian born?" based on the retrieved documents, and the reward model returns an overall score of 2 for this. Through multiple rounds of multi-step reasoning and retrieval-enhanced verification process iterations, the model finally outputs the correct answer (i.e., "Life Hits"). In the task solving process, LLM generates the answer to the current sub-query based on its internal knowledge, but due to the time limit of the pre-training corpus or memory errors, the answer may be biased. Therefore, external knowledge can effectively verify the correctness of LLM's internal knowledge, thereby guiding the model to plan a reasonable path.
[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A complex reasoning method based on retrieval-enhanced verification and improvement, characterized in that: The complex reasoning method comprises the following steps: S1, the root node is a multi-hop problem with input, and the initialization root node is an original problem with output, that is, the problem to be solved; S2. Start a search simulation: use the Monte Carlo tree search algorithm to perform a search simulation; S3, select a candidate node from the tree to explore; S4, expand the child nodes of the selected candidate node; S5. Calculate rewards for the expanded child nodes; S6, reward is sent back, the node value of the entire tree is updated, and a simulation ends; S7. Repeat the above process until the maximum number of simulations is reached or the end node is obtained.
2. The complex reasoning method based on retrieval-enhanced verification and improvement according to claim 1, characterized in that: Step S3 is specifically as follows: Select a node with the highest value from the child nodes obtained in the previous simulation, and then select the next best child node along the tree layer until the leaf node is reached, which is the final state representing the final answer; in order to better balance exploration and utilization, the UCT algorithm is used to calculate the value of the node based on the number of visits and expected rewards of each node. The calculation formula is as follows: ; in, Indicates the current node, Indicates the parent node of the current node. Represents node rewards, represents the exploration coefficient, Indicates the number of visits.
3. The complex reasoning method based on retrieval-enhanced verification and improvement according to claim 1, characterized in that: Step S4 is specifically as follows: After selecting a node, the search tree is expanded by repeatedly sampling multiple child nodes. Since each node includes subquery and answer attributes, the expansion process includes two steps: subquery generation and answer derivation. The subquery generation is specifically as follows: firstly, context information is constructed by concatenating state information from the root node to the currently selected node, and then the strategy model is instructed to sample the next subquery according to the context information; The answer derivation is specifically as follows: after generating the subqueries, further instructing the strategy model to generate answers to explore the intrinsic knowledge of the LLM; for each subquery, using the historical context and the subquery as input to prompt the strategy model to generate candidate answers by utilizing the intrinsic knowledge encoded in its parameters, and in this process, retrieval knowledge from the outside is not considered to avoid knowledge conflicts; After the expansion operation is completed, each parent node obtains multiple child nodes, each of which contains a subquery and its answer.
4. The complex reasoning method based on retrieval-enhanced verification and improvement according to claim 1 is characterized in that: Step S5 is specifically as follows: Node rewards involve two types of reward scores, including answer-aware rewards and query-aware rewards; The answer-aware reward is specifically as follows: given a node subquery, retrieve the top K documents from the external corpus; based on the retrieved documents, use the reward model to assign an answer-aware reward to the currently generated answer; specifically, for the knowledge consistency between the generated answer and the retrieved document, there are three cases and different reward values, as shown below: ; in, represents the answer-aware reward, Indicates The answer information stored in the layer nodes, Represents the retrieved document collection; In the second case, when the generated answer conflicts with the document, a medium score is assigned to the answer, and the new potential answer from external knowledge is used to correct the original answer, supporting the policy model to continue reasoning from the current node; however, if the generated answer cannot be verified by external knowledge, the lowest score will be assigned to avoid the policy model from exploring the solution space that may be risky; The query-aware reward is specifically: using the historical context information from the root node to the current node to measure the rationality of the subquery; if the subquery evaluated by the reward model is logically inconsistent with the historical plan, the score is 0; otherwise, the score is 1; The node reward is obtained based on the product of the above two rewards.
5. The complex reasoning method based on retrieval-enhanced verification and improvement according to claim 1 is characterized in that: Step S6 is specifically as follows: After obtaining the newly expanded node reward, the reward is back-propagated to update the value from the root node to the current node; for each node in the path, its visit count and node value will be updated according to the following formula: ; in, Represents the current node, Indicates the number of visits that have not been updated. Indicates the updated number of visits. Indicates the node rewards that have not been updated. Indicates the updated node rewards, Represents the child node reward.
Citation Information
Cited By
Large model reasoning system, method and equipment based on Monte Carlo tree search
CN120354953A
Multimodal error information detection method and device
CN120851146B
Electric power NL2SQL exploration optimization method based on Monte Carlo tree search
CN121560933A