Large model reasoning system, method and equipment based on Monte Carlo tree search
Through a large-scale model inference system based on Monte Carlo tree search, the explicit thinking chain guidance signal and model operation are used to coordinate operations, and the problems of high computing resource consumption and insufficient flexibility in complex inference tasks are solved, achieving efficient and stable inference performance improvement.
Patent Information
- Application Number
- CN202510828334.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing large language models have problems with high computational resource consumption, insufficient flexibility and lack of explicit guidance signals in mathematical and logical reasoning, resulting in poor performance in complex inference tasks.
A large-model inference system based on Monte Carlo tree search is adopted to filter candidate thinking cards through task disassembly modules, and an explicit thinking chain guidance signal is generated using Monte Carlo tree search, and the inference result evaluation is carried out through the collaborative operation of two large language models to achieve efficient and complex inference.
It improves the performance of large models in complex inference tasks, reduces sensitivity to sample selection, enhances the stability and generalization ability of inference, and improves the efficiency and accuracy of inference.
Smart Images

Figure CN120354953A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a large model inference system, method and device based on Monte Carlo tree search. Background Art
[0002] In recent years, with the improvement of computing power, the growth of data volume and the continuous optimization of algorithms, large language models (LLMs) have demonstrated excellent performance in multiple fields such as education, search engines, customer support and healthcare. Inference ability, especially mathematical and logical reasoning, has become an important indicator for evaluating the basic cognitive ability of large models. Mastering complex reasoning ability requires the model to strictly follow user instructions and adopt sophisticated problem-solving strategies, which poses a severe challenge to current large models. Existing technologies such as distillation or in-context learning technologies either require significant computing resources or lack flexibility and are difficult to scale. Summary of the Invention
[0003] The present invention provides a large model inference system, method and device based on Monte Carlo tree search to at least partially solve the above problems.
[0004] In a first aspect of the present invention, a large model inference system based on Monte Carlo tree search is provided, and the system includes: A task decomposition module, configured to match the problem to be inferred with each thinking card in the thinking card library, and screen out multiple candidate thinking cards; A collaborative inference and verification module, configured to use a first large language model to take the optimal effective inference path corresponding to each candidate thinking card as an explicit thinking chain guidance signal, perform inference on the problem to be inferred, obtain inference results corresponding to each of the multiple candidate thinking cards, and use a second large language model to evaluate the inference results corresponding to each of the multiple candidate thinking cards to obtain an optimal inference result as the answer to the problem to be inferred; Wherein, the optimal effective inference path corresponding to each candidate thinking card is an inference path obtained based on Monte Carlo tree search, including a plurality of chained inference actions, and each inference action represents a thinking mode.
[0005] Optionally, the system further includes: An inference action definition module, configured to define N inference actions; The thinking card construction module based on Monte Carlo tree search is used to abstract each seed problem into a tree search problem, obtain the optimal effective reasoning path corresponding to each seed problem, and use multiple seed problems sharing an optimal effective reasoning path as a thinking card to obtain a thinking card library. Among them, an optimal effective reasoning path means that the answer obtained by reasoning according to this reasoning path matches the correct answer of the seed problem and the reasoning cost is minimized. An optimal effective reasoning path includes a root node and at least one node, each node represents a reasoning action, and the root node represents the seed problem.
[0006] Optionally, abstracting each seed problem into a tree search problem and obtaining the optimal effective reasoning path corresponding to each seed problem includes: During the first Monte Carlo tree search, randomly select multiple reasoning actions as the children nodes of the root node to obtain a first intermediate reasoning result. The randomly selected multiple reasoning actions serve as the first-layer nodes; Based on the first intermediate reasoning result, continue to select multiple reasoning actions for reasoning until the reasoning ends to obtain a reasoning path. The multiple nodes selected for the nth time serve as the nth-layer nodes; According to whether the reasoning result obtained according to each reasoning path matches the correct answer of the seed problem, update the reward values of each node upward along this reasoning path from the leaf node to obtain the reward value of each path; When the reasoning result obtained by a reasoning path matches the correct answer of the seed problem, use this reasoning path as an effective reasoning path; According to the reward values and the number of nodes of multiple effective reasoning paths, screen out the optimal effective reasoning path of this seed problem.
[0007] Optionally, using multiple seed problems sharing an optimal effective reasoning path as a thinking card to obtain a thinking card library includes: Compare the optimal effective reasoning paths of multiple seed problems to determine the set of seed problems sharing the same optimal effective reasoning path; Determine the semantic representation, the number of sub-problems, and the conditional complexity of each seed problem in a set of seed problems; Use the mean of the semantic representations, the mean of the number of sub-problems, and the mean of the conditional complexities of each seed problem in a set of seed problems as the semantic representation, the number of sub-problems, and the conditional complexity of a thinking card.
[0008] Optionally, the task decomposition module is used to: Calculate the semantic representation, the number of sub-problems, and the conditional complexity of the problem to be reasoned; Compare the semantic representations, the number of sub - problems, and the condition complexity of each thinking card in the thinking card library with those of the problem to be inferred, to obtain multiple candidate thinking cards.
[0009] Optionally, for the problem to be inferred, use the first large - language model to perform inference according to the inference path corresponding to each candidate thinking card, to obtain the inference results corresponding to each of the multiple candidate thinking cards, including: Input the problem to be inferred and the text description of the inference path corresponding to each candidate thinking card into the first large - language model to obtain the inference result of each candidate thinking card; Use the second large - language model to evaluate the inference results corresponding to each of the multiple candidate thinking cards, to obtain the optimal inference result as the answer to the problem to be inferred, including: Input the inference result of each candidate thinking card into the second large - language model based on result verification to obtain the inference result score of each candidate thinking card, and take the inference result of the candidate thinking card with the highest score as the optimal inference result.
[0010] Optionally, for the problem to be inferred, use the first large - language model to perform inference according to the inference path corresponding to each candidate thinking card, to obtain the inference results corresponding to each of the multiple candidate thinking cards, including: Input the problem to be inferred and the text description of the inference path corresponding to each candidate thinking card into the first large - language model to obtain the inference result and the inference process of each candidate thinking card; Use the second large - language model to evaluate the inference results corresponding to each of the multiple candidate thinking cards, to obtain the optimal inference result as the answer to the problem to be inferred, including: Input the inference result and the inference process of each candidate thinking card into the second large - language model based on process verification to obtain the inference process score of each candidate thinking card, and take the inference result of the candidate thinking card with the highest score as the optimal inference result.
[0011] Optionally, for the problem to be inferred, use the first large - language model to perform inference according to the inference path corresponding to each candidate thinking card, to obtain the inference results corresponding to each of the multiple candidate thinking cards, including: Input the problem to be inferred and the text description of the inference path corresponding to each candidate thinking card into the first large - language model to obtain the inference result and the inference process of each candidate thinking card; Use the second large - language model to evaluate the inference results corresponding to each of the multiple candidate thinking cards, to obtain the optimal inference result as the answer to the problem to be inferred, including: Input the reasoning result and the reasoning process of each candidate thinking card into the second large language model based on consistency verification to obtain the reasoning process score of each candidate thinking card, and use the reasoning result of the candidate thinking card with the highest score as the optimal reasoning result.
[0012] The second aspect of the present invention provides a large model reasoning method based on Monte Carlo tree search, and the method includes: Match the problem to be reasoned with each thinking card in the thinking card library, and screen out multiple candidate thinking cards; Through the first large language model, use the optimal effective reasoning path corresponding to each candidate thinking card as an explicit thinking chain guiding signal to reason about the problem to be reasoned, obtain the reasoning results corresponding to each of the multiple candidate thinking cards, and, through the second large language model, evaluate the reasoning results corresponding to each of the multiple candidate thinking cards to obtain the optimal reasoning result as the answer to the problem to be reasoned; Among them, the optimal effective reasoning path corresponding to each candidate thinking card is a reasoning path obtained based on Monte Carlo tree search, including multiple reasoning actions connected in a chain, and each reasoning action represents a thinking mode.
[0013] The third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the large model reasoning method based on Monte Carlo tree search as described in the second aspect of the present invention.
[0014] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the large model reasoning method based on Monte Carlo tree search as described in the second aspect of the present invention.
[0015] The fifth aspect of the present invention provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, it implements the steps in the large model reasoning method based on Monte Carlo tree search as described in the second aspect of the present invention.
[0016] A large model reasoning system based on Monte Carlo tree search proposed by the present invention adaptively decomposes tasks for different tasks to obtain candidate thinking cards, uses the optimal effective reasoning path corresponding to each candidate thinking card as an explicit thinking chain guiding signal, reasons about the problem to be reasoned through a large model, and through the collaborative operation between two models, realizes efficient reasoning. This system is expected to improve the performance of large models in complex reasoning tasks and overcome the limitations of current models in mathematical, logical, and common sense reasoning. Description of the Drawings
[0017] To more clearly illustrate the technical solution of the present invention, the accompanying drawings required for the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 is the structural block diagram of the large model inference system based on Monte Carlo tree search provided by the present invention; Figure 2 is the structural schematic diagram of the large model inference system based on Monte Carlo tree search provided by the present invention; Figure 3 is the schematic diagram of the processing process of the Monte Carlo tree search-based thinking card construction module in the large model inference system based on Monte Carlo tree search provided by the present invention; Figure 4 is the schematic diagram of the process flow of the large model inference method based on Monte Carlo tree search provided by the present invention. Detailed implementation manners
[0019] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0020] Current in-context learning (ICL) methods, such as chain-of-thought (CoT), have achieved certain results in improving the reasoning ability of large language models (LLMs). However, these methods still have the following limitations: 1. Sensitivity to examples: Currently, in-context learning methods highly depend on the quality of the provided examples, and the model is very sensitive to changes in examples. This sensitivity usually requires manual careful design and selection of examples, increasing the labor cost.
[0021] 2. Lack of explicit guidance signals: Imitation learning based on examples lacks clear guidance signals, resulting in low reasoning efficiency. Without clear guidance, it is difficult for the model to effectively learn complex reasoning processes.
[0022] 3. Insufficient generalization ability: When the test samples are not in the same format as the provided examples, even if the reasoning ideas are the same, it is difficult for the model to reason correctly. This limits the applicability of the model in different tasks and scenarios.
[0023] To address the problems existing in existing In-Context Learning (ICL) methods, such as sensitivity to example selection, lack of explicit guidance signals, and insufficient generalization ability, an embodiment of the present invention proposes a large model inference system based on Monte Carlo Tree Search (MCTS). By expanding the concept of "context" in traditional ICL from specific samples to a higher-order abstract cognitive thinking mode, and utilizing the tree-like hierarchical structure of MCTS, explicit thinking chain guidance signals are provided, avoiding the sensitivity of example selection and the problem of insufficient generalization ability in imitation learning. Through the collaborative operation based on the tree structure between two models, the inference performance and efficiency are significantly improved.
[0024] Specifically, as Figure 1 shown, it shows the structural block diagram of the large model inference system based on Monte Carlo Tree Search provided by an embodiment of the present invention. The system includes: A task decomposition module 101, configured to match the problem to be inferred with each thinking card in the thinking card library, and screen out multiple candidate thinking cards.
[0025] Among them, the optimal effective inference path corresponding to each candidate thinking card is an inference path obtained based on Monte Carlo Tree Search, including a plurality of chained inference actions, and each inference action represents a thinking mode.
[0026] In an embodiment of the present invention, through the task decomposition module, candidate thinking cards corresponding to the problem to be inferred can be obtained, and then the corresponding optimal effective inference path can be obtained. Based on this optimal effective inference path, a plurality of chained inference actions can be obtained, and in the subsequent inference process, these chained inference actions can be used as explicit thinking chain guidance signals to provide a guiding idea for the subsequent collaborative inference and verification module, and guide the first large language model to perform inference.
[0027] A collaborative inference and verification module 102, configured to use the first large language model to take the optimal effective inference path corresponding to each candidate thinking card as an explicit thinking chain guidance signal, perform inference on the problem to be inferred, obtain the inference results corresponding to each of the multiple candidate thinking cards, and, use the second large language model to evaluate the inference results corresponding to each of the multiple candidate thinking cards to obtain the optimal inference result as the answer to the problem to be inferred.
[0028] In the embodiments of the present invention, the collaborative reasoning and verification module uses two models with approximately 7B parameters to work together. The first large language model is used to reason about the problem to be reasoned. This first large language model is responsible for generating a complete complex reasoning path based on the explicit chain-of-thought guidance signals provided by the candidate thought cards, and obtaining the reasoning result for the problem to be reasoned.
[0029] Specifically, the explicit chain-of-thought guidance signals provide a description of multiple chained reasoning actions. The first large language model can generate specific reasoning steps based on the description of the multiple chained reasoning actions and obtain the reasoning result.
[0030] The second large language model is used to verify and screen the reasoning results and processes generated by the first large language model. To solve the problem of how to select the optimal result from the reasoning results corresponding to each of the multiple candidate thought cards, in the embodiments of the present invention, a process supervision model can be used to score each step in the reasoning process corresponding to each reasoning result, and use the minimum score in the reasoning process as the overall score of this path. Finally, the one with the highest score is selected from the reasoning results corresponding to each of the multiple candidate thought cards as the final reasoning result.
[0031] In the embodiments of the present invention, the second large language model can adopt existing verification models in related technologies, such as: result verification model, process verification model, or consistency verification model.
[0032] Therefore, the technical solution provided by the embodiments of the present invention reduces the dependence on the selection of specific examples, improves the stability and generalization ability of the reasoning system; provides explicit chain-of-thought guidance for the first large language model, making the reasoning process more organized and controllable; through the multi-path generation and process supervision of the collaborative reasoning model, effectively balances the reasoning performance and computational efficiency, and significantly improves the overall effect of complex reasoning tasks.
[0033] In an alternative embodiment, a schematic structural diagram of another large model reasoning system based on Monte Carlo tree search provided by the embodiments of the present invention is as Figure 2 shown, and the system includes: A reasoning action definition module for defining N reasoning actions.
[0034] In the embodiments of the present invention, the most basic behavior definition of the reasoning process is carried out by constructing five types of human-like cognitive reasoning actions. These actions simulate the "System 2" thinking mode, that is, deliberate, logical and efficient cognitive reasoning. With the help of these strictly defined basic actions, complex problems can be decomposed into a series of structured, chained "thought cards", providing a standardized basic unit for subsequent tree search.
[0035] Specifically, the five types of human-like cognitive reasoning actions include: System Analysis (SA), One-Step Thought (OST), Chain-of-Thought (CoT), Divide and Conquer (DC), and Self-Reflection and Refine (SRR).
[0036] The thinking card construction module based on Monte Carlo tree search is used to abstract each seed problem into a tree search problem, obtain the optimal effective reasoning path corresponding to each seed problem, and use multiple seed problems sharing an optimal effective reasoning path as a thinking card to obtain a thinking card library. Among them, an optimal effective reasoning path means that the answer obtained by reasoning according to this reasoning path matches the correct answer of the seed problem and the reasoning cost is minimized. An optimal effective reasoning path includes a root node and at least one node, each node represents a reasoning action, and the root node represents the seed problem.
[0037] In the embodiment of the present invention, based on the basic cognitive reasoning actions defined by the reasoning action definition module, complex reasoning problems can be abstracted into a tree search problem, and the seed problems q are used as the root nodes of the search tree, and each subsequent node represents a specific reasoning action.
[0038] In the embodiment of the present invention, the number of seed problems is 200, which is used as seed data for the construction of thinking cards by the thinking card construction module based on Monte Carlo tree search.
[0039] In the embodiment of the present invention, Monte Carlo tree search (including stages such as selection, expansion, simulation, and backpropagation) is used to update the dynamic reward values of each node to obtain an effective reasoning path, and the effective reasoning path is scored by calculating the Value of Completion (VOC) to obtain the optimal effective reasoning path, so as to screen and construct a problem-optimal effective reasoning path repository to provide a high-quality template for subsequent reasoning.
[0040] Specifically, in the embodiment of the present invention, abstracting each seed problem into a tree search problem and obtaining the optimal effective reasoning path corresponding to each seed problem includes the following steps: S1. During the first Monte Carlo tree search, randomly select multiple reasoning actions as the child nodes of the root node to obtain a first intermediate reasoning result. The randomly selected multiple reasoning actions are used as the first-level nodes.
[0041] S2. Based on the first intermediate reasoning result, continue to select multiple reasoning actions for reasoning until the reasoning ends to obtain a reasoning path. The multiple nodes selected for the nth time are used as the nth-level nodes.
[0042] S3. According to whether the inference result obtained along each inference path matches the correct answer of the seed question, update the reward values of each node upward along the inference path from the leaf node to obtain the reward value of each path.
[0043] S4. When the inference result obtained along an inference path matches the correct answer of the seed question, regard this inference path as a valid inference path.
[0044] S5. According to the reward values and the number of nodes of multiple valid inference paths, screen out the optimal valid inference path of this seed question.
[0045] In the embodiment of the present invention, at the beginning of Monte Carlo tree search, the reward values of all unexplored nodes are all set to 0, and the iterative improvement of the reward value is realized through dynamic update subsequently. The update formula is:
[0046] wherein, is the future reward discount factor, is the current node 's parent node, is the newly expanded new node.
[0047] Specifically, as Figure 3 shown, it shows a schematic diagram of the processing process of the thinking card construction module based on Monte Carlo tree search. In the actual Monte Carlo tree search process, for each seed question (such as "A person writes three letters to each of two friends every two weeks. How many letters does he write in a year?"), the system executes a complete MCTS process, including the following four stages. Starting from the root node, use the tree policy (such as the UCT algorithm) to select child nodes until a node that can be expanded or terminated is found: (1) Selection: Based on the Upper Confidence Bound (UCT) formula: ; wherein, represents the s of the child node , is the exploration coefficient , used to control the balance degree , represents the number of visits of the parent node p , ) represents the number of visits of the child node s , , represents the reward value of the child node s , 。
[0048] Based on the UCT formula, at each decision point, the node with the highest UCT value can be selected as the child node to balance exploration (less visited nodes) and exploitation (nodes with high reward values).
[0049] (2) Expansion: For the selected child node, new child nodes are expanded according to the cognitive reasoning actions (SA, OST, CoT, DC, and SRR), and the reward value is initialized for each new node; (3) Simulation: Starting from the newly expanded node, simulation reasoning is carried out to obtain a complete reasoning path; (4) Backpropagation: According to the reasoning result corresponding to the reasoning path, it is matched with the correct answer of the seed problem, and the reward value and visit times of each node are updated upward along the search path from the leaf node according to the matching result.
[0050] As Figure 3 shown, for the seed problem Q, multiple reasoning actions are randomly selected as the child nodes of the root node for Monte Carlo tree search to obtain the first intermediate reasoning result, and the randomly selected multiple reasoning actions are used as the first-layer nodes. Each node in the Monte Carlo tree corresponds to a reasoning action. Based on the first intermediate reasoning result, multiple reasoning actions are continuously selected for reasoning to obtain the second-layer nodes until the final reasoning result is obtained, then the path reasoning ends, and a reasoning path is obtained. The multiple nodes selected for the nth time are used as the nth-layer nodes, and finally a reasoning path composed of multiple levels of nodes is obtained. Each node on this reasoning path corresponds to a reasoning action. By adopting the above steps, multiple reasoning paths can be obtained for a seed problem, and corresponding reasoning results can be obtained for each reasoning path. The reasoning path that matches the correct answer of the seed problem with the reasoning result is used as the valid reasoning path. Figure 3 In a 1~ a 5 respectively represent a kind of cognitive reasoning action (SA, OST, CoT, DC, and SRR).
[0051] After the Monte Carlo tree search stage, a seed can obtain one or more valid reasoning paths. Further, the embodiment of the present invention also proposes a path evaluation and optimal path selection stage for the valid reasoning paths. For each seed problem, multiple valid reasoning paths can be obtained through the MCTS process. To meet the requirements of adaptive reasoning, the embodiment of the present invention introduces an effective reasoning path evaluation mechanism based on the value of computation (VOC). The value of computation (VOC) represents the difference between the expected benefit increment of further computation and the computation cost of a certain valid reasoning path. Its calculation formula is:
[0052] Among them, is the seed problem, is the seed problem q The corresponding valid inference path, is the benefit - cost balance hyperparameter. For simplicity, the present invention defines is the path for the seed problem q The final reward value (i.e., the value of the root node of this path value), while is used to evaluate the number of cognitive inference actions in this valid inference path (i.e., the total number of nodes included in the valid inference path), thereby reflecting its inference cost. Finally, according to this score select the optimal path from the set of valid inference paths to obtain the optimal valid inference path , and construct a "seed problem - valid optimal path repository" based on the seed data.
[0053] In the embodiment of the present invention, the "seed problem - valid optimal path repository" is used to record each seed problem and the corresponding valid optimal path.
[0054] In the embodiment of the present invention, multiple seed problems sharing an optimal valid inference path are regarded as a thinking card, and a thinking card library is obtained, including: S11, compare the optimal valid inference paths of multiple seed problems to determine the set of seed problems sharing the same optimal valid inference path.
[0055] S12, determine the semantic representation, the number of sub - problems, and the conditional complexity of each seed problem in a set of seed problems.
[0056] S13, use the mean of the semantic representations, the mean of the number of sub - problems, and the mean of the conditional complexities of each seed problem in a set of seed problems as the semantic representation, the number of sub - problems, and the conditional complexity of a thinking card.
[0057] In the embodiments of the present invention, the thinking card construction module based on Monte Carlo tree search further refines multiple inference paths in the "seed problem - optimal effective inference path" into high-level thinking cards, and each thinking card records the characteristics of multiple seed problems sharing an optimal effective inference path. Specifically, each card abstractly represents an inference method (i.e., the optimal effective inference path) summarized from prior problems. The feature extraction process of the seed problems is based on the following three indicators: semantic representation, the number of sub-problems, and condition complexity. Specifically, the problem text of the seed problem can be converted into a semantic vector with a fixed dimension (embedding representation). The seed problem is parsed to obtain the number of sub-problems and the condition complexity of the conditions corresponding to the seed problem. After this extraction, each seed problem can be represented by corresponding features. Further, the mean of the semantic representations, the mean of the number of sub-problems, and the mean of the condition complexity of all seed problems in the set of seed problems sharing the same optimal effective inference path can be used as the semantic representation, the number of sub-problems, and the condition complexity of the corresponding thinking card, and used as the embedding representation on the thinking card. Thus, the high-level thinking card can be used as a template and referenced based on the semantic representation, the number of sub-problems, and the condition complexity of the thinking card in the subsequent inference stage, so as to provide a unique and efficient inference template for different problems to be inferred.
[0058] A task decomposition module, configured to match the problem to be inferred with each thinking card in the thinking card library, and screen out multiple candidate thinking cards.
[0059] In the embodiments of the present invention, the task decomposition module is specifically configured to: Calculate the semantic representation, the number of sub-problems, and the condition complexity of the problem to be inferred; compare the semantic representation, the number of sub-problems, and the condition complexity of each thinking card in the thinking card library with the semantic representation, the number of sub-problems, and the condition complexity of the problem to be inferred to obtain multiple candidate thinking cards.
[0060] In the embodiments of the present invention, for the problem to be inferred, its semantic representation, condition complexity, and the number of sub-problems can be determined, and based on the semantic representation, condition complexity, and the number of sub-problems, pattern matching is performed in the pre-constructed high-level thinking card set based on semantic similarity [R1], condition complexity, and the number of sub-problems; the corresponding candidate thinking cards are obtained. Thus, a personalized and adaptable thinking template is provided for each problem to be inferred, ensuring that the inference process can be adaptively decomposed and hierarchically processed for different tasks.
[0061] Specifically, three to five most matching cards can be selected from the thinking cards based on semantic similarity, condition complexity, and the number of sub-problems. This process can be formally described as:
[0062] Among them, represents the subset of cards closest to the problem to be inferred, and the distance function is used to measure similarity. The selected cards serve as a thinking mode template to provide guiding ideas for the subsequent collaborative reasoning and verification module.
[0063] In the embodiments of the present invention, based on semantic representation, condition complexity, and the number of sub-problems, candidate thinking cards with relatively high semantic similarity, similar condition complexity, and similar number of sub-problems to the problem to be inferred can be matched from the set of thinking cards.
[0064] The collaborative reasoning and verification module is used to, through the first large language model, take the optimal effective reasoning path corresponding to each candidate thinking card as an explicit thinking chain guiding signal, perform reasoning on the problem to be inferred, obtain the reasoning results corresponding to each candidate thinking card, and, through the second large language model, evaluate the reasoning results corresponding to each candidate thinking card to obtain the optimal reasoning result as the answer to the problem to be inferred.
[0065] In the embodiments of the present invention, for the problem to be inferred, through the first large language model, reasoning is performed according to the reasoning path corresponding to each candidate thinking card, and the reasoning results corresponding to each candidate thinking card are obtained, including: inputting the problem to be inferred and the text description of the reasoning path corresponding to each candidate thinking card into the first large language model to obtain the reasoning result and reasoning process of each candidate thinking card.
[0066] In the embodiments of the present invention, after determining the reasoning path corresponding to the candidate thinking card, the reasoning path can be used as an explicit thinking chain guiding signal to guide the first large language model to perform reasoning steps on the problem to be inferred according to multiple reasoning actions corresponding to the reasoning path to obtain the reasoning result.
[0067] In the embodiments of the present invention, by evaluating the reasoning results corresponding to each candidate thinking card through the second large language model to obtain the optimal reasoning result as the answer to the problem to be inferred, it includes any one of the following: Inputting the reasoning result of each candidate thinking card into the second large language model based on result verification to obtain the reasoning result score of each candidate thinking card, and taking the reasoning result of the candidate thinking card with the highest score as the optimal reasoning result; Inputting the reasoning result and reasoning process of each candidate thinking card into the second large language model based on process verification to obtain the reasoning process score of each candidate thinking card, and taking the reasoning result of the candidate thinking card with the highest score as the optimal reasoning result; Input the reasoning results and reasoning processes of each candidate thinking card into the second large language model based on consistency verification to obtain the reasoning process scores of each candidate thinking card, and use the reasoning result of the candidate thinking card with the highest score as the optimal reasoning result.
[0068] In the embodiments of the present invention, the second large language model for verifying multiple reasoning results and reasoning processes obtained by the first large language model can be a model based on result verification, a model based on process verification, or a model based on consistency (voting) verification.
[0069] In the embodiments of the present invention, a collaborative operation based on a tree structure between two models is adopted: the first large language model is responsible for generating multiple reasoning results and multiple reasoning steps, and the second large language model performs process supervision, scoring, and verification on each step of each reasoning step; finally, the optimal path is selected based on the overall scores of each reasoning step, significantly improving the reasoning performance and efficiency of the system and ensuring the accuracy and robustness of the reasoning results.
[0070] The overall design of the system provided by the embodiments of the present invention realizes the expansion of the specific sample "context" in traditional context learning to a higher-order abstract cognitive thinking mode, and realizes explicit thinking chain guidance in a tree-shaped hierarchical structure. Through the Monte Carlo tree search mechanism, it effectively overcomes the technical defects such as example selection sensitivity, lack of explicit guidance signals, and insufficient generalization ability existing in traditional context learning methods; it realizes the diversification, adaptive generation, fine screening, and dynamic update of reasoning paths in complex reasoning tasks, and has wide applicability and high promotion prospects.
[0071] Compared with the prior art, a large model reasoning system based on Monte Carlo tree search (MCTS) proposed in the embodiments of the present invention has the following beneficial effects: 1. Reduce sensitivity to example selection: By expanding the "context" concept in traditional context learning from specific samples to a higher-order abstract cognitive thinking mode, the dependence on specific examples is reduced, and the sensitivity of the model to example selection is reduced.
[0072] 2. Provide explicit thinking chain guidance: Utilize the tree-shaped hierarchical structure of MCTS to naturally provide an explicit thinking chain structure, enhancing the reasoning ability of the model.
[0073] 3. Improve generalization ability: Through the collaborative operation of the tree structure, the model can better adapt to different tasks and scenarios, overcoming the problem of insufficient generalization ability in traditional imitation learning.
[0074] 4. Improve reasoning performance and efficiency: Through the collaborative operation based on the tree structure between two models, the reasoning performance and efficiency are significantly improved.
[0075] Based on the same inventive concept, the present invention also provides a large model inference method based on Monte Carlo tree search, as Figure 4 shown, which shows a step flow chart of the large model inference method based on Monte Carlo tree search. The method includes the following steps: S401, match the problem to be inferred with each thinking card in the thinking card library, and screen out multiple candidate thinking cards; S402, through the first large language model, use the optimal effective inference path corresponding to each candidate thinking card as an explicit thinking chain guidance signal to infer the problem to be inferred, and obtain the inference results corresponding to each of the multiple candidate thinking cards. And, evaluate the inference results corresponding to each of the multiple candidate thinking cards through the second large language model to obtain the optimal inference result as the answer to the problem to be inferred; Among them, the optimal effective inference path corresponding to each candidate thinking card is an inference path obtained based on Monte Carlo tree search, including multiple inference actions connected in a chain, and each inference action represents a thinking mode.
[0076] Optionally, the method further includes: S41, define N inference actions; S42, abstract each seed problem into a tree search problem to obtain the optimal effective inference path corresponding to each seed problem, and use multiple seed problems sharing an optimal effective inference path as a thinking card to obtain a thinking card library. Among them, an optimal effective inference path means that the answer inferred according to this inference path matches the correct answer of the seed problem and the inference cost is the smallest. An optimal effective inference path includes a root node and at least one node, and each node represents an inference action, and the root node represents the seed problem.
[0077] In the embodiment of the present invention, the thinking card library can be obtained based on the steps S41~S42 and is executed before step S401. Generally speaking, the thinking card library can be obtained in advance based on S41~S42.
[0078] Optionally, abstracting each seed problem into a tree search problem to obtain the optimal effective inference path corresponding to each seed problem includes: During the first Monte Carlo tree search process, randomly select multiple inference actions as the children nodes of the root node to obtain a first intermediate inference result, and the randomly selected multiple inference actions are used as the first layer nodes; Based on the first intermediate inference result, continue to select multiple inference actions for inference until the inference ends to obtain an inference path, and the multiple nodes selected for the nth time are used as the nth layer nodes; According to whether the inference result obtained along each inference path matches the correct answer of the seed question, update the reward values of each node upward along the inference path from the leaf node to obtain the reward value of each path; In the case where the inference result obtained along an inference path matches the correct answer of the seed question, regard this inference path as a valid inference path; According to the reward values and the number of nodes of multiple valid inference paths, screen out the optimal valid inference path for this seed question.
[0079] Optionally, regard multiple seed questions sharing an optimal valid inference path as a thinking card to obtain a thinking card library, including: Compare the optimal valid inference paths of multiple seed questions to determine the set of seed questions sharing the same optimal valid inference path; Determine the semantic representation, the number of sub-questions, and the condition complexity of each seed question in a set of seed questions; Take the mean of the semantic representations, the mean of the number of sub-questions, and the mean of the condition complexity of each seed question in a set of seed questions as the semantic representation, the number of sub-questions, and the condition complexity of a thinking card.
[0080] Optionally, match the question to be inferred with each thinking card in the thinking card library to screen out multiple candidate thinking cards, including: Calculate the semantic representation, the number of sub-questions, and the condition complexity of the question to be inferred; Compare the semantic representation, the number of sub-questions, and the condition complexity of each thinking card in the thinking card library with the semantic representation, the number of sub-questions, and the condition complexity of the question to be inferred to obtain multiple candidate thinking cards.
[0081] Optionally, for the question to be inferred, through the first large language model, perform inference according to the inference path corresponding to each candidate thinking card to obtain the inference results corresponding to each candidate thinking card, including: Input the question to be inferred and the text description of the inference path corresponding to each candidate thinking card into the first large language model to obtain the inference result of each candidate thinking card; Evaluate the inference results corresponding to each candidate thinking card through the second large language model to obtain the optimal inference result as the answer to the question to be inferred, including: Input the inference result of each candidate thinking card into the second large language model based on result verification to obtain the inference result score of each candidate thinking card, and take the inference result of the candidate thinking card with the highest score as the optimal inference result.
[0082] Optionally, for the problem to be inferred, through the first large language model, reasoning is performed according to the reasoning paths corresponding to each candidate thinking card, and the reasoning results corresponding to each of the multiple candidate thinking cards are obtained, including: Input the problem to be inferred and the text descriptions of the reasoning paths corresponding to each candidate thinking card into the first large language model to obtain the reasoning results and reasoning processes of each candidate thinking card; Evaluate the reasoning results corresponding to each of the multiple candidate thinking cards through the second large language model to obtain the optimal reasoning result as the answer to the problem to be inferred, including: Input the reasoning results and reasoning processes of each candidate thinking card into the second large language model based on process verification to obtain the reasoning process scores of each candidate thinking card, and use the reasoning result of the candidate thinking card with the highest score as the optimal reasoning result.
[0083] Optionally, for the problem to be inferred, through the first large language model, reasoning is performed according to the reasoning paths corresponding to each candidate thinking card, and the reasoning results corresponding to each of the multiple candidate thinking cards are obtained, including: Input the problem to be inferred and the text descriptions of the reasoning paths corresponding to each candidate thinking card into the first large language model to obtain the reasoning results and reasoning processes of each candidate thinking card; Evaluate the reasoning results corresponding to each of the multiple candidate thinking cards through the second large language model to obtain the optimal reasoning result as the answer to the problem to be inferred, including: Input the reasoning results and reasoning processes of each candidate thinking card into the second large language model based on consistency verification to obtain the reasoning process scores of each candidate thinking card, and use the reasoning result of the candidate thinking card with the highest score as the optimal reasoning result.
[0084] A large model reasoning method based on Monte Carlo tree search proposed by the present invention adaptively decomposes tasks for different tasks to obtain candidate thinking cards, uses the optimal effective reasoning path corresponding to each candidate thinking card as an explicit thinking chain guidance signal, performs reasoning on the problem to be inferred through a large model, and realizes efficient reasoning through the collaborative operation between two models. This method is expected to improve the performance of large models in complex reasoning tasks and overcome the limitations of current models in mathematical and logical reasoning.
[0085] Based on the same inventive concept, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the steps in the large model reasoning method based on Monte Carlo tree search as described in any one of the above embodiments.
[0086] Based on the same inventive concept, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in the large model inference method based on Monte Carlo tree search described in any of the above embodiments are implemented.
[0087] Based on the same inventive concept, the present invention provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps in the large model inference method based on Monte Carlo tree search described in any of the above embodiments are implemented.
[0088] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0089] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, devices, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0090] The present invention is described with reference to the flowcharts and / or block diagrams of methods, terminal devices (devices), and computer program products according to the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable terminal devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable terminal devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0091] These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including instruction means, and the instruction means implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0092] These computer program instructions can also be loaded onto a computer or other programmable terminal device, so that a series of operation steps are performed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide for implementing the process Figure 1 one process or multiple processes and / or blocks Figure 1 steps for the functions specified in one block or multiple blocks.
[0093] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0094] Finally, it should also be noted that in the present invention, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the said element.
[0095] The above has introduced in detail a large model inference system based on Monte Carlo tree search provided by the present invention. Specific examples are used in the present invention to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A large model inference system based on Monte Carlo tree search, characterized in that, The system includes: A task decomposition module, configured to match the problem to be inferred with each thinking card in the thinking card library, and screen out multiple candidate thinking cards; A collaborative reasoning and verification module, configured to use a first large language model to use the optimal effective reasoning path corresponding to each candidate thinking card as an explicit thinking chain guidance signal to infer the problem to be inferred, obtain the reasoning results corresponding to each of the multiple candidate thinking cards, and use a second large language model to evaluate the reasoning results corresponding to each of the multiple candidate thinking cards to obtain the optimal reasoning result as the answer to the problem to be inferred; Wherein, the optimal effective reasoning path corresponding to each candidate thinking card is a reasoning path obtained based on Monte Carlo tree search, including multiple reasoning actions connected in a chain, and each reasoning action represents a thinking mode.
2. The large model inference system based on Monte Carlo tree search according to claim 1, characterized in that, The system further includes: A reasoning action definition module, configured to define N reasoning actions; A thinking card construction module based on Monte Carlo tree search, configured to abstract each seed problem into a tree search problem, obtain the optimal effective reasoning path corresponding to each seed problem, and use multiple seed problems sharing an optimal effective reasoning path as a thinking card to obtain a thinking card library, where an optimal effective reasoning path means that the answer obtained by reasoning according to this reasoning path matches the correct answer of the seed problem and the reasoning cost is minimized. An optimal effective reasoning path includes a root node and at least one node, each node represents a reasoning action, and the root node represents the seed problem.
3. The large model inference system based on Monte Carlo tree search according to claim 2, characterized in that, Abstracting each seed problem into a tree search problem to obtain the optimal effective reasoning path corresponding to each seed problem includes: During the first Monte Carlo tree search, randomly select multiple reasoning actions as the child nodes of the root node to obtain a first intermediate reasoning result, and the randomly selected multiple reasoning actions serve as the first layer of nodes; Based on the first intermediate reasoning result, continue to select multiple reasoning actions for reasoning until the reasoning ends to obtain a reasoning path, and the multiple nodes selected for the nth time serve as the nth layer of nodes; According to whether the reasoning result obtained according to each reasoning path matches the correct answer of the seed problem, update the reward value of each node upward along the reasoning path from the leaf node to obtain the reward value of each path; When the reasoning result obtained by a reasoning path matches the correct answer of the seed problem, use this reasoning path as an effective reasoning path; According to the reward values and the number of nodes of multiple effective reasoning paths, screen out the optimal effective reasoning path of the seed problem.
4. The large model inference system based on Monte Carlo tree search according to claim 2, wherein Using multiple seed problems sharing an optimal effective reasoning path as a thinking card to obtain a thinking card library includes: Compare the optimal effective reasoning paths of multiple seed problems to determine the set of seed problems sharing the same optimal effective reasoning path; Determine the semantic representation, number of sub-problems, and condition complexity of each seed problem in a set of seed problems; Use the mean of the semantic representations, the mean of the number of sub-problems, and the mean of the condition complexity of each seed problem in a set of seed problems as the semantic representation, number of sub-problems, and condition complexity of a thinking card.
5. The large model inference system based on Monte Carlo tree search according to claim 4, wherein The task decomposition module is used for: Calculating the semantic representation, the number of sub - problems, and the condition complexity of the problem to be inferred; Comparing the semantic representation, the number of sub - problems, and the condition complexity of each thinking card in the thinking card library with those of the problem to be inferred to obtain multiple candidate thinking cards.
6. The large model inference system based on Monte Carlo tree search according to any one of claims 1 to 5, characterized in that, For the problem to be inferred, through the first large - language model, reasoning is carried out according to the reasoning path corresponding to each candidate thinking card to obtain the reasoning results corresponding to each of the multiple candidate thinking cards, including: Inputting the problem to be inferred and the text description of the reasoning path corresponding to each candidate thinking card into the first large - language model to obtain the reasoning result of each candidate thinking card; Evaluating the reasoning results corresponding to each of the multiple candidate thinking cards through the second large - language model to obtain the optimal reasoning result as the answer to the problem to be inferred, including: Inputting the reasoning result of each candidate thinking card into the second large - language model based on result verification to obtain the reasoning result score of each candidate thinking card, and taking the reasoning result of the candidate thinking card with the highest score as the optimal reasoning result.
7. The large model inference system based on Monte Carlo tree search according to any one of claims 1 to 5, characterized in that For the problem to be inferred, through the first large - language model, reasoning is carried out according to the reasoning path corresponding to each candidate thinking card to obtain the reasoning results corresponding to each of the multiple candidate thinking cards, including: Inputting the problem to be inferred and the text description of the reasoning path corresponding to each candidate thinking card into the first large - language model to obtain the reasoning result and the reasoning process of each candidate thinking card; Evaluating the reasoning results corresponding to each of the multiple candidate thinking cards through the second large - language model to obtain the optimal reasoning result as the answer to the problem to be inferred, including: Inputting the reasoning result and the reasoning process of each candidate thinking card into the second large - language model based on process verification to obtain the reasoning process score of each candidate thinking card, and taking the reasoning result of the candidate thinking card with the highest score as the optimal reasoning result.
8. The large model inference system based on Monte Carlo tree search according to any one of claims 1 to 5, characterized in that For the problem to be inferred, through the first large - language model, reasoning is carried out according to the reasoning path corresponding to each candidate thinking card to obtain the reasoning results corresponding to each of the multiple candidate thinking cards, including: Inputting the problem to be inferred and the text description of the reasoning path corresponding to each candidate thinking card into the first large - language model to obtain the reasoning result and the reasoning process of each candidate thinking card; Evaluating the reasoning results corresponding to each of the multiple candidate thinking cards through the second large - language model to obtain the optimal reasoning result as the answer to the problem to be inferred, including: Inputting the reasoning result and the reasoning process of each candidate thinking card into the second large - language model based on consistency verification to obtain the reasoning process score of each candidate thinking card, and taking the reasoning result of the candidate thinking card with the highest score as the optimal reasoning result.
9. A large model inference method based on Monte Carlo tree search, characterized in that, The method includes: Matching the problem to be inferred with each thinking card in the thinking card library to screen out multiple candidate thinking cards; Through the first large language model, the optimal effective reasoning path corresponding to each candidate thinking card is used as an explicit chain-of-thought guiding signal to reason about the problem to be reasoned, obtaining the reasoning results corresponding to each of the multiple candidate thinking cards, and, through the second large language model, evaluating the reasoning results corresponding to each of the multiple candidate thinking cards to obtain the optimal reasoning result as the answer to the problem to be reasoned; Among them, the optimal effective reasoning path corresponding to each candidate thinking card is a reasoning path obtained based on Monte Carlo tree search, including multiple reasoning actions connected in a chain, and each reasoning action represents a thinking mode.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the large model reasoning method based on Monte Carlo tree search described in claim 9.
Citation Information
Patent Citations
Multimodal reasoning method based on Monte Carlo tree and dynamic retrieval
CN119808941A
Intelligent decision-making method, computer equipment and computer readable storage medium
CN119808966A
Complex reasoning method based on retrieval enhanced verification and improvement
CN119990302A
Enhanced large language model processing method, system and platform based on Monte Carlo tree search and storage medium
CN120011510A
Interpreting Natural Language Comparisons During Visual Analysis
US20240362261A1
Cited By
Large-model multi-chain reasoning optimization method and system based on truncation-overwriting
CN120996199A
A method and system for truncated-overwritten large model multi-chain reasoning optimization
CN120996199B
Interactive canvas content generation method and device based on large model
CN121597808A