Monte carlo tree search method and device based on science question solving task

CN117933392BActive Publication Date: 2026-09-29TSINGHUA UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410065460.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2026-09-29
Estimated Expiration
2044-01-16

AI Technical Summary

Benefits of technology

[0035](1)易于实现。相较于基于PPO和RLHF的方法,SVMC不需要进行复杂的策略梯度更新,只需要训练一个stepwise的价值模型Value Model,并通过准确的价值评估及算法实现来引导模型进行多步推断。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117933392B_ABST
    Figure CN117933392B_ABST
Patent Text Reader

Abstract

The application discloses a monte carlo tree search method and device based on a science question solving task, the method comprises the following steps: obtaining a science question data set with step-by-step solving annotations; inputting the science question data set into a value model to train the value model by using a stepwise regression method, and using a monte carlo tree strategy model to search for a solution for each question to determine a corresponding search tree to train the strategy model; wherein the tree nodes of the search tree are solutions to the science questions composed of a plurality of reasoning steps, and the tree edges are the number of single-step reasoning steps; constructing a search model based on the trained value model and the trained monte carlo tree strategy model; inputting real-time science question data into the search model to perform tree search based on the node state value evaluation results output by the trained value model to obtain data search solution results. The application can greatly improve the reasoning performance of the model in the task of solving difficult university science questions, and effectively solves the problem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network information technology, and in particular to a Monte Carlo tree search method and apparatus for solving science problems. Background Technology

[0002] In various tasks, specific optimization methods are often required to maximize the inference capabilities of large models. Solving step-by-step problems using large models is an important task, with solving university-level science problems being the most challenging. University-level science problems demand higher levels of completeness and accuracy in logical reasoning, with clear logical relationships between each step and involving rich background knowledge, which places a significant burden on the inference of large models. In this area, common model optimization methods based on PPO+RLHF are costly and complex to implement, requiring an easier-to-implement and more efficient approach. One feasible solution is to logically divide the solution of science problems into multiple steps, allowing the model to reason one step at a time, and then improve the model's search pattern for single-step solutions, guiding the model's inference through a specific search-feedback mechanism. Commonly used search patterns include CoT and ToT, but these frameworks suffer from problems such as a relatively small search level and incomplete search of the solution space.

[0003] Currently, most language models require optimization for a given task using methods such as SFT, RLHF, and RLAIF, and the search method in the model output space is also a key focus. For tasks involving logically complex STEM problems, CoT and ToT are commonly used search methods. These methods often employ basic BFS and DFS algorithms, which typically have two limitations: 1. The search level is relatively small, resulting in an incomplete search of the solution space. 2. There is a lack of feedback mechanisms for each search step, and the search process lacks proper guidance. Therefore, relying solely on these basic search algorithms cannot guide the model and is insufficient for solving more difficult problems. Currently, one attempt at a search framework is to use Monte Carlo Tree Search (MCTS). This method utilizes the reward model generated during the PPO process from a pre-trained model, employs MCTS to search the problem's solution space, and uses the reward model to evaluate the state, guiding the search process. However, this requires executing a complete RL process, which is costly. Summary of the Invention

[0004] The present invention aims to at least partially solve one of the technical problems in the related art.

[0005] To address this, this invention proposes a Monte Carlo tree search method for solving science problems. It improves upon the Monte Carlo tree search method for datasets composed of science problems, employing a value model to guide the search for problem-solving steps. This method is easy to implement and has significant optimization effects, greatly improving the inference performance of the model in solving challenging university science problems.

[0006] Another objective of this invention is to provide a Monte Carlo tree search device based on a science problem-solving task.

[0007] To achieve the above objectives, this invention proposes a Monte Carlo tree search method based on a science problem-solving task, comprising:

[0008] Obtain a dataset of science problems with step-by-step solution annotations;

[0009] The dataset of science questions is input into the value model to train the value model using stepwise regression. The Monte Carlo tree strategy model is used to search for solutions to each question and determine the corresponding search tree for training the strategy model. The tree nodes of the search tree are solutions to science questions consisting of several reasoning steps, and the tree edges are the number of single-step reasoning steps performed.

[0010] A search model is constructed based on a pre-trained value model and a pre-trained Monte Carlo tree strategy model;

[0011] Real-time science problem data is input into the search model, and tree search is performed based on the node state value evaluation results output by the trained value model to obtain the data search solution results.

[0012] The Monte Carlo tree search method for solving science problems in this embodiment of the invention may also have the following additional technical features:

[0013] In one embodiment of the present invention, obtaining a dataset of science problems with step-by-step solution annotations includes:

[0014] Obtain a basic dataset based on a science problem-solving task, and construct a first data point based on the relevant solution results of the basic dataset; wherein, the relevant solution results include the step-by-step solution of each science problem in the basic dataset and annotations reflecting whether the solution is correct;

[0015] The corresponding solution in the first data point is divided into several steps to form the second data point;

[0016] The first k data steps of the second data point are labeled with value using linear or curvilinear assignment methods to obtain the science problem dataset.

[0017] In one embodiment of the present invention, the value model is a ChatGLM3-6B pre-trained model.

[0018] In one embodiment of the present invention, each search round corresponding to the search tree includes a node selection phase, an expansion phase, an MC simulation search phase, and a backpropagation phase.

[0019] In one embodiment of the present invention, the node selection stage: let T be the search tree of a certain stage, and let S be a certain state node of the t-th layer in T. t C S Let v be the set of child nodes of node S. S V represents the number of visits to node S. S Let S be the value of node S, and ε be the search rate. The node selection phase starts from the root node and selects child nodes until an unexpanded node is reached or a node meets the termination condition. For the current node S... t For all child nodes s, select the child nodes that satisfy the following formula. If the termination condition is met, end the entire search process and output the results:

[0020]

[0021] The expansion phase: Let S be the node selected in the node selection phase. t Then, using the Monte Carlo tree strategy model P, the inference is repeated n_search times, and the resulting single-step is denoted as a1, a2, ..., a n_search This generates n_search new states s1,...,s n_search As S t The child nodes of T are included, and the value of all child nodes is evaluated:

[0022]

[0023] In the MC simulation search phase: the highest-value child node of the node expanded during the expansion phase is denoted as S. T The forward exploration steps are n_forward, the initial maximum value M = 0, starting from S T The forward simulation search is repeated n_forward times. In each search, the current state is denoted as S. t Let P generate several next actions, and use a greedy or random strategy to select and execute an action to obtain a new state S. t+1 If the value of the new state is higher than the maximum value M, update M;

[0024] The backpropagation phase: Let γ be the simulated weight parameter, then for S T Value update:

[0025]

[0026] From S T Begin updating the reverse value and access count, assuming the current node is S. t If the current node is not found, the access count is incremented by one, the value of the parent node is updated in a weighted manner, and then the current node becomes the parent node. This process continues until the root node is reached.

[0027]

[0028] To achieve the above objectives, another aspect of the present invention proposes a Monte Carlo tree search device based on a science problem-solving task, comprising:

[0029] The dataset acquisition module is used to acquire datasets of science problems with step-by-step solution annotations;

[0030] The multi-model training module is used to input the science problem dataset into the value model to train the value model using the stepwise regression method, and to use the Monte Carlo tree strategy model to search for solutions to each problem and determine the corresponding search tree for strategy model training; wherein, the tree nodes of the search tree are solutions to science problems consisting of several reasoning steps, and the tree edges are the number of single-step reasoning steps performed.

[0031] The search model building module is used to build a search model based on a trained value model and a trained Monte Carlo tree strategy model.

[0032] The search solution output module is used to input real-time science question data into the search model and perform tree search based on the node state value evaluation results output by the trained value model to obtain the data search solution results.

[0033] The Monte Carlo tree search method and apparatus based on the problem-solving task in the present invention explores the solution space and optimizes the search direction of each solution step by combining a value model. It can achieve a certain optimization effect without performing the RL process or making any fine-tuning to the strategy model, so that the optimized model can predict the solution more accurately and effectively.

[0034] The beneficial effects of this invention are as follows:

[0035] (1) Easy to implement. Compared with PPO and RLHF-based methods, SVMC does not require complex policy gradient updates. It only needs to train a stepwise value model and guide the model to perform multi-step inference through accurate value evaluation and algorithm implementation.

[0036] (2) Significant optimization effect. Without performing task-specific fine-tuning on the pre-trained model, SVMC alone can improve the accuracy of ChatGLM2 by about 10% on a given university math dataset, which is a very significant optimization effect.

[0037] (3) Wide applicability. The SVMC optimization method has the ability to be easily transferred to a wider range of tasks, such as problem solving in other subjects, situational reasoning, and strategy development. SVMC does not require setting a fixed number of reasoning steps, but only determines whether the reasoning ends based on value judgments, thus making it suitable for a wider range of tasks.

[0038] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0039] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0040] Figure 1 This is a flowchart of a Monte Carlo tree search method based on a science problem-solving task according to an embodiment of the present invention;

[0041] Figure 2 This is a framework diagram of a Monte Carlo tree search method based on a science problem-solving task according to an embodiment of the present invention;

[0042] Figure 3 This is a schematic diagram of the SVMC algorithm according to an embodiment of the present invention;

[0043] Figure 4 This is a schematic diagram of a Monte Carlo tree search device based on a science problem-solving task according to an embodiment of the present invention. Detailed Implementation

[0044] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0045] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0046] The Monte Carlo tree search method and apparatus based on a science problem-solving task, according to embodiments of the present invention, are described below with reference to the accompanying drawings.

[0047] Figure 1 This is a flowchart of a Monte Carlo tree search method based on a science problem-solving task according to an embodiment of the present invention.

[0048] like Figure 1 As shown, the method of the present invention includes, but is not limited to, the following steps:

[0049] S1, Obtain a dataset of science problems with step-by-step solution annotations;

[0050] S2, input the dataset of science questions into the value model to train the value model using the stepwise regression method, and use the Monte Carlo tree strategy model to search for solutions to each question to determine the corresponding search tree for training the strategy model; where the tree nodes of the search tree are solutions to science questions consisting of several reasoning steps, and the tree edges are the number of single-step reasoning steps performed.

[0051] S3, a search model is constructed based on a trained value model and a trained Monte Carlo tree strategy model;

[0052] S4 inputs real-time science problem data into the search model and performs tree search based on the node state value evaluation results output by the trained value model to obtain the data search solution results.

[0053] Figure 2 This is a flowchart illustrating the method of an embodiment of the present invention. Figure 2 As shown, the SVMC framework implementation consists of three main modules: acquiring a refined dataset of problems, implementing the MCTS algorithm, and training the value model. The first step is refining the dataset, which involves step-wise value allocation of a dataset of university science problems labeled with step-by-step solutions. The result of this step serves as the training and testing data source for the Value Model. Subsequently, the Value Model is trained and tested on the refined dataset, while the MCTS algorithm is implemented. Finally, the Value Model and Policy Model are embedded into the SVMC inference framework, enabling step-by-step inference search for solutions to specific science problems.

[0054] In one embodiment of the present invention, for a given science problem-solving task, a basic dataset is required. This dataset contains a step-by-step solution to each university science problem, along with annotations indicating whether the solution is correct, forming a data point. Based on this, the corresponding solution in each data point can be divided into several smaller steps to form new data points. Then, a linear or curvilinear assignment method can be used to assign value annotations to the first k data steps of each new data point, thereby obtaining a refined science problem dataset.

[0055] Specifically, university science problems often have clear steps and logical relationships, making it possible to construct refined and high-quality step-by-step datasets. MCTS tree search is performed at the step level (node ​​level) of the problem. Each tree node represents a solution to a problem that may not be complete, consisting of several reasoning steps. The value model needs to evaluate the state value of the step-level tree nodes. To train the value model, a step-level dataset is required. For science problem datasets with insufficient refinement (labeling whether the entire solution is correct but not evaluating each step), the state value can be approximated using a certain pattern based on the actual final reward Rn of each data point (often depending on the data label). For simplicity, a linear approximation can be used. Suppose that the solution A for data point X can be decomposed into n steps a1, a2, ... a n The initial state is s0 (i.e., the state without any solution reasoning), and executions a1 to a... t The obtained state is s t (The state represented by the incomplete solution composed of the first t steps of reasoning), then s t The state value can be approximated as:

[0056]

[0057] This allows us to obtain a refined dataset of science problems for the value model to learn.

[0058] For example,

[0059] Q: Find the value of (2+3)*4-5;

[0060] A: 1. First calculate 2 + 3 = 5;

[0061] 2. Then calculate 5 * 4 = 20;

[0062] 3. Finally, the result is 20-5=15;

[0063] Tag: 1 (indicates correct)

[0064] The value is assigned using a linear method, which yields four sub-state samples.

[0065] Sample1: {Q: Find the value of (2+3)*4-5, A: , label: 0};

[0066] Sample2: {Q: Find the value of (2+3)*4-5, A: First calculate 2+3=5, label: 1 / 3};

[0067] Sample3: {Q: Find the value of (2+3)*4-5, A: First calculate 2+3=5, then calculate 5*4=20, label: 2 / 3};

[0068] Sample4: {Q: Find the value of (2+3)*4-5, A: First calculate 2+3=5, then calculate 5*4=20, finally, the result is 20-5=15, label: 1}.

[0069] In one embodiment of the present invention, a value model is trained by dividing an existing refined question dataset into training, validation and test sets, so that it outputs the corresponding value for the entire solution state after completing several steps of reasoning.

[0070] Understandably, the accuracy of the value model evaluation will significantly impact the search process results. Based on a refined dataset, stepwise training was performed using a pre-trained model of appropriate size as the basis (the ChatGLM3-6B pre-trained model was used in the experiment).

[0071] For example, Input = 'Q: Calculate the value of (2+3)*4-5, A: First calculate 2+3=5'; Output = 0.3.

[0072] In one embodiment of the present invention, the process of searching for a solution to each problem can be viewed as a search tree. A tree node represents a solution to a problem that may not be complete, consisting of several reasoning steps. Tree edges represent single-step reasoning steps. Deeper nodes have more reasoning steps, and child nodes are always states reached by performing additional reasoning steps based on the existing reasoning of their parent nodes. The search process is conducted in rounds, each round consisting of four steps: node selection phase, expansion phase, McLeod simulation search phase, and backpropagation phase.

[0073] For example, selecting a node phase:

[0074] For convenience, we will use a uniform notation below, denoting the search tree at a certain stage as T, and the state node at the t-th level of T as S. t C S Let v be the set of child nodes of node S. S V represents the number of visits to node S. SLet ε be the value of node S, and ε be the search rate. The selection phase starts from the root node and continuously selects child nodes according to the following rules until an unexpanded node is reached or a node meets the termination condition (value greater than a set threshold). For the current node S... t For all child nodes s, select the child node that satisfies the following formula. If the termination condition is met, the entire search process can be terminated and the result can be output.

[0075]

[0076] For example, the expansion phase:

[0077] Let S be the node selected during the node selection phase. t Then, the policy model P is used to repeatedly perform n_search inferences, and the resulting single-step is denoted as a1, a2, ..., a n_search This generates n_search new states s1,...,s. n_search , take it as S t The child nodes are included in T. The value of all child nodes is evaluated.

[0078]

[0079] For example, the MC simulation search phase:

[0080] Let S be the highest-value child node of the node that is expanded during the expansion phase. T The forward exploration steps are n_forward. Initialize the maximum value M = 0, starting from S. T Begin repeating the forward simulation search n_forward times. In each search, denote the current state as S. t Let P generate several next actions, and use a greedy or random strategy to select and execute an action to obtain a new state S. t+1 If the value of the new state is higher than the maximum value M, then update M.

[0081] For example, the backpropagation phase:

[0082] Let γ be the simulated weight parameter, then for S T Value update:

[0083]

[0084] From S T Begin updating the reverse value and access count. Let the current node be S. t If the current node is found to be a node, its access count is incremented by one, the value of its parent node is updated in a weighted manner, and then the current node becomes the parent node. This process continues until the root node is reached.

[0085]

[0086] The core of this invention lies in simplifying the evaluation of the action value function during the search process, while using a state value model for step-wise evaluation. Simultaneously, a greedy strategy is employed to simulate node exploration, resulting in a more accurate estimation of node value and correcting biases in the value model output. The algorithm does not require specific levels or numbers of nodes in the search tree, making it adaptable to a wide range of task needs. Figure 3 As shown.

[0087] In one embodiment of the present invention, a search framework is constructed. Based on the implemented SVMC search algorithm, the present invention embeds a strategy model P and a value model V, enabling inference for a wider range of science problems.

[0088] Furthermore, the experimental results of this invention are as follows:

[0089] Using ChatGLM3-6B as the strategy and value model, this invention achieved a 2%-10% accuracy improvement on a dataset consisting of 60 labeled university math problems. The evaluation criterion was zero-sample testing, ultimately achieving performance comparable to GPT-4, which has undergone a ToT test. (See Table 1.)

[0090] Table 1

[0091] ChatGLM2 1.67% ToT Prompt ChatGLM2 3.33% SVMC ChatGLM2 13.33% GPT-3.5-turbo 6.67% ToT Prompt GPT-3.5-turbo 10% GPT-4 10% ToT Prompt GPT-4 13.33%

[0092] The Monte Carlo tree search method based on the problem-solving task in this invention improves upon the Monte Carlo tree search method for datasets composed of university science problems. It uses a value model to guide the search of problem-solving steps and has the advantages of being easy to implement and having significant optimization effects. It can greatly improve the inference performance of the model in the task of solving difficult university science problems.

[0093] To achieve the above embodiments, such as Figure 4 As shown, this embodiment also provides a Monte Carlo tree search device 10 based on a science problem-solving task. The device 10 includes a dataset acquisition module 100, a multi-model training module 200, a search model construction module 300, and a search solution output module 400.

[0094] The dataset acquisition module 100 is used to acquire a dataset of science problems with step-by-step solution annotations;

[0095] The multi-model training module 200 is used to input the science problem dataset into the value model to train the value model using the stepwise regression method, and to use the Monte Carlo tree strategy model to search for solutions to each problem to determine the corresponding search tree for strategy model training; wherein, the tree nodes of the search tree are solutions to science problems consisting of several reasoning steps, and the tree edges are the number of single-step reasoning steps performed.

[0096] Search model building module 300 is used to build a search model based on a trained value model and a trained Monte Carlo tree strategy model;

[0097] The search solution output module 400 is used to input real-time science problem data into the search model and perform tree search based on the node state value evaluation results output by the trained value model to obtain the data search solution results.

[0098] Furthermore, the aforementioned dataset acquisition module 100 is also used for:

[0099] Obtain a basic dataset based on a science problem-solving task, and construct a first data point based on the relevant solution results of the basic dataset; wherein, the relevant solution results include the step-by-step solution of each science problem in the basic dataset and annotations reflecting whether the solution is correct;

[0100] The corresponding solution in the first data point is divided into several steps to form the second data point;

[0101] The first k data steps of the second data point are labeled with value using linear or curvilinear assignment methods to obtain the science problem dataset.

[0102] Furthermore, the value model is a pre-trained ChatGLM3-6B model.

[0103] Furthermore, each search round corresponding to the search tree includes the node selection phase, the expansion phase, the MC simulation search phase, and the backpropagation phase.

[0104] Furthermore, the aforementioned multi-model training module 200 is also used for:

[0105] The node selection phase: Let T be the search tree of a certain phase, and let S be a state node of the t-th layer in T. t C S Let v be the set of child nodes of node S. S V represents the number of visits to node S. S Let S be the value of node S, and ε be the search rate. The node selection phase starts from the root node and selects child nodes until an unexpanded node is reached or a node meets the termination condition. For the current node S... tFor all child nodes s, select the child nodes that satisfy the following formula. If the termination condition is met, end the entire search process and output the results:

[0106]

[0107] The expansion phase: Let S be the node selected in the node selection phase. t Then, using the Monte Carlo tree strategy model P, the inference is repeated n_search times, and the resulting single-step is denoted as a1, a2, ..., a n_search This generates n_search new states s1,...,s n_search As S t The child nodes of T are included, and the value of all child nodes is evaluated:

[0108]

[0109] In the MC simulation search phase: the highest-value child node of the node expanded during the expansion phase is denoted as S. T The forward exploration steps are n_forward, the initial maximum value M = 0, starting from S T The forward simulation search is repeated n_forward times. In each search, the current state is denoted as S. t Let P generate several next actions, and use a greedy or random strategy to select and execute an action to obtain a new state S. t+1 If the value of the new state is higher than the maximum value M, update M;

[0110] The backpropagation phase: Let γ be the simulated weight parameter, then for S T Value update:

[0111]

[0112] From S T Begin updating the reverse value and access count, assuming the current node is S. t If the current node is not found, the access count is incremented by one, the value of the parent node is updated in a weighted manner, and then the current node becomes the parent node. This process continues until the root node is reached.

[0113]

[0114] The Monte Carlo tree search device based on the problem-solving task of science problems in this invention improves the Monte Carlo tree search method for datasets composed of university science problems. It adopts a value model to guide the search of problem-solving steps and has the advantages of being easy to implement and having significant optimization effects. It can greatly improve the inference performance of the model in the task of solving difficult university science problems.

[0115] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0116] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

Claims

1. A Monte Carlo tree search method based on a science problem-solving task, characterized in that, The method includes the following steps: Obtain a dataset of science problems with step-by-step solution annotations; The dataset of science problems is input into the value model to train the value model using stepwise regression. A Monte Carlo tree strategy model is then used to search for solutions to each problem and determine the corresponding search tree for training the strategy model. The tree nodes of the search tree represent solutions to science problems consisting of several reasoning steps, and the tree edges represent the number of single-step reasoning steps. Each search round corresponding to the search tree includes a node selection phase, an expansion phase, a Monte Carlo simulation search phase, and a backpropagation phase. During the node selection phase: starting from the root node, for the current node... All child nodes Select child nodes that satisfy the following formula until an unexpanded node is reached or a node satisfies the termination condition: in, This represents the number of visits to node S. The value of node S, Search rate; During the expansion phase: Let the node selected in the node selection phase be denoted as . Then, the Monte Carlo tree strategy model P is used to repeatedly perform inference n_search times, generating n_search new states. As The child nodes are incorporated into the search tree T, and the value of all child nodes is evaluated using a value model: In the MC simulation search phase: Let the highest-value child node of the node expanded during the expansion phase be denoted as... Initialize the maximum value ,from The forward simulation search is repeated n_forward times. In each search, the Monte Carlo tree policy model P generates several next actions. A greedy or random policy is used to select and execute an action to obtain a new state. If the value of the new state is higher than the maximum value M, then update M; During the backpropagation phase: based on the simulated weight parameters ,right Value update: and from Let the current node be... Increment its access count by one for the current node. The value of the parent node is updated in a weighted manner, and then the current node becomes the parent node, continuing until the root node is reached: ; A search model is constructed based on a pre-trained value model and a pre-trained Monte Carlo tree strategy model; Real-time science problem data is input into the search model, and tree search is performed based on the node state value evaluation results output by the trained value model to obtain the data search solution results.

2. The method according to claim 1, characterized in that, The process of obtaining a dataset of science problems with step-by-step solution annotations includes: Obtain a basic dataset based on a science problem-solving task, and construct a first data point based on the relevant solution results of the basic dataset; wherein, the relevant solution results include the step-by-step solution of each science problem in the basic dataset and annotations reflecting whether the solution is correct; The corresponding solution in the first data point is divided into several steps to form the second data point; The first k data steps of the second data point are labeled with value using linear or curvilinear assignment methods to obtain the science problem dataset.

3. The method according to claim 2, characterized in that, The value model is a ChatGLM3-6B pre-trained model.

4. A Monte Carlo tree search device based on a science problem-solving task, characterized in that, include: The dataset acquisition module is used to acquire datasets of science problems with step-by-step solution annotations; The multi-model training module is used to input the science problem dataset into the value model for stepwise regression training, and to use the Monte Carlo tree strategy model to search for solutions to each problem and determine the corresponding search tree for strategy model training. The tree nodes of the search tree represent solutions to science problems consisting of several reasoning steps, and the tree edges represent the number of single-step reasoning steps. Each search round corresponding to the search tree includes a node selection phase, an expansion phase, a Monte Carlo simulation search phase, and a backpropagation phase. During the node selection phase: starting from the root node, for the current node... All child nodes Select child nodes that satisfy the following formula until an unexpanded node is reached or a node satisfies the termination condition: in, This represents the number of visits to node S. The value of node S, Search rate; During the expansion phase: Let the node selected in the node selection phase be denoted as . Then, the Monte Carlo tree strategy model P is used to repeatedly perform inference n_search times, generating n_search new states. As The child nodes are incorporated into the search tree T, and the value of all child nodes is evaluated using a value model: In the MC simulation search phase: Let the highest-value child node of the node expanded during the expansion phase be denoted as... Initialize the maximum value ,from The forward simulation search is repeated n_forward times. In each search, the Monte Carlo tree policy model P generates several next actions. A greedy or random policy is used to select and execute an action to obtain a new state. If the value of the new state is higher than the maximum value M, then update M; During the backpropagation phase: based on the simulated weight parameters ,right Value update: and from Let the current node be... Increment its access count by one for the current node. The value of the parent node is updated in a weighted manner, and then the current node becomes the parent node, continuing until the root node is reached: ; The search model building module is used to build a search model based on a trained value model and a trained Monte Carlo tree strategy model. The search solution output module is used to input real-time science question data into the search model and perform tree search based on the node state value evaluation results output by the trained value model to obtain the data search solution results.

5. The apparatus according to claim 4, characterized in that, The dataset acquisition module is also used for: Obtain a basic dataset based on a science problem-solving task, and construct a first data point based on the relevant solution results of the basic dataset; wherein, the relevant solution results include the step-by-step solution of each science problem in the basic dataset and annotations reflecting whether the solution is correct; The corresponding solution in the first data point is divided into several steps to form the second data point; The first k data steps of the second data point are labeled with value using linear or curvilinear assignment methods to obtain the science problem dataset.

6. The apparatus according to claim 4, characterized in that, The value model is a ChatGLM3-6B pre-trained model.

Citation Information

Patent Citations

  • Chess game method based on opponent modeling and Monte Carlo reinforcement learning

    CN116128060A