Training method and device based on online tree search, equipment and medium

Through entropy-guided online tree search and Monte Carlo method combined with the global local advantage reward mechanism, the problem of low reinforcement learning efficiency of large language models is solved, and the performance and efficiency of complex inference tasks are improved.

CN120338059APending Publication Date: 2025-07-18TSINGHUA UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510414845.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing reinforcement learning methods for large language models are inefficient and poor in complex inference tasks. Traditional Monte Carlo Tree Search (MCTS) performs poorly and inefficiently at the same inference cost, making it difficult to fully utilize the model's information in the inference process.

Method used

The online tree search method based on entropy guidance is adopted, and the bifurcated point is selected through the entropy value for tree expansion. The node value is calculated in combination with the Monte Carlo method and the global and local advantage reward mechanism is introduced, and the model is optimized by using the strategy gradient method.

Benefits of technology

It significantly improves search efficiency and response diversity, forms a closed-loop optimization system, and improves the ability of large language models in complex inference tasks such as mathematics and programming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338059A_ABST
    Figure CN120338059A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network information, in particular to a training method and device based on online tree search, equipment and a medium, and the method comprises the steps: carrying out the initialization processing of given prompt information based on entropy guide tree search, and generating a guide tree root; selecting a bifurcation point of a guide tree root according to the entropy value, and performing extension processing on the bifurcation point of the guide tree to obtain a tree structure; and calculating a node value in the tree structure by using a Monte Carlo method, calculating a reward signal based on the node value in the tree structure, and strengthening the tree search strategy model. The exploration diversity is enhanced through tree search, the learning efficiency is improved through process supervision, a closed-loop optimization system is formed, the capacity of a large language model on complex reasoning tasks such as mathematics and programming is remarkably improved, and the method has wide application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network information technology, and in particular, to a training method, device, equipment, and medium based on online tree search. Background Art

[0002] In recent years, large language models have demonstrated remarkable capabilities in various complex reasoning tasks such as mathematics, programming, and autonomous agents. With the continuous development of model scale and architecture, how to further improve the reasoning ability of these models has become the focus of research. Reinforcement learning has been proven to be an effective method for significantly improving the reasoning ability of large language models, which continuously optimizes the model's decision-making strategy through reward feedback. Current large language model reinforcement learning methods usually adopt the method of independently sampling multiple trajectories and obtaining reward signals based on the correctness of the final answer. Although this method is simple and direct, it often fails to fully utilize the rich information generated by the model during the reasoning process and is prone to falling into local optimal solutions. In contrast, tree search, as an RL technology that has achieved remarkable success in other fields such as AlphaZero, performs excellently in traditional reasoning tasks such as reinforcement chess, but has not been fully developed and utilized in the field of reinforcement learning for large language model reasoning. Some existing attempts mainly focus on using tree search to improve performance during reasoning or generate data for offline training. For example, some studies have tried to use tree search combined with external rewards to enhance performance during reasoning, or to generate offline training data. However, the potential of directly combining online policy tree search with reinforcement learning to improve large language model reasoning ability has not been fully explored. Implementing online policy RL training combined with tree search to improve LLM reasoning faces two main challenges: First, traditional Monte Carlo tree search may not be as effective as independently sampling multiple responses at the same reasoning cost, which may hinder the performance of reinforcement learning. Experiments show that at the same reasoning cost, the PassRate of MCTS is relatively poor, making it less effective in searching for the correct answer. Second, MCTS is significantly inefficient, requiring a large number of iterations and step-by-step generation, resulting in low parallelism and not being suitable for the characteristics and advantages of modern LLM reasoning engines.

[0003] In summary, how to design an efficient and accurate online tree search training method is an urgent problem to be solved at present. Summary of the Invention

[0004] This application aims to solve at least one of the technical problems in the related art to some extent.

[0005] For this reason, the first object of this application is to propose a training method based on online tree search to solve the problems of large limitations, low efficiency, and poor performance of existing technical means.

[0006] The second object of this application is to propose a device.

[0007] The third object of the present application is to propose an electronic device.

[0008] The fourth object of the present application is to propose a computer-readable storage medium.

[0009] To achieve the above object, an embodiment of the first aspect of the present application proposes a training method based on online tree search, including:

[0010] Performing an initialization process on given prompt information based on entropy-guided tree search to generate a guiding tree root;

[0011] Selecting a bifurcation point of the guiding tree root according to the entropy value, and performing an expansion process on the bifurcation point of the guiding tree to obtain a tree structure;

[0012] Calculating the node values in the tree structure by using the Monte Carlo method, calculating a reward signal based on the node values in the tree structure, and strengthening the tree search policy model.

[0013] Preferably, the performing an initialization process on given prompt information based on entropy-guided tree search to generate a guiding tree root includes:

[0014] For given prompt information, the entropy-guided tree generates multiple initial complete responses as the initialization of multiple independent trees, and the multiple initial complete responses form a tree set to obtain the guiding tree root.

[0015] Preferably, the selecting a bifurcation point of the guiding tree root according to the entropy value, and performing an expansion process on the bifurcation point of the guiding tree to obtain a tree structure includes:

[0016] Using the entropy-guided tree search to identify the token with the highest entropy value in the existing tree as the bifurcation point, calculating the entropy value of the token, and expanding the tree structure based on the entropy value.

[0017] Preferably, the expanding the tree structure based on the entropy value includes:

[0018] Continuing to generate multiple different candidate responses based on the bifurcation point, adding the multiple different candidate responses to the original tree structure, and repeating the bifurcation expansion operation to obtain the tree structure.

[0019] Preferably, the calculating the node values in the tree structure by using the Monte Carlo method, calculating a reward signal based on the node values in the tree structure, and strengthening the tree search policy model includes:

[0020] Calculate the value of each step in the tree structure based on the Monte Carlo method, and optimize the calculation of the value of each step in the tree structure by adding global advantage and local advantage, where the global advantage is calculated by comparing the difference between the node value and the root node value, and the local advantage is calculated by comparing the difference between the node value and its parent node value.

[0021] Preferably, the method of using the Monte Carlo method to calculate the node value in the tree structure, calculating the reward signal based on the node value in the tree structure, and strengthening the tree search policy model further includes:

[0022] Introduce a reward reweighting mechanism based on the Monte Carlo method. For the rewards of non-leaf node steps, reweight them by dividing by the square root of the number of leaf nodes in its subtree to strengthen the tree search policy model.

[0023] Preferably, the method of using the Monte Carlo method to calculate the node value in the tree structure, calculating the reward signal based on the node value in the tree structure, and strengthening the tree search policy model further includes:

[0024] Use the policy gradient method for model optimization processing, and its calculation formula is:

[0025] E[r(x,y)-β*log(πθ(y|x) / πref(y|x))]

[0026] Where r(·) is the reward function based on the process supervision signal, πref is the reference model, and β is the KL efficiency parameter.

[0027] To achieve the above object, the second aspect embodiment of the present application proposes a training device based on online tree search, including:

[0028] An initialization module, which performs initialization processing on the given prompt information based on entropy-guided tree search to generate a guiding tree root;

[0029] A tree structure generation module, which selects the bifurcation point of the guiding tree root according to the entropy value, and performs expansion processing on the bifurcation point of the guiding tree to obtain a tree structure;

[0030] A reinforcement learning module, which uses the Monte Carlo method to calculate the node value in the tree structure, calculates the reward signal based on the node value in the tree structure, and strengthens the tree search policy model.

[0031] To achieve the above object, the third aspect embodiment of the present application proposes an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0032] The memory stores computer execution instructions;

[0033] The processor executes the computer-executable instructions stored in the memory to implement the method described in any of the above.

[0034] To achieve the above object, an embodiment of the fourth aspect of the present application provides a computer-readable storage medium, including computer-executable instructions stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in any of the above.

[0035] A training method based on online tree search provided by the present application introduces an entropy-guided tree search strategy. By calculating and selecting the token with the highest entropy value as the bifurcation point for tree expansion, it generates more diverse responses under the same inference cost. Different from the step-by-step generation method relied on by traditional MCTS, it significantly improves the search efficiency and response diversity. Innovatively combines global advantages and local advantages to calculate a comprehensive reward for each inference step in the tree, and at the same time adopts a reward reweighting mechanism based on node importance to prevent overfitting, providing high-quality fine-grained supervision signals without an additional reward model. Enhances exploration diversity through tree search, improves learning efficiency using process supervision, forms a closed-loop optimization system, significantly enhances the capabilities of large language models in complex inference tasks such as mathematics and programming, and has broad application value.

[0036] Some of the additional aspects and advantages of the present application will be given in the following description, some will become obvious from the following description, or will be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0038] Figure 1 is a flowchart of the first specific embodiment of a training method based on online tree search provided by the present invention;

[0039] Figure 2 is a flowchart of the second specific embodiment of a training method based on online tree search provided by the present invention;

[0040] Figure 3 is an evaluation data graph for mathematics and code tasks;

[0041] Figure 4 is a first linear result data graph;

[0042] Figure 5 is a second linear result data graph;

[0043] Figure 6 is a structural block diagram of a training device based on online tree search provided by an embodiment of the present invention. Detailed implementation manners

[0044] The core of the present invention is to provide a training method, device, electronic device and medium based on online tree search, which enhances exploration diversity through tree search, improves learning efficiency by using process supervision, forms a closed-loop optimization system, and significantly improves the capabilities of large language models in complex reasoning tasks such as mathematics and programming.

[0045] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the drawings and specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] Please refer to Figure 1 , Figure 1 which is a flowchart of the first specific embodiment of a training method based on online tree search provided by the present invention; the specific operation steps are as follows:

[0047] Step S101: Initialize the given prompt information based on entropy-guided tree search to generate a guiding tree root;

[0048] For the given prompt information, the entropy-guided tree generates multiple initial complete responses as the initialization of multiple independent trees, and the multiple initial complete responses form a tree set to obtain the guiding tree root.

[0049] Step S102: Select a bifurcation point of the guiding tree root according to the entropy value, and perform an expansion process on the bifurcation point of the guiding tree to obtain a tree structure;

[0050] Use entropy-guided tree search to identify the token with the highest entropy value in the existing tree as the bifurcation point, calculate the entropy value of the token, and expand the tree structure based on the entropy value.

[0051] Based on the bifurcation point, continue to generate multiple different candidate responses, add the multiple different candidate responses to the original tree structure and repeat the bifurcation expansion operation to obtain the tree structure.

[0052] Step S103: Calculate the node values in the tree structure using the Monte Carlo method, calculate the reward signal based on the node values in the tree structure, and strengthen the tree search policy model.

[0053] Calculate the value of each step in the tree structure based on the Monte Carlo method, and optimize the calculation of the value of each step in the tree structure by adding global advantage and local advantage, where the global advantage is calculated by comparing the difference between the node value and the root node value, and the local advantage is calculated by comparing the difference between the node value and its parent node value.

[0054] Introduce a reward reweighting mechanism based on the Monte Carlo method. For the rewards of non-leaf node steps, reweight them by dividing by the square root of the number of leaf nodes in its subtree to strengthen the tree search strategy model.

[0055] Use the policy gradient method for model optimization processing, and its calculation formula is:

[0056] E[r(x,y)-β*log(πθ(y|x) / πref(y|x))]

[0057] Among them, r(·) is the reward function based on the process supervision signal, πref is the reference model, and β is the KL efficiency parameter.

[0058] This embodiment provides a training method based on online tree search. Based on the entropy-guided tree search strategy, calculate and select the token with the highest entropy value as the bifurcation point for tree expansion, which significantly improves the search efficiency and response diversity. It is a complete training framework that combines online policy tree search and reinforcement learning. Enhance exploration diversity through tree search, and improve learning efficiency using process supervision to form a closed-loop optimization system, significantly enhancing the capabilities of large language models in complex reasoning tasks such as mathematics and programming, and having broad application value.

[0059] Based on the above embodiment, this embodiment describes the training method based on online tree search, as Figure 2 shown, specifically as follows:

[0060] Input a natural language inference question q;

[0061] Step 1: Entropy-guided tree search (EPTree);

[0062] Generate multiple responses to the natural language inference question q, use the model to evaluate the responses according to the reference answer, and summarize them into an inference thought chain.

[0063] Initialization: Generate M initial complete responses for the given prompt as the roots of M trees. Among them, the EPTree algorithm is one of the core innovations of this method. It aims to optimize the tree search process for reinforcement learning training, especially focusing on the PassRate metric (i.e., the ability to generate diverse and correct answers under a given inference budget), while ensuring the algorithm is efficient and highly parallelizable to adapt to the characteristics of modern large language model inference engines.

[0064] For the given prompt x, EPTree first generates M initial complete responses in parallel, and these responses serve as the initialization of M independent trees. Specifically, for each tree T i , use the policy model π θ to generate a complete response y iThis step provides diverse starting points for subsequent tree search, enabling the final tree structure to cover a broader solution space.

[0065] These initial responses form a set of trees T = {T i}, where each tree contains only one complete path at the initial stage. This parallel initialization strategy not only improves efficiency but also increases the diversity of the final tree structure, facilitating the model to explore a wider range of solutions.

[0066] Fork point selection: Calculate and select the top-N tokens in each tree as fork points based on the entropy value (representing the degree of uncertainty of the model's prediction at a certain position). Among them, EPTree expands the tree structure by identifying the token with the highest entropy value (uncertainty) in the existing tree as the fork point. For each token in tree T i , calculate its entropy value, which represents the degree of uncertainty of the model at that position. The higher the entropy value, the more likely it is that this position is a decision point where the model believes there are multiple reasonable continuations, and it is also an ideal fork point for exploring different reasoning paths. In actual operation, the algorithm will select the top-N tokens with the highest entropy values as fork points. To promote more effective exploration, EPTree also masks the tokens near the end of the sequence, ensuring that the model explores different reasoning paths rather than just trying different final answer expressions. This entropy-based selection mechanism can specifically create branches at the most uncertain decision points, greatly improving the efficiency and diversity of tree search.

[0067] Expansion: Continue to generate new content from the selected fork points to form new reasoning branches until a complete response is generated.

[0068] Iteration: Repeat the above fork and expansion processes L times to finally form a tree structure containing rich and diverse reasoning paths. Among them, after determining the fork points, EPTree continues to generate T different candidate responses from each selected fork point until a complete solution is obtained. This process creates multiple different reasoning paths for each fork point, enriching the structure and diversity of the tree. For each tree T i, all newly generated branches are added to the original tree structure, forming a more enriched and diverse tree. This expansion method allows the model to explore multiple possible reasoning paths from intermediate states without having to start from scratch, significantly improving the exploration efficiency. After initialization, the EPTree repeats the forking and expansion process L times. Each iteration continues to expand based on the high-entropy points in the current tree structure, ultimately forming a rich tree structure containing M×(N×L×T + 1) leaf nodes (complete responses). In practice, the EPTree usually only requires about 2 iterations to construct a tree that is diverse and information-rich enough, giving it a significant advantage in terms of computational efficiency. Compared with independent multi-chain sampling, the EPTree can generate approximately twice as many different responses at the same reasoning cost, which significantly enhances the exploration efficiency and the quality of the learning signal during the reinforcement learning process.

[0069] Step 2: Reinforcement Learning Based on Process Supervision and Tree Search

[0070] Generate a tree using the above tree search method.

[0071] Value Estimation: Use the Monte Carlo method to estimate the value of each node in the tree, calculated as the proportion of correct answers among all leaf nodes reachable from that node. For each step s n (corresponding to a node in the tree), first use the Monte Carlo method to estimate its value. Let L(s n ) represent the set of all descendant leaf nodes of node s n (including node s n itself if it is a leaf node). The value V(s n ) of node s n is calculated as:

[0072]

[0073] This value definition intuitively reflects the likelihood of reaching the correct answer from this node and provides a basis for subsequent reward calculation.

[0074] In a specific embodiment, the method also incorporates two complementary advantage metrics, including:

[0075] Global Advantage (G A ): Represents the potential advantage of the current step relative to the overall correct rate of all samples. The global advantage is calculated by comparing the difference between the node value and the value of the root node (representing the average correct rate of all responses). This metric evaluates the superiority of a certain reasoning step compared to the overall average level and reflects the value of this step at the macroscopic level.

[0076] G A (s n ) = V(sn ) - V(root)

[0077] Among them, V(root) represents the value of the root node, which represents the average correct rate of all generated responses.

[0078] Local Advantage, L A ): Quantify the improvement brought by the current step relative to its direct parent step. The local advantage is calculated by comparing the difference between the node value and its parent node value. A positive local advantage indicates that this step is more likely to lead to the correct answer compared to its parent node and should be encouraged; a negative local advantage indicates that this step reduces the possibility of reaching the correct answer and should be suppressed.

[0079] L A (s n ) = V(s n ) - V(p(s n ))

[0080] A positive local advantage indicates that this step is more likely to lead to the correct answer compared to its parent node and should be encouraged; a negative local advantage indicates that this step reduces the possibility of reaching the correct answer and should be suppressed. By combining the global and local advantages, this patent calculates a comprehensive reward for each reasoning step. This dual - advantage mechanism can comprehensively evaluate the contributions of each step at different scales and provide a more accurate and effective supervision signal.

[0081] R(s n ) = G A (s n ) + L A (s n )

[0082] To prevent overfitting during the training process, this patent introduces an innovative reward re - weighting mechanism. Specifically, for the rewards of non - leaf - node steps, they are re - weighted by dividing by the square root of the number of leaf nodes in their sub - trees. This weighting strategy takes into account the importance and influence scope of the nodes, ensuring that key steps shared by multiple paths are not over - reinforced due to frequent occurrence, thus maintaining the balance and stability of training. For the rewards of non - leaf - node steps, they are re - weighted by dividing by the square root of the number of leaf nodes in their sub - trees:

[0083]

[0084] This weighting strategy takes into account the importance and influence scope of the nodes, ensuring that key steps shared by multiple paths are not over - reinforced due to frequent occurrence, thus maintaining the balance and stability of training.

[0085] Process supervision signal generation: Calculate the reward signal for each inference step by combining global advantages and local advantages.

[0086] Policy update: Use the generated process supervision signal to update the policy model through the policy gradient method.

[0087] The policy gradient method is used for model optimization, and the training objective is:

[0088] E[r(x,y)-β*log(πθ(y|x) / πref(y|x))]

[0089] where r(·) is the reward function based on the process supervision signal, π ref is the reference model, and β is the KL efficiency parameter (default value is 10 -4 ).

[0090] In each training iteration, first use the EPTree to generate a tree structure for a batch of prompts, and then assign process supervision signals calculated based on global and local advantages to each step in the tree. The complete sequence extracted from the tree is used to update the policy model, and the model parameters are optimized by the gradient descent method.

[0091] A training method based on online tree search provided by an embodiment of the present invention introduces an entropy-guided tree search strategy. By calculating and selecting the token with the highest entropy value as the bifurcation point for tree expansion, it generates more diverse responses under the same inference cost, which is different from the step-by-step generation method relied on by traditional MCTS, significantly improving the search efficiency and response diversity. Innovatively combine global advantages and local advantages to calculate comprehensive rewards for each inference step in the tree, and at the same time adopt a reward reweighting mechanism based on node importance to prevent overfitting, providing high-quality fine-grained supervision signals without an additional reward model. Enhance exploration diversity through tree search, utilize process supervision to improve learning efficiency, form a closed-loop optimization system, significantly enhance the capabilities of large language models in complex reasoning tasks such as mathematics and programming, and have broad application value.

[0092] Based on the above embodiments, this embodiment uses experimental results to elaborate on this method, specifically as follows:

[0093] Reinforcement learning based on tree search

[0094] Such as Figure 3As shown, the proposed method in this method is evaluated on mathematical and code tasks. MATH500 is a dataset of high school-level math problems, Omni-math-500, AMC, and Olympiad cover high school and college-level math problems, and AIME mainly includes high-difficulty math competition problems. LiveCodeBench is mainly a collection of code problems. It can be seen that the method proposed in this patent achieved remarkable results on these datasets in experiments based on the open-source models GLM and Qwen.

[0095] As Figure 4 , Figure 5 shown, the proposed method in this method measures the effectiveness of the model's inference expansion ability. It can be seen that after using the method proposed in this method, as the amount of reinforcement learning training increases, the performance of the model shows a near-logarithmic linear inference expansion trend as the thinking length increases. This result indicates that the model's inference expansion ability continues to enhance after reinforcement learning, and at the same time, this inference expansion ability can be effectively measured.

[0096] Please refer to Figure 6 , Figure 6 , which is the structural block diagram of a training device based on online tree search provided by an embodiment of the present invention; the specific device may include:

[0097] An initialization module 100, which initializes the given prompt information based on entropy-guided tree search to generate a guiding tree root;

[0098] A tree structure generation module 200, which selects the bifurcation points of the guiding tree root according to the entropy value and expands the bifurcation points of the guiding tree to obtain a tree structure;

[0099] A reinforcement learning module 300, which uses the Monte Carlo method to calculate the node values in the tree structure, calculates the reward signal based on the node values in the tree structure, and strengthens the tree search policy model.

[0100] A training device based on online tree search in this embodiment is used to implement the foregoing training method based on online tree search. Therefore, the specific implementation manners in a training device based on online tree search can be seen in the embodiment part of the foregoing training method based on online tree search. For example, the initialization module 100, the tree structure generation module 200, and the reinforcement learning module 300 are respectively used to implement steps S101, S102, and S103 in the foregoing training method based on online tree search. Therefore, its specific implementation manners can refer to the descriptions of the corresponding respective part embodiments and will not be elaborated herein.

[0101] To implement the above embodiments, the present application also provides an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided in the foregoing embodiments.

[0102] To implement the above embodiments, the present application also provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the method provided in the foregoing embodiments when executed by a processor.

[0103] To implement the above embodiments, the present application also provides a computer program product including a computer program, which implements the method provided in the foregoing embodiments when executed by a processor.

[0104] The collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in the present application all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0105] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of these legitimate uses. In addition, such collection / sharing should be carried out after obtaining the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization including authorizing relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.

[0106] The present application anticipates providing embodiments for users to selectively block the use or access of personal information data. That is, the present disclosure anticipates providing hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of users.

[0107] In the descriptions of the foregoing embodiments, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0108] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0109] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0110] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or other suitable processing as necessary, and then storing it in a computer memory.

[0111] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), and the like.

[0112] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0113] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0114] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.

Claims

1. A training method based on online tree search, characterized in that Comprising: Performing initialization processing on given prompt information based on entropy-guided tree search to generate a guiding tree root; Selecting a bifurcation point of the guiding tree root according to the entropy value, and performing expansion processing on the bifurcation point of the guiding tree to obtain a tree structure; Using the Monte Carlo method to calculate the node values in the tree structure, calculating a reward signal based on the node values in the tree structure, and strengthening the tree search policy model.

2. The training method based on online tree search according to claim 1, wherein, The performing initialization processing on given prompt information based on entropy-guided tree search to generate a guiding tree root includes: For given prompt information, the entropy-guided tree generates multiple initial complete responses as the initialization of multiple independent trees, and the multiple initial complete responses form a tree set to obtain a guiding tree root.

3. The training method based on online tree search according to claim 1, wherein The selecting a bifurcation point of the guiding tree root according to the entropy value, and performing expansion processing on the bifurcation point of the guiding tree to obtain a tree structure includes: Using the entropy-guided tree search to identify the token with the highest entropy value in the existing tree as the bifurcation point, calculating the entropy value of the token, and expanding the tree structure based on the entropy value.

4. The training method based on online tree search according to claim 3, wherein The expanding the tree structure based on the entropy value includes: Continuing to generate multiple different candidate responses based on the bifurcation point, adding the multiple different candidate responses to the original tree structure, and repeating the bifurcation expansion operation to obtain a tree structure.

5. The training method based on online tree search according to claim 1, wherein The using the Monte Carlo method to calculate the node values in the tree structure, calculating a reward signal based on the node values in the tree structure, and strengthening the tree search policy model includes: Calculating the value of each step in the tree structure based on the Monte Carlo method, and optimizing the calculation of the value of each step in the tree structure by adding the global advantage and the local advantage, where the global advantage is calculated by comparing the difference between the node value and the root node value, and the local advantage is calculated by comparing the difference between the node value and its parent node value.

6. The training method based on online tree search according to claim 5, characterized in that, The using the Monte Carlo method to calculate the node values in the tree structure, calculating a reward signal based on the node values in the tree structure, and strengthening the tree search policy model further includes: Introducing a reward reweighting mechanism based on the Monte Carlo method. For the reward of non-leaf node steps, reweighting is performed by dividing by the square root of the number of leaf nodes in its subtree to strengthen the tree search policy model.

7. The training method based on online tree search according to claim 6, wherein The using the Monte Carlo method to calculate the node values in the tree structure, calculating a reward signal based on the node values in the tree structure, and strengthening the tree search policy model further includes: Adopting a policy gradient method for model optimization processing, and its calculation formula is: E[r(x,y)-β*log(πθ(y|x) / πref(y|x))] where r(·) is a reward function based on the process supervision signal, πref is a reference model, and β is a KL efficiency parameter.

8. A training device based on online tree search, characterized in that, Comprising: An initialization module that performs initialization processing on given prompt information based on entropy-guided tree search to generate a guiding tree root; A tree structure generation module that selects a bifurcation point of the guiding tree root according to the entropy value, and performs expansion processing on the bifurcation point of the guiding tree to obtain a tree structure; A reinforcement learning module that uses the Monte Carlo method to calculate the node values in the tree structure, calculates a reward signal based on the node values in the tree structure, and strengthens the tree search policy model.

9. An electronic device, characterized in that, Comprising: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1-7.

Citation Information

Cited By

  • Large-model multi-type tool collaborative reasoning method, system and equipment

    CN120745849A

  • Question and answer task processing model training method and device, equipment and medium

    CN120873610A

  • A question and answer task processing model training method and device, equipment and medium

    CN120873610B