Mathematical problem reasoning trajectory generation method, system and equipment based on large model
Through the combination of probability roadmap and Monte Carlo tree search algorithm, the problems of low efficiency and poor quality of inference trajectory generation in the existing technology are solved, and efficient and high-quality inference trajectory generation are achieved.
Patent Information
- Application Number
- CN202510827622.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the inference trajectory generation method based on a large model is inefficient and the generated trajectory quality cannot be guaranteed to be optimal.
Probability roadmap guidance and Monte Carlo tree search algorithm are used to predict the probability of solution steps through the trained probability roadmap prediction, and path search is performed in combination with Monte Carlo tree search algorithm to generate high-quality mathematical problem inference trajectory.
It significantly improves the generation efficiency and quality of inference trajectories, and can quickly find high-quality inference trajectories in the first round of iterations, avoiding the inefficiency problem caused by random sampling traversing all paths.
Smart Images

Figure CN120354952A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, system, and device for generating an inference trajectory of a mathematical problem based on a large model. Background Art
[0002] Currently, the system generates COT (Chain of Thought) inference trajectories by means of random sampling. Specifically, after a problem is given, the system will use a large model to generate a COT inference trajectory with a correct answer and directly use this COT inference trajectory. However, this method requires random sampling in the tree search space to traverse all paths and obtain the path with the correct final answer as the generated COT inference trajectory with the correct answer. Although this method can generate a COT inference trajectory with the correct answer, its efficiency is low, and the quality of the generated trajectory cannot be guaranteed to be optimal. Therefore, it is urgent to solve this technical problem. Summary of the Invention
[0003] In view of the above problems, this application is proposed to provide a method, system, device, storage medium, and computer program product for generating an inference trajectory of a mathematical problem based on a large model that overcomes the above problems or at least partially solves the above problems. The technical solutions are as follows: In a first aspect, a method for generating an inference trajectory of a mathematical problem based on a large model is provided, including: Obtain a pre-trained probability roadmap, where the probability roadmap is used to predict the exploration probability of the solution steps of a mathematical problem; Receive an instruction from a user to solve a mathematical problem, and use a pre-fine-tuned large model to parse the instruction to parse out the target mathematical problem; Use a trained mathematical problem inference model, take the target mathematical problem as the root node, construct a multi-level inference tree, use the Monte Carlo tree search algorithm for path search to search for the inference trajectory of the target mathematical problem. During the search process, for multiple candidate solution steps at each level, use the probability roadmap to predict the exploration probability of each candidate solution step among the multiple candidate solution steps at this level; based on the predicted exploration probabilities of each candidate solution step, select a target solution step from the multiple candidate solution steps as a sampling node to enter the search for the next level, and finally generate the inference trajectory of the target mathematical problem and provide it to the user.
[0004] In a possible implementation manner, the probability roadmap is trained through the following steps: Construct a training sample set, where the training sample set includes mathematical problems, the solution processes of solving the mathematical problems, and annotation data on whether each solution step in the solution process of solving the mathematical problems is correct; Obtain a preset basic large model; Use the training sample set to train the basic large model to obtain a trained probability roadmap.
[0005] In a possible implementation, using the training sample set to train the basic large model to obtain a trained probability roadmap includes: Use the training sample set and adopt the supervised fine-tuning method to train the basic large model to maximize the probability output of the correct answer steps; Optimize the model parameters through the cross-entropy loss function to obtain a trained probability roadmap.
[0006] In a possible implementation, adopting the probability roadmap to predict the exploration probability of each candidate answer step among multiple candidate answer steps at this level includes: For each candidate answer step among the multiple candidate answer steps at this level, combine the candidate answer step with the answer step of the parent node of the node where the candidate answer step is located, input the combination result into the probability roadmap, and output the correct probability score of the candidate answer step as the exploration probability of the candidate answer step.
[0007] In a possible implementation, based on the predicted exploration probabilities of each candidate answer step, select a target answer step from the multiple candidate answer steps as a sampling node to enter the search of the next level, including: Based on the predicted exploration probabilities of each candidate answer step and combined with the value scores of the nodes where each candidate answer step is located predefined in the Monte Carlo tree search algorithm, calculate the comprehensive probability of each candidate answer step; According to the comprehensive probabilities of each candidate answer step, select a target answer step from the multiple candidate answer steps as a sampling node to enter the search of the next level.
[0008] In a possible implementation, train a mathematical problem reasoning model through the following steps: Preset an initial base large model and perform supervised fine-tuning to fine-tune the format of the output mathematical problem reasoning trajectory to obtain a fine-tuned large model; Receive a sample mathematical problem input by the user, and parse and generate a structured problem representation through the fine-tuned large model; Use the fine-tuned large model as the initial mathematical problem reasoning model. Taking the parsed mathematical problem as the root node, construct a multi-level reasoning tree, and use the Monte Carlo tree search algorithm to perform path search to search for the reasoning trajectory of the parsed mathematical problem. During the search process, for multiple candidate solution steps at each level, use the probability roadmap to predict the exploration probability of each candidate solution step among the multiple candidate solution steps at this level; based on the predicted exploration probabilities of each candidate solution step, select the target solution step from the multiple candidate solution steps as the sampling node to enter the search of the next level, and finally generate and output the reasoning trajectory of the parsed mathematical problem. According to the way of reinforcement learning optimized by the group relative strategy, combined with the output reasoning trajectory of the parsed mathematical problem, train the initial mathematical problem reasoning model to complete the first round of model training, and repeat this process iteratively until the model converges to obtain the trained mathematical problem reasoning model.
[0009] In a possible implementation manner, based on the predicted exploration probabilities of each candidate solution step, selecting the target solution step from the multiple candidate solution steps as the sampling node to enter the search of the next level includes: Based on the predicted exploration probabilities of each candidate solution step, and combined with the value scores of the nodes where each candidate solution step is located predefined in the Monte Carlo tree search algorithm, calculate the comprehensive probability of each candidate solution step. According to the comprehensive probabilities of each candidate solution step, select the target solution step from the multiple candidate solution steps as the sampling node to enter the search of the next level.
[0010] In a possible implementation manner, according to the way of reinforcement learning optimized by the group relative strategy, combined with the output reasoning trajectory of the parsed mathematical problem, training the initial mathematical problem reasoning model to complete the first round of model training includes: According to the way of reinforcement learning optimized by the group relative strategy, calculate the relative advantage of the output reasoning trajectory of the parsed mathematical problem according to the preset reward function, and divide it into the maximum reward group and the minimum reward group. By maximizing the relative advantage of the maximum reward group, update the parameters of the initial mathematical problem reasoning model to complete the first round of model training.
[0011] In a second aspect, a system for generating a reasoning trajectory of a mathematical problem based on a large model is provided, including: An acquisition unit, configured to acquire a pre-trained probability roadmap, where the probability roadmap is used to predict the exploration probability of the solution steps of the mathematical problem; An analysis unit, configured to receive an instruction from a user to solve a mathematical problem, and use a pre-fine-tuned large model to analyze the instruction to analyze the target mathematical problem; A generation unit, which is used to use the trained mathematical problem reasoning model, take the target mathematical problem as the root node, construct a multi-level inference tree, use the Monte Carlo tree search algorithm to perform path search, search for the inference trajectory of the target mathematical problem, and during the search process, for multiple candidate solution steps at each level, use the probability roadmap to predict the exploration probability of each candidate solution step among the multiple candidate solution steps at this level; based on the predicted exploration probabilities of each candidate solution step, select the target solution step from the multiple candidate solution steps as the sampling node to enter the search of the next level, and finally generate the inference trajectory of the target mathematical problem and provide it to the user.
[0012] In a possible implementation manner, the system further includes a first training unit, which is used to train the probability roadmap through the following steps: Construct a training sample set, where the training sample set includes mathematical problems, the solution processes for solving the mathematical problems, and the annotation data on whether each solution step in the solution process of solving the mathematical problems is correct; Obtain a preset basic large model; Use the training sample set to train the basic large model to obtain a trained probability roadmap.
[0013] In a possible implementation manner, the first training unit is further used for: Use the training sample set and adopt the supervised fine-tuning method to train the basic large model to maximize the probability output of the correct solution steps; Optimize the model parameters through the cross-entropy loss function to obtain a trained probability roadmap.
[0014] In a possible implementation manner, the generation unit is further used for: For each candidate solution step among the multiple candidate solution steps at this level, combine the candidate solution step with the solution step of the parent node of the node where the candidate solution step is located, input the combination result into the probability roadmap, and output the correct probability score of the candidate solution step as the exploration probability of the candidate solution step.
[0015] In a possible implementation manner, the generation unit is further used for: Based on the predicted exploration probabilities of each candidate solution step, and in combination with the value scores of the nodes where each candidate solution step is located predefined in the Monte Carlo tree search algorithm, calculate the comprehensive probability of each candidate solution step; According to the comprehensive probabilities of each candidate solution step, select the target solution step from the multiple candidate solution steps as the sampling node to enter the search of the next level.
[0016] In a possible implementation, the system further includes a second training unit for training a mathematical problem reasoning model through the following steps: Preset an initial base large model and perform supervised fine-tuning to fine-tune the format of the output mathematical problem reasoning trajectory to obtain a fine-tuned large model; Receive a sample mathematical problem input by the user, and parse and generate a structured problem representation through the fine-tuned large model; Use the fine-tuned large model as the initial mathematical problem reasoning model, take the parsed mathematical problem as the root node, construct a multi-level inference tree, and use the Monte Carlo tree search algorithm to perform path search to search for the reasoning trajectory of the parsed mathematical problem. During the search process, for multiple candidate solution steps at each level, use the probability roadmap to predict the exploration probability of each candidate solution step among the multiple candidate solution steps at that level; based on the predicted exploration probabilities of each candidate solution step, select the target solution step from the multiple candidate solution steps as the sampling node to enter the search for the next level, and finally generate and output the reasoning trajectory of the parsed mathematical problem. According to the method of reinforcement learning optimized by group relative strategy, combine the output reasoning trajectory of the parsed mathematical problem to train the initial mathematical problem reasoning model to complete the first round of model training, and repeat this process until the model converges to obtain a trained mathematical problem reasoning model.
[0017] In a possible implementation, the second training unit is further configured to: Based on the predicted exploration probabilities of each candidate solution step, and in combination with the value scores of the nodes where each candidate solution step is located predefined in the Monte Carlo tree search algorithm, calculate the comprehensive probability of each candidate solution step; According to the comprehensive probabilities of each candidate solution step, select the target solution step from the multiple candidate solution steps as the sampling node to enter the search for the next level.
[0018] In a possible implementation, the second training unit is further configured to: According to the method of reinforcement learning optimized by group relative strategy, calculate the relative advantage of the output reasoning trajectory of the parsed mathematical problem according to a preset reward function, and divide it into a maximum reward group and a minimum reward group. By maximizing the relative advantage of the maximum reward group, update the parameters of the initial mathematical problem reasoning model to complete the first round of model training.
[0019] In a third aspect, an electronic device is provided. The electronic device includes a processor and a memory. Among them, a computer program is stored in the memory, and the processor is configured to run the computer program to execute the method for generating a mathematical problem reasoning trajectory based on a large model described in any one of the above.
[0020] In a fourth aspect, a storage medium is provided, which stores a computer program, wherein the computer program is configured to execute the method for generating an inference trajectory of a mathematical problem based on a large model as described in any one of the above when running.
[0021] In a fifth aspect, a computer program product is provided, including a computer program, wherein the computer program is configured to execute the method for generating an inference trajectory of a mathematical problem based on a large model as described in any one of the above when running.
[0022] By means of the above technical solutions, the method, system, device, storage medium and computer program product for generating an inference trajectory of a mathematical problem based on a large model provided by the embodiments of the present application. The method for generating an inference trajectory of a mathematical problem based on a large model can quickly find an inference trajectory of a high-quality target mathematical problem in the first round of iteration through the guidance of a probabilistic roadmap and the search strategy of the Monte Carlo tree search algorithm, significantly improving the generation efficiency and quality of the inference trajectory. Specifically, the technical effects are as follows: (1) High efficiency: Through the guidance of the probabilistic roadmap, it is possible to quickly locate a high-quality inference trajectory in the search space, avoiding the inefficiency problem caused by randomly sampling and traversing all paths. (2) High quality: By using the search strategy of the Monte Carlo tree search algorithm, it is possible to find the optimal inference trajectory in the tree search space, ensuring that the quality of the generated inference trajectory is higher than that of the random sampling method. (3) Dynamic adjustment: The combination of the probabilistic roadmap and the Monte Carlo tree search algorithm enables the system to dynamically adjust the path selection during the search process, further improving the search efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for describing the embodiments of the present application will be briefly introduced below.
[0024] Figure 1 Shows a schematic diagram of the model framework provided by the embodiments of the present application; Figure 2 Shows a flowchart of the method for generating an inference trajectory of a mathematical problem based on a large model provided by the embodiments of the present application; Figure 3 Shows a schematic diagram of the guidance of a probabilistic roadmap and the search strategy of the Monte Carlo tree search algorithm provided by the embodiments of the present application; Figure 4 Shows a structural diagram of the system for generating an inference trajectory of a mathematical problem based on a large model provided by the embodiments of the present application; Figure 5 Shows a structural diagram of the system for generating an inference trajectory of a mathematical problem based on a large model provided by another embodiment of the present application; Figure 6 The structure diagram of an electronic device provided by an embodiment of the present application is shown. Detailed implementation manners
[0025] Hereinafter, the exemplary embodiments of the present application will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be completely conveyed to those skilled in the art.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such use can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the term "including" and its variants should be interpreted as open-ended terms meaning "including but not limited to".
[0027] To solve the above technical problems, an embodiment of the present application provides a method for generating a reasoning trajectory of a mathematical problem based on a large model. Here, the large model (Large Model, LM), first, as the name implies, is large in scale, with network parameters reaching tens of billions, hundreds of billions or even more; second, generality, which means not limited to a specific problem or field; third, emergence, that is, the emergence of unexpected new capabilities.
[0028] First, Figure 1 The schematic diagram of the model framework provided by an embodiment of the present application is shown. In Figure 1 , the preset initial base large model is fine-tuned through SFT (Supervised Fine-Tuning) for the format of the reasoning trajectory of the output mathematical problem to obtain a fine-tuned large model; then, under the guidance of the probability roadmap and the search strategy of the Monte Carlo tree search algorithm, according to the reinforcement learning method of GRPO (Group Relative Policy Optimization), the model is trained to complete the first round of model training, and this process is repeated iteratively until the model converges to obtain a trained mathematical problem reasoning model.
[0029] As Figure 2 shown, the method for generating a reasoning trajectory of a mathematical problem based on a large model may include the following steps S201 to S203: Step S201, obtain a pre-trained probability roadmap, where the probability roadmap is used to predict the exploration probability of the solution steps of the mathematical problem; Step S202: Receive an instruction from the user to solve a math problem, and use a pre-fine-tuned large model to parse the instruction to parse out the target math problem; Step S203: Use the trained math problem reasoning model, take the target math problem as the root node, construct a multi-level reasoning tree, use the Monte Carlo tree search algorithm for path search, search for the reasoning trajectory of the target math problem. During the search process, for multiple candidate solution steps at each level, use the probabilistic roadmap to predict the exploration probability of each candidate solution step among the multiple candidate solution steps at this level; based on the predicted exploration probability of each candidate solution step, select the target solution step from the multiple candidate solution steps as the sampling node to enter the search of the next level, and finally generate the reasoning trajectory of the target math problem and provide it to the user.
[0030] In this step, the Monte Carlo tree search (MCTS) algorithm is a simulation-based search algorithm used to find the optimal path in a tree structure.
[0031] Through the guidance of the probabilistic roadmap and the search strategy of the Monte Carlo tree search algorithm in this embodiment, it is possible to quickly find the reasoning trajectory of the high-quality target math problem in the first round of iteration, significantly improving the generation efficiency and quality of the reasoning trajectory.
[0032] A possible implementation manner is provided in the embodiment of the present application. The probabilistic roadmap mentioned in step S201 above can be trained through the following steps A1 to A3: Step A1: Construct a training sample set, where the training sample set includes math problems, the solution processes for solving the math problems, and the labeled data indicating whether each solution step in the solution process of solving the math problems is correct; Step A2: Obtain a preset basic large model; Step A3: Use the training sample set to train the basic large model to obtain a trained probabilistic roadmap.
[0033] In this embodiment, by annotating the correctness of each solution step in the math problem, such as step 1 is correct, step 2 is wrong, step 3 is correct, etc., a dataset with fine-grained supervision signals is constructed, enabling the model to learn local errors in the problem-solving logic chain, rather than simply judging whether the final answer is correct or wrong. This facilitates subsequent use of the probabilistic roadmap to predict the exploration probability of each candidate solution step among multiple candidate solution steps, providing guidance for the search of the Monte Carlo tree search algorithm and making up for the problem of large exploration space caused by the random exploration of the Monte Carlo tree search algorithm.
[0034] In an embodiment of the present application, a possible implementation is provided. In the above step A3, the basic large model is trained using a training sample set to obtain a trained probability roadmap, which may specifically include the following steps A3-1 and A3-2: Step A3-1: Use the training sample set to train the basic large model by means of supervised fine-tuning to maximize the probability output of the correct answer step; Step A3-2: Optimize the model parameters through the cross-entropy loss function to obtain a trained probability roadmap.
[0035] In this embodiment, the probability output of the correct answer step is maximized through supervised fine-tuning, making the model more inclined to generate logical steps that meet expectations, significantly improving the accuracy and reliability of the problem-solving process; and the cross-entropy loss function is used to directly optimize the model parameters, which can effectively measure the difference between the predicted probability distribution and the true distribution, accelerate model convergence, and improve training efficiency.
[0036] In an embodiment of the present application, a possible implementation is provided. In the above step S203, the exploration probability of each candidate answer step among multiple candidate answer steps at this level is predicted using a probability roadmap, which may specifically include the following step B1: Step B1: For each candidate answer step among multiple candidate answer steps at this level, combine the candidate answer step with the answer step of the parent node of the node where the candidate answer step is located, input the combination result into the probability roadmap, and output the correct probability score of the candidate answer step as the exploration probability of the candidate answer step.
[0037] In this embodiment, through the modeling and evaluation of the step sequence by the probability roadmap, a more intelligent search strategy is realized, balancing the trade-off between exploration (trying new steps) and exploitation (prioritizing high-probability steps), and ultimately improving the efficiency and robustness of complex problem-solving.
[0038] In an embodiment of the present application, a possible implementation is provided. In the above step S203, based on the exploration probabilities of the predicted candidate answer steps, a target answer step is selected from multiple candidate answer steps as a sampling node to enter the search of the next level, which may specifically include the following steps C1 and C2: Step C1: Based on the exploration probabilities of the predicted candidate answer steps and in combination with the value scores of the nodes where each candidate answer step is located predefined in the Monte Carlo tree search algorithm, calculate the comprehensive probability of each candidate answer step; Step C2: According to the comprehensive probabilities of each candidate answer step, select a target answer step from multiple candidate answer steps as a sampling node to enter the search of the next level.
[0039] In this embodiment, in the tree search structure (such as multi-step reasoning), instead of blindly traversing all possible branches, steps with high comprehensive probability (i.e., nodes with both high exploration potential and high historical value) are preferentially selected to accelerate convergence to the optimal solution.
[0040] A possible implementation is provided in the embodiment of the present application. Through the following formula, based on the exploration probabilities of each candidate solution step predicted and combined with the value scores of the nodes where each candidate solution step is located, which are predefined in the Monte Carlo tree search algorithm, the comprehensive probability of each candidate solution step is calculated:
[0041]
[0042] Among them, represents the comprehensive probability of selecting the -th solution step in the state ; represents the parent node of the -th solution step; represents the value score of the node where the -th solution step is located in the state , measuring the expectation of future rewards under this step; represents the number of visits to the parent node of the -th solution step; represents the number of visits to the -th solution step with as the parent node; represents the current reward contribution value of selecting the -th solution step in the state ; represents the exploration probability of selecting the -th solution step in the state ; represents the exploration probability of selecting the -th solution step in the state ; represents the solution step, which is a positive integer.
[0043] This embodiment comprehensively considers the exploration probabilities of each candidate solution step and the value scores of the nodes where each candidate solution step is located, which are predefined in the Monte Carlo tree search algorithm, enabling the system to dynamically adjust the path selection during the search process and further improving the search efficiency.
[0044] A possible implementation is provided in the embodiment of the present application. The mathematical problem reasoning model mentioned in step S203 above can be trained through the following steps D1 to D4: Step D1: Preset an initial base large model and perform supervised fine-tuning on the format of the inferred trajectory of the output math problems to obtain a fine-tuned large model; Step D2: Receive a sample math problem input by the user, and parse and generate a structured problem representation through the fine-tuned large model; Step D3: Use the fine-tuned large model as the initial math problem inference model, construct a multi-level inference tree with the parsed math problem as the root node, and use the Monte Carlo tree search algorithm to perform path search to search for the inference trajectory of the parsed math problem. During the search process, for multiple candidate solution steps at each level, use a probabilistic roadmap to predict the exploration probability of each candidate solution step among the multiple candidate solution steps at that level; based on the predicted exploration probabilities of each candidate solution step, select a target solution step from the multiple candidate solution steps as a sampling node to enter the search of the next level, and finally generate and output the inference trajectory of the parsed math problem; Step D4: According to the method of reinforcement learning optimized by the group relative strategy, combine the output inference trajectory of the parsed math problem to train the initial math problem inference model to complete the first round of model training, and repeat and iterate this process until the model converges to obtain a trained math problem inference model.
[0045] In this embodiment, logical consistency verification is performed on the generated inference trajectory, and paths containing contradictory steps are eliminated. The parameters of the initial math problem inference model can be updated through the reinforcement learning feedback mechanism to optimize the subsequent search efficiency.
[0046] A possible implementation manner is provided in the embodiment of the present application. In the above step D3, based on the predicted exploration probabilities of each candidate solution step, a target solution step is selected from the multiple candidate solution steps as a sampling node to enter the search of the next level, which may specifically include the following steps D3-1 and D3-2: Step D3-1: Based on the predicted exploration probabilities of each candidate solution step, and in combination with the value scores of the nodes where each candidate solution step is located predefined in the Monte Carlo tree search algorithm, calculate the comprehensive probability of each candidate solution step; Step D3-2: According to the comprehensive probabilities of each candidate solution step, select a target solution step from the multiple candidate solution steps as a sampling node to enter the search of the next level.
[0047] In this embodiment, in a tree search structure (such as multi-step inference), blindly traversing all possible branches is avoided, but steps with high comprehensive probabilities (i.e., nodes with both high exploration potential and high historical value) are preferentially selected to accelerate convergence to the optimal solution.
[0048] In an embodiment of the present application, a possible implementation manner is provided. In the above step D4, according to the reinforcement learning method optimized by the group relative strategy, combined with the inference trajectory of the parsed mathematical problem output, the initial mathematical problem inference model is trained to complete the first round of model training. Specifically, according to the reinforcement learning method optimized by the group relative strategy, the inference trajectory of the parsed mathematical problem output is used to calculate the relative advantage according to a preset reward function, and the maximum reward group and the minimum reward group are divided. By maximizing the relative advantage of the maximum reward group, the parameters of the initial mathematical problem inference model are updated to complete the first round of model training. Here, the preset reward function can be set according to actual needs, and this embodiment does not limit this.
[0049] The above has introduced Figure 2 Multiple implementation manners of each link of the illustrated embodiment will be further described below through specific embodiments for the method for generating the inference trajectory of the mathematical problem based on the large model in the embodiment of the present application.
[0050] In a specific embodiment, it includes the main process, the training probability roadmap, and the Monte Carlo tree search algorithm optimization.
[0051] 1) Main process.
[0052] After Figure 1 As shown, the preset initial base large model is fine-tuned through SFT for the format of the output inference trajectory of the mathematical problem to obtain a fine-tuned large model; then, under the guidance of the probability roadmap and the search strategy of the Monte Carlo tree search algorithm, according to the reinforcement learning method of GRPO, the model is trained to complete the first round of model training, and this process is repeated iteratively until the model converges to obtain a trained mathematical problem inference model.
[0053] 2) Training probability roadmap.
[0054] Let the mathematical problem be Q, and the solution process for solving this mathematical problem be Step. Step includes multiple solution steps, such as n solution steps. Let Step = { , n where is a positive integer.
[0055]
[0056] Among them, represents the probability roadmap to be trained, Q represents the mathematical problem, represents the th solution step, represents the exploration probability of selecting the solution step.
[0057] 2.1) Probabilistic roadmap: Annotate data according to steps and train a probabilistic roadmap detector so that it can score the intermediate steps of problem-solving, giving high scores to correct steps and low scores to incorrect steps.
[0058] 2.2) The probabilistic roadmap can be trained as a binary classification model that scores each step, and the score is the logits (probability vector) of the last token of the step.
[0059] 2.3) The role of the probabilistic roadmap is to provide an exploration probability during the search of the Monte Carlo tree search algorithm. Through the score of the probabilistic roadmap, it can be known that the correctness of this step is relatively high, so a larger exploration probability is given. On the contrary, a smaller exploration probability is given to the step with a low score of the probabilistic roadmap to make up for the problem of large exploration space caused by the random exploration of the Monte Carlo tree search.
[0060] 3) Optimization of the Monte Carlo tree search (MCTS) algorithm 3.1) Traditional MCTS starts from a node in the first round and conducts multiple rollouts (simulated explorations) of the entire path, resulting in an overly large exploration space and low efficiency in obtaining high-quality paths. Moreover, the rollouts in subsequent rounds are also sampling explorations based on the value scores of the nodes. In this embodiment, an exploration probability based on the probabilistic roadmap is introduced.
[0061] As Figure 3 shown, use the trained mathematical problem reasoning model, take the target mathematical problem as the root node, construct a multi-level inference tree, and use the Monte Carlo tree search algorithm to perform path search to search for the inference trajectory of the target mathematical problem. The circles are nodes representing solution steps, and the numbers in the circles are the value scores of each node; during the search process, for multiple candidate solution steps at each level, use the probabilistic roadmap to predict the exploration probability of each candidate solution step among the multiple candidate solution steps at this level; based on the predicted exploration probabilities of each candidate solution step and combined with the value scores of the nodes where each candidate solution step is located predefined in the Monte Carlo tree search algorithm, calculate the comprehensive probability of each candidate solution step; according to the comprehensive probabilities of each candidate solution step, select the target solution step from the multiple candidate solution steps as the sampling node to enter the search of the next level, and finally generate the inference trajectory of the target mathematical problem and provide it to the user.
[0062] For example, for the three solution steps at the first level, a probabilistic roadmap is used to predict the exploration probabilities of each candidate solution step among the three solution steps at this level. Based on the predicted exploration probabilities of each candidate solution step and combined with the value scores of the nodes where each candidate solution step is located as predefined in the Monte Carlo tree search algorithm, the comprehensive probability of each candidate solution step is calculated; according to the comprehensive probabilities of each candidate solution step, a target solution step is selected from multiple candidate solution steps as a sampling node to enter the search for the next level.
[0063] Similarly, for the two solution steps in the left path at the second level, a probabilistic roadmap is used to predict the exploration probabilities of each candidate solution step among the two solution steps at this level. Based on the predicted exploration probabilities of each candidate solution step and combined with the value scores of the nodes where each candidate solution step is located as predefined in the Monte Carlo tree search algorithm, the comprehensive probability of each candidate solution step is calculated; according to the comprehensive probabilities of each candidate solution step, a target solution step is selected from multiple candidate solution steps as a sampling node to enter the search for the next level.
[0064] It should be noted that Figure 3 What is shown is only illustrative and does not limit this embodiment.
[0065] 3.2) Through the following formula, based on the predicted exploration probabilities of each candidate solution step and combined with the value scores of the nodes where each candidate solution step is located as predefined in the Monte Carlo tree search algorithm, the comprehensive probability of each candidate solution step is calculated:
[0066]
[0067] Wherein, represents the comprehensive probability of selecting the th solution step in state , represents the parent node of the th solution step; represents the value score of the node where the th solution step is located in state , measuring the expectation of future rewards under this step; represents the number of visits to the parent node of the th solution step; represents the number of visits to the th solution step with as the parent node; represents the current reward contribution value of selecting the th solution step in state ; Indicates in the state under which the exploration probability of selecting the th solution step is; Indicates in the state under which the exploration probability of selecting the th solution step is; Indicates the solution step, which is a positive integer.
[0068] The present embodiment can achieve the following technical effects: (1) High efficiency: Guided by the probability roadmap, it can quickly locate high-quality reasoning trajectories in the search space, avoiding the inefficiency problem caused by randomly sampling and traversing all paths; (2) High quality: Using the search strategy of the Monte Carlo tree search algorithm, it can find the optimal reasoning trajectory in the tree search space, ensuring that the quality of the generated reasoning trajectory is higher than that of the random sampling method; (3) Dynamic adjustment: The combination of the probability roadmap and the Monte Carlo tree search algorithm enables the system to dynamically adjust the path selection during the search process, further improving the search efficiency.
[0069] It should be noted that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. In practical applications, all the above possible implementation manners can be combined arbitrarily to form possible embodiments of the present application, which will not be elaborated herein one by one.
[0070] Based on the method for generating a reasoning trajectory of a mathematical problem based on a large model provided in the above embodiments, based on the same inventive concept, the embodiments of the present application also provide a system for generating a reasoning trajectory of a mathematical problem based on a large model.
[0071] Figure 4 is the structural diagram of the system for generating a reasoning trajectory of a mathematical problem based on a large model provided by the embodiments of the present application. As Figure 4 shown, the system for generating a reasoning trajectory of a mathematical problem based on a large model may specifically include an acquisition unit 410, an analysis unit 420, and a generation unit 430.
[0072] The acquisition unit 410 is configured to acquire a pre-trained probability roadmap, where the probability roadmap is used to predict the exploration probability of the solution steps of the mathematical problem; The analysis unit 420 is configured to receive an instruction from a user to solve a mathematical problem, and use a pre-fine-tuned large model to analyze the instruction to parse out the target mathematical problem; A generation unit 430 is configured to use the trained mathematical problem reasoning model, take the target mathematical problem as the root node, construct a multi-level inference tree, perform path search using the Monte Carlo tree search algorithm, search for the inference trajectory of the target mathematical problem, and during the search process, for multiple candidate solution steps at each level, use the probability roadmap to predict the exploration probability of each candidate solution step among the multiple candidate solution steps at this level; based on the predicted exploration probabilities of each candidate solution step, select a target solution step from the multiple candidate solution steps as a sampling node to enter the search for the next level, and finally generate the inference trajectory of the target mathematical problem and provide it to the user.
[0073] In an embodiment of the present application, a possible implementation manner is provided, as Figure 5 shown, the system shown above Figure 4 may further include a first training unit 510, configured to train the probability roadmap through the following steps: Construct a training sample set, where the training sample set includes mathematical problems, solution processes for solving the mathematical problems, and annotation data on whether each solution step in the solution process of solving the mathematical problems is correct; Obtain a preset basic large model; Use the training sample set to train the basic large model to obtain a trained probability roadmap.
[0074] In an embodiment of the present application, a possible implementation manner is provided, and the first training unit 510 is further configured to: Use the training sample set to train the basic large model by using a supervised fine-tuning method to maximize the probability output of correct solution steps; Optimize the model parameters through a cross-entropy loss function to obtain a trained probability roadmap.
[0075] In an embodiment of the present application, a possible implementation manner is provided, and the generation unit 430 is further configured to: For each candidate solution step among the multiple candidate solution steps at this level, combine the candidate solution step with the solution step of the parent node of the node where the candidate solution step is located, input the combination result into the probability roadmap, and output the correct probability score of the candidate solution step as the exploration probability of the candidate solution step.
[0076] In an embodiment of the present application, a possible implementation manner is provided, and the generation unit 430 is further configured to: Based on the predicted exploration probabilities of each candidate solution step and in combination with the value scores of the nodes where each candidate solution step is located predefined in the Monte Carlo tree search algorithm, calculate the comprehensive probability of each candidate solution step; Select a target solution step from the multiple candidate solution steps as a sampling node to enter the search of the next level according to the comprehensive probability of each candidate solution step.
[0077] A possible implementation manner is provided in the embodiments of the present application. For example, Figure 5 as shown above Figure 4 the system shown may further include a second training unit 520, which is used to train a mathematical problem reasoning model through the following steps: Preset an initial base large model and perform supervised fine-tuning to fine-tune the format of the output mathematical problem reasoning trajectory to obtain a fine-tuned large model; Receive a sample mathematical problem input by the user, and parse and generate a structured problem representation through the fine-tuned large model; Use the fine-tuned large model as an initial mathematical problem reasoning model, construct a multi-level inference tree with the parsed mathematical problem as the root node, and use the Monte Carlo tree search algorithm to perform path search to search for the reasoning trajectory of the parsed mathematical problem. During the search process, for multiple candidate solution steps at each level, use the probability roadmap to predict the exploration probability of each candidate solution step among the multiple candidate solution steps at this level; based on the predicted exploration probability of each candidate solution step, select a target solution step from the multiple candidate solution steps as a sampling node to enter the search of the next level, and finally generate and output the reasoning trajectory of the parsed mathematical problem. According to the method of reinforcement learning optimized by the group relative strategy, combine the output reasoning trajectory of the parsed mathematical problem to train the initial mathematical problem reasoning model to complete the first round of model training, and repeat this process iteratively until the model converges to obtain a trained mathematical problem reasoning model.
[0078] A possible implementation manner is provided in the embodiments of the present application. The second training unit 520 is further used for: Based on the predicted exploration probability of each candidate solution step, and in combination with the value score of the node where each candidate solution step is located predefined in the Monte Carlo tree search algorithm, calculate the comprehensive probability of each candidate solution step; Select a target solution step from the multiple candidate solution steps as a sampling node to enter the search of the next level according to the comprehensive probability of each candidate solution step.
[0079] A possible implementation manner is provided in the embodiments of the present application. The second training unit 520 is further used for: According to the method of reinforcement learning optimized based on group relative strategies, the inference trajectory of the parsed mathematical problem output is used to calculate the relative advantage according to a preset reward function, and the maximum reward group and the minimum reward group are divided. By maximizing the relative advantage of the maximum reward group, the parameters of the initial mathematical problem inference model are updated to complete the first round of model training.
[0080] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, including a processor and a memory. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method for generating the inference trajectory of the mathematical problem based on the large model in any one of the above embodiments.
[0081] In an exemplary embodiment, an electronic device is provided, as Figure 6 shown. Figure 6 The electronic device 600 shown in the figure includes: a processor 601 and a memory 603. Among them, the processor 601 and the memory 603 are connected, such as connected through a bus 602. Optionally, the electronic device 600 may further include a transceiver 604. It should be noted that in practical applications, the transceiver 604 is not limited to one, and the structure of the electronic device 600 does not constitute a limitation to the embodiments of the present application.
[0082] The processor 601 may be a CPU (Central Processing Unit, central processor), GPU (Graphics Processing Unit, graphics processor), DSP (Digital Signal Processor, data signal processor), ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), FPGA or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in combination with the disclosure of the present application. The processor 601 may also be a combination that realizes a computing function, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0083] The bus 602 may include a path for transmitting information between the above components. The bus 602 may be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard structure) bus, etc. The bus 602 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 6 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0084] The memory 603 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0085] The memory 603 is used to store the computer program code for executing the solution of this application, and is controlled and executed by the processor 601. The processor 601 is used to execute the computer program code stored in the memory 603 to implement the content shown in the foregoing method embodiments.
[0086] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The illustrated electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.
[0087] Based on the same inventive concept, the embodiments of this application also provide a storage medium in which a computer program is stored. The computer program is set to execute the method for generating the inference trajectory of a mathematical problem based on a large model in any of the foregoing embodiments when running.
[0088] Based on the same inventive concept, the embodiments of this application also provide a computer program product, including a computer program, where the computer program is configured to execute the method for generating the inference trajectory of a mathematical problem based on a large model in any of the foregoing embodiments when running.
[0089] Those skilled in the art can clearly understand the specific working processes of the above-described systems, devices, and modules, and can refer to the corresponding processes in the foregoing method embodiments. For the sake of brevity, they will not be described in detail here.
[0090] Those of ordinary skill in the art can understand that the technical solution of this application can be embodied in the form of a software product in essence, or in whole or in part. The computer software product is stored in a storage medium, which includes a number of program instructions for causing an electronic device (such as a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of this application when the program instructions are running. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0091] Alternatively, all or part of the steps of implementing the foregoing method embodiments can be completed by hardware related to program instructions (such as an electronic device such as a personal computer, a server, or a network device), and the program instructions can be stored in a computer-readable storage medium. When the program instructions are executed by the processor of the electronic device, the electronic device executes all or part of the steps of the methods described in the embodiments of this application.
[0092] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that within the spirit and principle of this application, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the corresponding technical solutions to deviate from the protection scope of this application.
Claims
1. A method for generating the reasoning trajectory of mathematical problems based on large models, characterized in that, Including: Obtain a pre-trained probability roadmap, where the probability roadmap is used to predict the exploration probability of the solution steps of a math problem; Receive an instruction from the user to solve a math problem, and use a pre-fine-tuned large model to parse the instruction to parse out the target math problem; Use the trained math problem reasoning model, use the target math problem as the root node to construct a multi-level inference tree, use the Monte Carlo tree search algorithm for path search, search for the inference trajectory of the target math problem. During the search, for multiple candidate solution steps at each level, use the probability roadmap to predict the exploration probability of each candidate solution step among the multiple candidate solution steps at this level; based on the predicted exploration probabilities of each candidate solution step, select the target solution step from the multiple candidate solution steps as the sampling node to enter the search of the next level, and finally generate the inference trajectory of the target math problem and provide it to the user.
2. The method according to claim 1, wherein Train the probability roadmap through the following steps: Construct a training sample set, where the training sample set includes math problems, the solution process of solving the math problems, and the labeled data on whether each solution step in the solution process of solving the math problems is correct; Obtain a preset basic large model; Use the training sample set to train the basic large model to obtain a trained probability roadmap.
3. The method according to claim 2, characterized in that, Using the training sample set to train the basic large model to obtain a trained probability roadmap includes: Use the training sample set and adopt the supervised fine-tuning method to train the basic large model to maximize the probability output of the correct solution steps; Optimize the model parameters through the cross-entropy loss function to obtain a trained probability roadmap.
4. The method according to claim 2, characterized in that, Using the probability roadmap to predict the exploration probability of each candidate solution step among the multiple candidate solution steps at this level includes: For each candidate solution step among the multiple candidate solution steps at this level, combine the candidate solution step with the solution step of the parent node of the node where the candidate solution step is located, input the combination result into the probability roadmap, and output the correct probability score of the candidate solution step as the exploration probability of the candidate solution step.
5. The method according to any one of claims 1 to 4, characterized in that Based on the predicted exploration probabilities of each candidate solution step, select the target solution step from the multiple candidate solution steps as the sampling node to enter the search of the next level, including: Based on the predicted exploration probabilities of each candidate solution step, and combining the value scores of the nodes where each candidate solution step is located predefined in the Monte Carlo tree search algorithm, calculate the comprehensive probability of each candidate solution step; According to the comprehensive probability of each candidate solution step, select the target solution step from the multiple candidate solution steps as the sampling node to enter the search of the next level.
6. The method according to claim 1, wherein Train the math problem reasoning model through the following steps: The preset initial base large model is fine-tuned through supervised fine-tuning to fine-tune the format of the output math problem inference trajectory to obtain a fine-tuned large model; Receive the sample math problem input by the user, and parse and generate a structured problem representation through the fine-tuned large model; Use the fine-tuned large model as the initial mathematical problem reasoning model. Taking the parsed mathematical problem as the root node, construct a multi-level reasoning tree, and use the Monte Carlo tree search algorithm to perform path search to search for the reasoning trajectory of the parsed mathematical problem. During the search process, for multiple candidate solution steps at each level, use the probability roadmap to predict the exploration probability of each candidate solution step among the multiple candidate solution steps at that level; based on the predicted exploration probabilities of each candidate solution step, select the target solution step from the multiple candidate solution steps as the sampling node to enter the search of the next level, and finally generate and output the reasoning trajectory of the parsed mathematical problem; According to the reinforcement learning method optimized by the group relative strategy, combined with the output reasoning trajectory of the parsed mathematical problem, train the initial mathematical problem reasoning model to complete the first round of model training. Repeat and iterate this process until the model converges to obtain the trained mathematical problem reasoning model.
7. The method according to claim 6, wherein Based on the predicted exploration probabilities of each candidate solution step, selecting the target solution step from the multiple candidate solution steps as the sampling node to enter the search of the next level includes: Based on the predicted exploration probabilities of each candidate solution step, and combined with the value scores of the nodes where each candidate solution step is located predefined in the Monte Carlo tree search algorithm, calculate the comprehensive probability of each candidate solution step; According to the comprehensive probability of each candidate solution step, select the target solution step from the multiple candidate solution steps as the sampling node to enter the search of the next level.
8. The method according to claim 6, characterized in that, According to the reinforcement learning method optimized by the group relative strategy, combined with the output reasoning trajectory of the parsed mathematical problem, training the initial mathematical problem reasoning model to complete the first round of model training includes: According to the reinforcement learning method optimized by the group relative strategy, calculate the relative advantage of the output reasoning trajectory of the parsed mathematical problem according to the preset reward function, and divide it into the maximum reward group and the minimum reward group. By maximizing the relative advantage of the maximum reward group, update the parameters of the initial mathematical problem reasoning model to complete the first round of model training.
9. A generation system for the inference trajectory of mathematical problems based on large models, characterized in that, Includes: An acquisition unit for acquiring a pre-trained probability roadmap, where the probability roadmap is used to predict the exploration probability of the solution steps of the mathematical problem; An analysis unit for receiving an instruction from the user to solve a mathematical problem, and using a pre-fine-tuned large model to analyze the instruction to analyze the target mathematical problem; A generation unit for using the trained mathematical problem reasoning model, taking the target mathematical problem as the root node, constructing a multi-level reasoning tree, using the Monte Carlo tree search algorithm to perform path search to search for the reasoning trajectory of the target mathematical problem. During the search process, for multiple candidate solution steps at each level, use the probability roadmap to predict the exploration probability of each candidate solution step among the multiple candidate solution steps at that level; based on the predicted exploration probabilities of each candidate solution step, select the target solution step from the multiple candidate solution steps as the sampling node to enter the search of the next level, and finally generate the reasoning trajectory of the target mathematical problem and provide it to the user.
10. The system according to claim 9, wherein It further includes a first training unit for training a probability roadmap through the following steps: Construct a training sample set, where the training sample set includes math problems, solution processes for solving the math problems, and labeled data indicating whether each step in the solution process of solving the math problems is correct; Obtain a preset basic large model; Use the training sample set to train the basic large model to obtain a trained probability roadmap.
11. The system according to claim 10, wherein, The first training unit is further used for: Use the training sample set to train the basic large model by means of supervised fine-tuning to maximize the probability output of correct solution steps; Optimize the model parameters through a cross-entropy loss function to obtain a trained probability roadmap.
12. The system according to claim 10, wherein The generation unit is further used for: For each candidate solution step among multiple candidate solution steps at this level, combine the candidate solution step with the solution step of the parent node of the node where the candidate solution step is located, input the combination result into the probability roadmap, and output the correct probability score of the candidate solution step as the exploration probability of the candidate solution step.
13. The system according to any one of claims 9 to 12, characterized in that, The generation unit is further used for: Based on the exploration probabilities of the predicted candidate solution steps and in combination with the value scores predefined for the nodes where the candidate solution steps are located in the Monte Carlo tree search algorithm, calculate the comprehensive probabilities of the candidate solution steps; According to the comprehensive probabilities of the candidate solution steps, select a target solution step from the multiple candidate solution steps as a sampling node to enter the search of the next level.
14. The system according to claim 9, wherein It further includes a second training unit for training a math problem reasoning model through the following steps: Preset an initial base large model and perform supervised fine-tuning to fine-tune the format of the output math problem reasoning trajectory to obtain a fine-tuned large model; Receive a sample math problem input by the user, parse and generate a structured problem representation through the fine-tuned large model; Use the fine-tuned large model as the initial math problem reasoning model, construct a multi-level inference tree with the parsed math problem as the root node, use the Monte Carlo tree search algorithm to perform path search to search for the reasoning trajectory of the parsed math problem. During the search process, for multiple candidate solution steps at each level, use the probability roadmap to predict the exploration probability of each candidate solution step among the multiple candidate solution steps at this level; based on the exploration probabilities of the predicted candidate solution steps, select a target solution step from the multiple candidate solution steps as a sampling node to enter the search of the next level, and finally generate and output the reasoning trajectory of the parsed math problem; According to the method of reinforcement learning optimized by group relative strategy, in combination with the output reasoning trajectory of the parsed math problem, train the initial math problem reasoning model to complete the first round of model training, and repeat this process iteratively until the model converges to obtain a trained math problem reasoning model.
15. The system according to claim 14, wherein The second training unit is further used for: Based on the exploration probabilities of the predicted candidate solution steps and in combination with the value scores predefined for the nodes where the candidate solution steps are located in the Monte Carlo tree search algorithm, calculate the comprehensive probabilities of the candidate solution steps; According to the comprehensive probabilities of each candidate solution step, a target solution step is selected from the multiple candidate solution steps as a sampling node to enter the search at the next level.
16. The system according to claim 14, wherein The second training unit is further configured to: According to the reinforcement learning method optimized by the group relative strategy, calculate the relative advantage of the inference trajectory of the parsed math problem output according to a preset reward function, divide the maximum reward group and the minimum reward group, and update the parameters of the initial math problem inference model by maximizing the relative advantage of the maximum reward group to complete the first round of model training.
17. An electronic device, characterized in that, It includes a processor and a memory. Among them, a computer program is stored in the memory, and the processor is configured to run the computer program to execute the method for generating the inference trajectory of the math problem based on the large model according to any one of claims 1 to 8.
18. A storage medium, characterized in that, A computer program is stored in the storage medium, wherein the computer program is configured to execute the method for generating the inference trajectory of the math problem based on the large model according to any one of claims 1 to 8 when running.
19. A computer program product comprising a computer program, characterized in that, The computer program is configured to execute the method for generating the inference trajectory of the math problem based on the large model according to any one of claims 1 to 8 when running.
Citation Information
Patent Citations
Theorem prediction method and device, electronic equipment and storage medium
CN118069915A
Question answering method and device, equipment, medium and program product
CN119990314A
Systems and methods for artificial intelligence agents
US20250139411A1
Cited By
Question and answer task processing model training method and device, equipment and medium
CN120873610A