Training method and device for model for solving mathematical problem
By performing hierarchical training on mathematical models and reinforcement learning with fine-grained reward data, the problem of training instability caused by sparse rewards was solved, and the accuracy and efficiency of the model in solving mathematical problems were improved.
Patent Information
- Application Number
- CN202510833529.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-09
AI Technical Summary
In existing technologies, when using reinforcement learning to train mathematical models, the reward data is sparse, resulting in inefficient and unstable training. The model is prone to falling into local optimality or catastrophic forgetting, making it difficult to improve the model's performance in mathematical reasoning and logical deduction.
After the training model solves multiple math problems, the incorrect questions are identified and trained in stages. The reinforcement learning algorithm is used to calculate rewards based on the weights of the correct solution steps, gradually improving the model performance. Simple problems are trained first and then complex problems to avoid training instability caused by directly facing complex problems.
This improves the stability and efficiency of model training, enabling the model to learn fine-grained problem-solving steps and improve the accuracy and efficiency of solving mathematical problems.
Smart Images

Figure CN120611779A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this specification relate to the field of computer technology, and more particularly to a method for training a model for solving mathematical problems. One or more embodiments of this specification also relate to an apparatus for training a model for solving mathematical problems, a computing device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Significant progress has been made in using large language models (LLMs) to solve mathematical problems. However, further improving these models' performance in mathematical reasoning, logical deduction, and precise calculations, enabling them to stably and efficiently output high-quality solutions to mathematical problems, remains a technical challenge. In particular, using reinforcement learning to optimize the mathematical capabilities of these models presents several key challenges that need to be addressed.
[0003] For example, reward data for reinforcement learning is relatively sparse. Specifically, the effectiveness of reinforcement learning relies heavily on high-quality, fine-grained reward data. Related technologies only provide sparse, binary rewards based on whether the final answer is correct or not—for example, a reward of 1 for a correct answer and 0 for an incorrect answer. This sparse reward data makes it difficult for the model to effectively learn the correct problem-solving paths and strategies, resulting in inefficient training and unstable convergence.
[0004] Furthermore, when training models using complex mathematical problems, they can easily become trapped in local optima, experience training instability, and experience reward oscillation. Because the model has been fine-tuned and has already learned some mathematical problem-solving knowledge, it can experience catastrophic forgetting when faced with the sparse reward data and complex mathematical examples required for reinforcement learning.
[0005] Therefore, how to use reinforcement learning to stably and efficiently train models used to solve mathematical problems to improve model performance is an urgent problem that needs to be solved. Summary of the Invention
[0006] In light of this, embodiments of this specification provide a method for training a model for solving math problems. One or more embodiments of this specification also relate to a training apparatus for a model for solving math problems, a computing device, a computer-readable storage medium, and a computer program product, which utilize reinforcement learning to stably and efficiently train a model for solving math problems, thereby improving model performance.
[0007] According to a first aspect of an embodiment of this specification, a method for training a model for solving mathematical problems is provided, comprising: Solving a plurality of sample math problems using the model to be trained to obtain a first solution result; the first solution result includes each model solution step for each of the sample math problems; the model to be trained is a fine-tuned model; Determining, from the plurality of sample math problems, sample math problems that the model to be trained has answered incorrectly based on the first answer result; Determining a first sample math problem and a second sample math problem from sample math problems that are incorrectly answered by the model to be trained; the number of correct answer steps included in the model answering steps of the first sample math problem is greater than the number of correct answer steps included in the model answering steps of the second sample math problem; Based on the first sample math problem, the model to be trained is trained using a preset reinforcement learning algorithm to obtain a pre-trained model; reward data of the preset reinforcement learning algorithm is calculated based on the weights of correct answer steps included in each model solution step; Based on the second sample math problem, the pre-trained model is trained using the preset reinforcement learning algorithm to obtain a trained model.
[0008] According to a second aspect of an embodiment of this specification, there is provided a training device for a model for solving mathematical problems, comprising: a model solution module configured to use the model to be trained to solve a plurality of sample math problems and obtain a first solution result; the first solution result includes each model solution step for each of the sample math problems; the model to be trained is a fine-tuned model; an incorrect math problem determining module, configured to determine, from the plurality of sample math problems, sample math problems that the model to be trained has answered incorrectly, based on the first answer result; a sample math problem determining module configured to determine a first sample math problem and a second sample math problem from the sample math problems that the model to be trained incorrectly answers; wherein the number of correct answer steps included in the model answering steps of the first sample math problem is greater than the number of correct answer steps included in the model answering steps of the second sample math problem; A first training module is configured to train the model to be trained using a preset reinforcement learning algorithm based on the first sample math problem to obtain a pre-trained model; reward data of the preset reinforcement learning algorithm is calculated based on the weights of correct answer steps included in each model solution step; The second training module is configured to train the pre-trained model based on the second sample math problem using the preset reinforcement learning algorithm to obtain a trained model.
[0009] According to a third aspect of an embodiment of this specification, a computing device is provided, including: memory and processor; The memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions. When the computer programs or instructions are executed by the processor, the steps of the above method are implemented.
[0010] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program or instructions, and the computer program or instructions implement the steps of the above method when executed by a processor.
[0011] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.
[0012] At least one embodiment of the present specification achieves the following beneficial effects: the embodiments of the present specification can use the model to be trained to solve multiple sample math problems, determine the sample math problems that the model has answered incorrectly, and obtain first sample math problems containing a larger number of correct answer steps and second sample math problems containing a smaller number of correct answer steps from the incorrectly answered sample math problems. Then, based on the first sample math problems, the model to be trained can be trained using a reinforcement learning algorithm to obtain a pre-trained model, and then based on the second sample math problems, the pre-trained model can be trained using a reinforcement learning algorithm to obtain a trained model, thereby gradually improving the model performance.
[0013] On the other hand, the reward data of the reinforcement learning algorithm of the embodiment of this specification is calculated based on the weights of the correct solution steps contained in each model solution step. Compared with the binary reward data in the related art based on the correctness of the model solution result, the model can obtain fine-grained reward data. For example, the model can learn the various solution steps of mathematical problems, thereby improving the model training efficiency and training stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is a flowchart of a method for training a model for solving math problems provided by one embodiment of this specification; Figure 2 This is a flow chart of calculating reward data provided by one embodiment of this specification; Figure 3This is a schematic diagram of a structure of a training device for a model for solving mathematical problems provided by one embodiment of this specification; Figure 4 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION
[0015] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0016] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0017] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0018] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0019] In this specification, a method for training a model for solving mathematical problems is provided. This specification also relates to a training device for a model for solving mathematical problems, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.
[0020] See also Figure 1 , Figure 1 This is a flowchart of a method for training a model for solving mathematical problems, provided in one embodiment of this specification. From a program perspective, the execution body of the process can be a program installed on a server or model training platform. From a hardware perspective, the execution body of the process can be a server or model training platform capable of model training. The method may specifically include the following steps.
[0021] Step 102: Solve a plurality of sample math problems using the model to be trained to obtain a first solution result; the first solution result includes each model solution step for each of the sample math problems; the model to be trained is a fine-tuned model.
[0022] In the embodiments of this specification, the model to be trained may be a model for solving math problems, and the model may be a fine-tuned model.
[0023] Optionally, you can fine-tune a pre-trained model to obtain the model to be trained. For example, you can use a dataset of labeled sample math problems to fine-tune the pre-trained model. Specifically, the pre-trained model can be a BERT model, a GPT model, or other similar models.
[0024] Optionally, multiple sample math problems can be input into the model to be trained to obtain the solution results of the model to be trained for the multiple sample math problems. For example, multiple sample math problems can be input into the model to be trained in sequence. Specifically, for any sample math problem, it can be input into the model to be trained once, or it can be input into the model to be trained multiple times. As a specific implementation method, multiple sample math problems can also be input into the model to be trained at one time. Specifically, when multiple math problems are input into the model to be trained at one time, they can be input into the model to be trained once, or they can be input into the model to be trained multiple times.
[0025] In the embodiments of this specification, the first solution result includes each model solution step for each sample math problem. Each sample math problem may correspond to at least one model solution step. In practical applications, prompt words can be constructed to instruct the model to be trained to output each model solution step for each sample math problem, and the prompt words can be input into the model to be trained to obtain each model solution step for each sample math problem.
[0026] Step 104: Based on the first answer result, determine, from the plurality of sample math problems, sample math problems that the model to be trained has answered incorrectly.
[0027] In the embodiments of this specification, a sample math problem incorrectly answered by the model to be trained may be a problem that does not include all correct solution steps for the sample math problem. For example, if the model solution steps for a sample math problem do not include all correct solution steps for the sample math problem, then the sample math problem may be a sample math problem incorrectly answered by the model to be trained. Alternatively, the correct solution steps may be steps that can correctly answer the math problem.
[0028] In practical applications, expert experience can be used to determine sample math problems that the model to be trained incorrectly answers from a number of sample math problems. As a specific implementation, a machine algorithm can also be used to determine sample math problems that the model to be trained incorrectly answers. For example, the correct solution steps for multiple sample math problems can be obtained in advance, and then a machine algorithm can be used to compare the model's solution steps for the sample math problems with the correct solution steps for the sample math problems to determine the sample math problems that the model to be trained incorrectly answers.
[0029] Optionally, the sample math problem that the model to be trained incorrectly solves may be one math problem or multiple math problems, which is not limited here.
[0030] In the embodiments of this specification, incorrectly answered sample math problems can be used to train the model to be trained. As a specific implementation, the embodiments of this specification can train the model using only incorrectly answered sample math problems, without using correctly answered sample math problems to train the model. Because the model to be trained already has the ability to answer correctly answered sample math problems, training the model using only incorrectly answered sample math problems can improve model training efficiency. The correctly answered sample math problems can be math problems other than the incorrectly answered sample math problems among the multiple sample math problems.
[0031] Step 106: Determine a first sample math problem and a second sample math problem from the sample math problems that the model to be trained incorrectly answers; the number of correct answer steps included in the model answering steps of the first sample math problem is greater than the number of correct answer steps included in the model answering steps of the second sample math problem.
[0032] In practical applications, in the process of determining sample math problems that are answered incorrectly by the model to be trained, the number of correct solution steps contained in the incorrectly answered sample math problems can also be determined, so that the first sample math problem and the second sample math problem can be determined based on the number of correct solution steps contained in the incorrectly answered sample math problems.
[0033] Optionally, the number of correct solution steps contained in the incorrectly answered sample math problem can be determined based on expert experience, or can be determined by comparing the model solution steps of the sample math problem with the correct solution steps of the sample math problem through a machine learning algorithm.
[0034] Step 108: Based on the first sample math problem, the model to be trained is trained using a preset reinforcement learning algorithm to obtain a pre-trained model; the reward data of the preset reinforcement learning algorithm is calculated based on the weights of the correct answer steps included in each model answer step.
[0035] In practical applications, the capabilities of a fine-tuned model are limited by the quality and coverage of the labeled data. Therefore, reinforcement learning algorithms can be used to train the fine-tuned model, enabling it to find more efficient and robust solutions than the labeled data, thereby improving model performance.
[0036] Reinforcement Learning (RL) is a machine learning method that aims to enable an agent to learn how to make optimal decisions to maximize cumulative reward data by interacting with the environment. The core idea of RL is to learn through trial and error. The agent chooses actions based on the current state of the environment. The environment will provide corresponding feedback, such as reward data or penalty data, and the agent will adjust its strategy based on the feedback. The preset reinforcement learning algorithm used in the embodiments of this specification can be at least one of the Q-Learning algorithm, the Deep Q-Network (DQN) algorithm, and the Policy Gradient Methods.
[0037] In an embodiment of the present specification, a model to be trained can be trained based on a first sample math problem using a preset reinforcement learning algorithm to obtain a pre-trained model. The reward data of the preset reinforcement learning algorithm is calculated based on the weights of the correct answer steps contained in each model answer step. For example, for a sample math problem, the sample math problem includes 5 correct answer steps. If the model obtains 3 correct answer steps, the reward represented by the reward data can be 0.6 or 60%, etc. Alternatively, if the 3 correct answer steps obtained by the model account for 50% of the score of the 5 correct answer steps, the reward represented by the reward data can be 0.5 or 50%.
[0038] In related technologies, when using reinforcement learning to train models based on math problems, the model provides sparse reward data. For example, a binary reward is given based on whether the model's final answer is correct or not. If the model's final answer is correct, the reward data represents a reward of 1; if the model's final answer is incorrect, the reward data represents a reward of 0. This sparse reward data makes it difficult to effectively guide the model in learning the correct problem-solving path and strategy, resulting in low training efficiency and unstable convergence.
[0039] In the embodiments of this specification, however, reward data is not calculated based on whether the final answer of the model solution is correct, such as the reward data indicating a reward of 1 or 0. Instead, rewards are calculated based on the weights of the correct answer steps included in the model solution, such as the reward data indicating a reward of 50%, 60%, etc. This allows the embodiments of this specification to provide fine-grained reward data, thereby improving model training efficiency and training stability.
[0040] Step 110: Based on the second sample math problem, the pre-trained model is trained using the preset reinforcement learning algorithm to obtain a trained model.
[0041] In the embodiments of this specification, after a pre-trained model is obtained by training a model to be trained based on a first sample math problem, the pre-trained model can be trained based on a second sample math problem. Because the model obtains fewer correct solution steps for the second sample math problem and more correct solution steps for the first sample math problem, the second sample math problem is more complex for the model, while the first sample math problem is relatively simple for the model. Therefore, the model can be trained using the first sample math problem first, allowing the model to learn simple math knowledge first. Then, the model can be trained using the second sample math problem, and then the model can be trained using the second sample math problem, allowing the model to learn complex math knowledge. This can gradually improve the model performance, thereby avoiding the problems of training instability and catastrophic forgetting that arise from directly training the model with complex math problems.
[0042] Optionally, the pre-trained model can be trained based on the second sample math problem using the preset reinforcement learning algorithm mentioned above to obtain a trained model.
[0043] In the embodiments of this specification, the trained model can be used to solve math problems.
[0044] For example, when users such as online teaching institutions or book publishers need to generate step-by-step solutions to a large number of math problems in a database, they can input these problems into the trained model, which then generates the solution steps based on the trained model. This avoids the tedious process of manually solving math problems, thereby improving the efficiency and accuracy of problem solving.
[0045] In addition, when users such as teachers, students or parents use teaching software for online learning and need to obtain the solution steps for a certain math problem, the teaching software can be loaded with the above-mentioned trained model or can call the above-mentioned trained model. Then, after the user enters the math problem in the teaching software, the teaching software can use the trained problem-solving model to quickly and accurately generate the solution steps for the above problem and display them, thereby improving the user experience.
[0046] In the embodiments of the present specification, a model to be trained can be used to solve multiple sample math problems, thereby determining, based on the solution results, a first sample math problem containing a larger number of correct solution steps and a second sample math problem containing a smaller number of correct solution steps. This allows for rapid and accurate determination of the first sample math problem that is relatively simple for the model and the second sample math problem that is relatively complex for the model, and further, based on the first sample math problem and the second sample math problem, the model to be trained can be rapidly and accurately trained.
[0047] In the embodiments of this specification, the model can be first trained using a first sample math problem whose model answering steps include a relatively large number of correct answering steps, and then the model can be trained using a second sample math problem whose model answering steps include a relatively small number of correct answering steps, so that the model first learns simple math knowledge and then learns complex math knowledge, thereby gradually improving the model performance and improving the stability of model training.
[0048] In the embodiments of this specification, when a reinforcement learning algorithm is used to train a model, the algorithm's reward data is calculated based on the weights of the correct solutions included in each model's solution steps. Compared to the binary reward data based on the correctness of the model's solution in related technologies, the model can obtain fine-grained reward data. This fine-grained reward data allows the model to learn the correct solution steps for sample math problems, improving model training efficiency and stability.
[0049] based on Figure 1 The present specification also provides some specific implementation plans of the method, which are described below.
[0050] In the embodiment of this specification, solving a plurality of sample math problems using the model to be trained to obtain a first solution may specifically include: Constructing prompt words for inputting the model to be trained; the prompt words include the multiple sample math problems and task instruction information, and the task instruction information is used to instruct the model to be trained to solve the multiple sample math problems.
[0051] The prompt words are input into the model to be trained to obtain a first answer result of the model to be trained for each of the sample math problems.
[0052] In practical applications, a prompt word can be constructed and then input into the model to be trained to obtain a first answer. Specifically, the prompt word can include multiple sample math problems and task instructions, where the task instructions are used to instruct the model to be trained to solve the multiple sample math problems. Optionally, a prompt word can be constructed for each sample math problem, with one prompt word corresponding to each sample math problem. Alternatively, a single prompt word can be constructed for multiple sample math problems, and the single prompt word can include multiple sample math problems.
[0053] As a specific implementation, the prompt word can be input into the model to be trained multiple times to obtain multiple first answer results, or the prompt word can be input into the model to be trained once to obtain a single answer result.
[0054] In an embodiment of the present specification, sample math problems that the model to be trained incorrectly answers can be determined based on the model's solution steps for the sample math problems in the first solution result. Optionally, determining the sample math problems that the model to be trained incorrectly answers from the plurality of sample math problems based on the first solution result can specifically include: Each model solution step of any sample math problem among the plurality of sample math problems is obtained from the first solution result.
[0055] The model solution steps of any sample math problem are compared with the correct solution steps of any sample math problem to obtain a comparison result.
[0056] If the comparison result indicates that the various model solution steps of the any sample math problem do not include the various correct solution steps of the any sample math problem, then the any sample math problem is determined to be a sample math problem that the model to be trained solves incorrectly.
[0057] Optionally, a sample math problem may correspond to one model solution step or multiple model solution steps.
[0058] In the embodiments of this specification, the various model solution steps of any sample math problem can be manually compared with the various correct solution steps of any sample math problem to obtain a comparison result. Optionally, a machine learning algorithm can be used to compare the various model solution steps of any sample math problem with the various correct solution steps of any sample math problem to obtain a comparison result. For example, the model solution steps can be structured or semi-structured parsed using preset grammatical rules, keyword matching algorithms, pattern recognition algorithms, and the like, and then the solution steps in the model solution steps that correspond to the correct solution steps can be identified. For example, if the correct solution steps include a certain algebraic transformation or apply a certain geometric theorem, it can be determined whether the model solution steps include these algebraic transformations, geometric theorems, etc. Furthermore, algorithms based on text similarity, such as edit distance algorithms, cosine similarity algorithms, and Jaccard similarity algorithms, or algorithms based on sequence matching, such as dynamic time warping (DTW) and longest common subsequence (LCS), can be used to identify the solution steps in the model solution steps that correspond to the correct solution steps.
[0059] Optionally, if the comparison result indicates that each model solution step of any sample math problem includes each correct solution step of any sample math problem, then the any sample math problem can be determined as a sample math problem that is correctly solved by the model to be trained.
[0060] Optionally, if the comparison results indicate that the model's solution steps for any sample math problem do not include all correct solution steps for any sample math problem, then the sample math problem is determined to be a sample math problem that the model to be trained incorrectly answers. For example, if a sample math problem includes five correct solution steps, and the model's solution steps for the sample math problem do not include any of the five correct solution steps, then the sample math problem is determined to be a sample math problem that the model to be trained incorrectly answers.
[0061] In practical applications, there may be sample math problems in the sample math problems that the model to be trained cannot understand at all, or that are too complex for the model to be trained. If these sample math problems are used directly to train the model, the model may fall into a local optimum or unstable training may occur. In the embodiment of this specification, sample math problems that the model can understand but not completely answer correctly can be obtained from the sample math problems, so that the model can be trained using these sample math problems, which can improve the stability of model training and the model training effect. Optionally, the sample math problems that the model to be trained answers incorrectly include a third sample math problem and a fourth sample math problem, the model answering step of the third sample math problem includes at least one correct answering step of the sample math problem, and the model answering step of the fourth sample math problem does not include any correct answering step of the sample math problem. Determining the first sample math problem and the second sample math problem from the sample math problems that the model to be trained answers incorrectly may specifically include: A first sample math problem and a second sample math problem are determined from the third sample math problem.
[0062] In the embodiments of this specification, the third sample math problem may be a math problem whose model solution steps include at least one correct solution step. For example, if a sample math problem has a model solution step that includes at least one correct solution step of the sample math problem, then the sample math problem may be the third sample math problem. Optionally, the third sample math problem may include one correct solution step or multiple correct solution steps.
[0063] As a specific implementation method, the fourth sample math problem may be a math problem that does not include any correct answering steps in the model answering steps. For example, for a sample math problem, if the model answering steps for the sample math problem do not include any correct answering steps for the sample math problem, then the sample math problem may be the fourth sample math problem. Optionally, the fourth sample math problem may be a sample math problem that the model to be trained cannot understand at all or is too complex for the model to be trained. In an embodiment of this specification, the first sample math problem and the second sample math problem can be determined from the sample math problems that the model to be trained answered incorrectly, excluding the fourth sample math problem, that is, from the third sample math problem.
[0064] In actual applications, after the model to be trained is trained to obtain a pre-trained model based on the first sample math problems, the math problem-solving ability of the pre-trained model will be higher than that of the model to be trained. Therefore, for the sample math problems that the model to be trained answers incorrectly, the pre-trained model may have the ability to answer them, that is, the pre-trained model may be able to answer these problems correctly. In order to improve the training efficiency and training stability of the pre-trained model and thus obtain a trained model, the embodiment of this specification may use the pre-trained model to answer the sample math problems that the model to be trained answers incorrectly, so as to determine the sample math problems used to train the pre-trained model. Optionally, based on the second sample math problems, the pre-trained model is trained using the preset reinforcement learning algorithm to obtain a trained model, which may specifically include: The second sample math problem and the fourth sample math problem are solved using the pre-trained model to obtain a second solution result; the second solution result includes each model solution step for the second sample math problem and the fourth sample math problem.
[0065] According to the second answer result, sample math problems that are incorrectly answered by the pre-trained model are determined from the second sample math problems and the fourth sample math problems.
[0066] A fifth sample math problem is determined from the sample math problems that are incorrectly answered by the pre-trained model; the number of correct answer steps included in the model answer steps of the fifth sample math problem is greater than a preset number threshold.
[0067] Based on the fifth sample math problem, the pre-trained model is trained using a preset reinforcement learning algorithm to obtain a trained model.
[0068] Optionally, since the pre-trained model is trained based on the first sample math problems and has the ability to solve the first sample math problems, the pre-trained model can be trained by using the sample math problems to be trained that are answered incorrectly by the model except the first sample math problems.
[0069] Furthermore, the pre-trained model can be used to answer the second sample math problem and the fourth sample math problem. Specifically, the pre-trained model can be used to answer the second sample math problem and the fourth sample math problem in the manner mentioned above of answering multiple sample math problems using the model to be trained. For example, a prompt word for inputting the pre-trained model is constructed, and the prompt word includes the second sample math problem and the fourth sample math problem and task instruction information, and the task instruction information is used to instruct the pre-trained model to answer the second sample math problem and the fourth sample math problem. Prompt words can be constructed separately for each sample math problem in the second sample math problem and the fourth sample math problem, or one prompt word can be constructed for the second sample math problem and the fourth sample math problem. The prompt word can be input into the pre-trained model multiple times, or the prompt word can be input into the pre-trained model once.
[0070] In practical applications, for any of the second and fourth sample math problems, the model solution steps for the sample math problem can be compared with the correct solution steps for the sample math problem to determine the sample math problem that the pre-trained model incorrectly answers.
[0071] Furthermore, the model can be trained using sample math problems that contain a large number of correct solution steps among the sample math problems that the pre-trained model incorrectly solves, thereby improving the efficiency and stability of model training. Specifically, a fifth sample math problem can be determined from the sample math problems that the pre-trained model incorrectly solves, and the number of correct solution steps contained in the model solution steps of the fifth sample math problem is greater than a preset number threshold. Then, based on the fifth sample math problem, the pre-trained model is trained using a preset reinforcement learning algorithm to obtain a trained model. The preset number threshold can be set according to actual needs, for example, it can be set to 2, 3, 5, etc., and is not specifically limited here.
[0072] As a specific implementation method, in order to improve the performance of the trained model, the trained model can also be trained. Optionally, after the pre-trained model is trained based on the fifth sample math problem using a preset reinforcement learning algorithm to obtain the trained model, the following steps can also be included: A sixth sample math problem is determined from the sample math problems that are incorrectly answered by the pre-trained model; the number of correct answer steps included in the model answer steps of the sixth sample math problem is not greater than the preset number threshold.
[0073] Based on the sixth sample math problem, the trained model is trained using a preset reinforcement learning algorithm.
[0074] In the embodiments of this specification, the trained model can be trained using sample math problems that contain a number of correct solution steps that is no greater than a preset threshold. For example, the trained model can be used to solve the sixth sample math problem, determine which sample math problems the trained model incorrectly solves, and then train the trained model based on the sample math problems the trained model incorrectly solves.
[0075] In practical applications, the effectiveness of model training using reinforcement learning depends largely on high-quality, fine-grained reward data. The embodiments of this specification can train the model based on high-quality, fine-grained reward data to improve the model training effect. Optionally, based on the first sample math problem, the model to be trained is trained using a preset reinforcement learning algorithm to obtain a pre-trained model, which may specifically include: The first sample math problem is input into the model to be trained to obtain a first model problem-solving step output by the model to be trained.
[0076] Calculate a first weight of the correct solution steps of the first sample math problem included in the first model problem-solving steps.
[0077] Based on the first weight, reward data of the preset reinforcement learning algorithm is determined.
[0078] Based on the reward data of the preset reinforcement learning algorithm, the model to be trained is trained to obtain a pre-trained model.
[0079] In an embodiment of the present specification, a first sample math problem can be input into a model to be trained, so that the model to be trained solves the first sample math problem to obtain a first model problem-solving step of the first sample math model. The first model problem-solving step may include one problem-solving step or multiple problem-solving steps.
[0080] Furthermore, the first model problem-solving steps can be compared with the correct problem-solving steps of the first sample math problem to determine the correct problem-solving steps of the first sample math problem included in the first model problem-solving steps, and then a first weight of the correct problem-solving steps of the first sample math problem included in the first model problem-solving steps can be calculated.
[0081] Optionally, calculating a first weight of a correct solution step of the first sample math problem included in the first model problem-solving step may specifically include: The first weight is calculated based on at least one of a first calculation method and a second calculation method; the first calculation method is to calculate the first weight based on a first proportion of the number of correct solution steps of the first sample math problem contained in the first model problem-solving steps in the number of correct solution steps of the first sample math problem; the second calculation method is to calculate the first weight based on a second proportion of the score of the correct solution steps of the first sample math problem contained in the first model problem-solving steps in the total score of the first sample math problem.
[0082] In an embodiment of this specification, a first weight can be calculated based on a first proportion of the number of correct solution steps of the first sample math problem included in the first model problem-solving steps to the number of correct solution steps of the first sample math problem. Specifically, the first proportion can be used as the first weight. For example, for a sample math problem, the sample math problem includes 10 correct solution steps. If the first model problem-solving steps include 4 correct solution steps, the first weight can be 0.4 or 40%, etc.
[0083] Optionally, the first weight can also be calculated based on the second proportion of the score of the correct solution steps of the first sample math problem contained in the first model problem-solving steps in the total score of the first sample math problem. Specifically, the second proportion can be used as the first weight. For example, for a certain sample math problem, the score of the correct solution steps contained in the sample math problem is 4 points, and the total score of the sample math problem is 20 points, then the first weight can be 0.2 or 20%, etc. Optionally, the proportion of the score of each correct solution step in the total score can be set based on expert experience.
[0084] As a specific implementation, the first weight can also be calculated based on the first proportion and the second proportion. For example, proportion weights can be pre-set for the first proportion and the second proportion. For example, the proportion weight of the first proportion is 0.6, and the proportion weight of the second proportion is 0.4. The first weight is then calculated based on the first proportion, the second proportion, the proportion weight of the first proportion, and the proportion weight of the second proportion.
[0085] In the embodiment of this specification, the reward data of the preset reinforcement learning algorithm can be determined based on the first weight. For example, the first weight can be used as the reward data of the preset reinforcement learning algorithm.
[0086] In practical applications, in order to increase the upper limit of the model's capabilities and enable the model to solve more complex mathematical problems, the embodiments of this specification may also determine the reward data of the preset reinforcement learning algorithm based on the difficulty coefficient of the sample mathematical problems. Optionally, the determination of the reward data of the preset reinforcement learning algorithm based on the first weight may specifically include: Obtain a first difficulty coefficient of the first sample math problem.
[0087] Based on the first weight and the first difficulty coefficient, the reward data of the preset reinforcement learning algorithm is determined; if the first sample math problem includes a seventh sample math problem and an eighth sample math problem with the same first weight, and the first difficulty coefficient of the seventh sample math problem is higher than the first difficulty coefficient of the eighth sample math problem, then the reward represented by the reward data corresponding to the seventh sample math problem is higher than the reward represented by the reward data corresponding to the eighth sample math problem.
[0088] Optionally, the first difficulty coefficient of the first sample math problem can be determined based on expert experience, such as by a math teacher or student. As a specific embodiment, the first difficulty coefficient of the first sample math problem can also be determined based on a machine model, which can be trained based on sample math problems and difficulty coefficients corresponding to the sample math problems. Optionally, sample math problems of different difficulty levels have different difficulty coefficients. The difficulty coefficient of a sample math problem with a higher difficulty level is greater than the difficulty coefficient of a sample math problem with a lower difficulty level.
[0089] In an embodiment of the present specification, reward data for a preset reinforcement learning algorithm can be determined based on a first weight and a first difficulty coefficient. The reward represented by the reward data corresponding to a sample math problem with a higher difficulty level is higher than the reward represented by the reward data corresponding to a sample math problem with a lower difficulty level. For example, if the first sample math problem includes a seventh sample math problem and an eighth sample math problem with the same first weight, and the first difficulty coefficient of the seventh sample math problem is higher than the first difficulty coefficient of the eighth sample math problem, then the reward represented by the reward data corresponding to the seventh sample math problem is higher than the reward represented by the reward data corresponding to the eighth sample math problem.
[0090] Specifically, the product of the first weight and the first difficulty coefficient can be used as the reward represented by the preset reward data of the reinforcement learning algorithm. For example, if the first weight of the first sample math problem is 50% and the first difficulty coefficient is 2, then the reward for the first sample math problem is 40%*2, that is, 80%.
[0091] The embodiments of this specification can use high-quality, fine-grained reward data to train the pre-trained model to improve the model training effect. Optionally, based on the second sample math problem, the pre-trained model is trained using the preset reinforcement learning algorithm to obtain a trained model, which may specifically include: The second sample math problem is input into the pre-trained model to obtain a second model problem-solving step output by the pre-trained model.
[0092] Calculate a second weight of the correct solution steps of the second sample math problem included in the second model problem-solving steps.
[0093] Based on the second weight, reward data of the preset reinforcement learning algorithm is determined.
[0094] Based on the reward data of the preset reinforcement learning algorithm, the pre-trained model is trained to obtain a trained model.
[0095] In an embodiment of the present specification, a second sample math problem can be input into a pre-trained model, and the pre-trained model can be used to solve the second sample math problem to obtain a second model problem-solving step of the second sample math model. The second model problem-solving step can include one problem-solving step or multiple problem-solving steps.
[0096] Furthermore, the second model problem-solving steps can be compared with the correct problem-solving steps of the second sample math problem to determine the correct problem-solving steps of the second sample math problem included in the second model problem-solving steps, and then a second weight of the correct problem-solving steps of the second sample math problem included in the second model problem-solving steps can be calculated.
[0097] Optionally, calculating the second weight of the correct solution steps of the second sample math problem included in the second model problem-solving step may specifically include: The second weight is calculated based on at least one of a third calculation method and a fourth calculation method; the third calculation method is to calculate the second weight based on a third proportion of the number of correct solution steps of the second sample math problem contained in the second model problem-solving steps in the number of correct solution steps of the second sample math problem; the fourth calculation method is to calculate the second weight based on a fourth proportion of the score of the correct solution steps of the second sample math problem contained in the second model problem-solving steps in the total score of the second sample math problem.
[0098] In an embodiment of this specification, the second weight can be calculated based on the third proportion of the number of correct solution steps of the second sample math problem included in the second model problem-solving steps to the number of correct solution steps of the second sample math problem. Specifically, the third proportion can be used as the first weight. For example, for a sample math problem, the sample math problem includes 8 correct solution steps. If the first model problem-solving steps include 6 correct solution steps, the second weight can be 0.75 or 75%, etc.
[0099] Optionally, the second weight can also be calculated based on the fourth proportion of the score of the correct solution steps of the second sample math problem contained in the second model problem-solving steps in the total score of the second sample math problem. Specifically, the fourth proportion can be used as the second weight. For example, for a certain sample math problem, the score of the correct solution steps contained in the sample math problem is 2 points, and the total score of the sample math problem is 5 points, then the first weight can be 0.4 or 40%, etc. Optionally, the proportion of the score of each correct solution step in the total score can be set based on expert experience.
[0100] As a specific implementation, the second weight can also be calculated based on the third proportion and the fourth proportion. For example, the proportion weights can be pre-set for the third proportion and the fourth proportion. For example, the proportion weight of the third proportion is 0.6, and the proportion weight of the fourth proportion is 0.4. Or the proportion weight of the third proportion is 0.7, and the proportion weight of the fourth proportion is 0.3. Then, the second weight is calculated based on the third proportion, the fourth proportion, the proportion weight of the third proportion, and the proportion weight of the fourth proportion.
[0101] In the embodiment of this specification, the reward data of the preset reinforcement learning algorithm can be determined based on the second weight. For example, the second weight can be used as the reward data of the preset reinforcement learning algorithm.
[0102] In practical applications, in order to increase the upper limit of the model's capabilities and enable the model to solve complex mathematical problems, the embodiments of this specification may also determine the reward data of the preset reinforcement learning algorithm based on the difficulty coefficient of the sample mathematical problems. Optionally, the determination of the reward data of the preset reinforcement learning algorithm based on the second weight may specifically include: A second difficulty coefficient of the second sample math problem is obtained.
[0103] Based on the second weight and the second difficulty coefficient, the reward data of the preset reinforcement learning algorithm is determined; if the second sample math problems include a ninth sample math problem and a tenth sample math problem with the same second weight, and the second difficulty coefficient of the ninth sample math problem is higher than the second difficulty coefficient of the tenth sample math problem, then the reward represented by the reward data corresponding to the ninth sample math problem is higher than the reward represented by the reward data corresponding to the tenth sample math problem.
[0104] Optionally, the second difficulty coefficient of the second sample math problem can be determined based on expert experience, such as by a math teacher or student. As a specific implementation, the second difficulty coefficient of the second sample math problem can also be determined based on a machine model, which can be trained based on sample math problems and difficulty coefficients corresponding to the sample math problems. Optionally, sample math problems of different difficulty levels have different difficulty coefficients. The difficulty coefficient of a sample math problem with a higher difficulty level is greater than the difficulty coefficient of a sample math problem with a lower difficulty level.
[0105] In an embodiment of the present specification, the reward data of the preset reinforcement learning algorithm can be determined based on the second weight and the second difficulty coefficient. The reward represented by the reward data corresponding to the sample math problem with a higher difficulty level is higher than the reward represented by the reward data corresponding to the sample math problem with a lower difficulty level. For example, if the second sample math problem includes a ninth sample math problem and a tenth sample math problem with the same second weight, and the second difficulty coefficient of the ninth sample math problem is higher than the second difficulty coefficient of the tenth sample math problem, then the reward represented by the reward data corresponding to the ninth sample math problem is higher than the reward represented by the reward data corresponding to the tenth sample math problem.
[0106] Specifically, the product of the second weight and the second difficulty coefficient can be used as the reward represented by the preset reward data of the reinforcement learning algorithm. For example, if the second weight of the second sample math problem is 30% and the first difficulty coefficient is 3, then the reward for the second sample math problem is 30%*3, which is 90%.
[0107] In order to further illustrate the training scheme of the model for solving math problems, the embodiments of this specification also provide the following specific examples.
[0108] During the model training process, at the beginning of each training cycle or each training stage, the current model is used to solve some or all sample math problems in the database once or multiple times. The sample math problems are divided into different groups based on the answer results, including a full correct group, a full wrong group, and a partial error group. The full correct group includes sample math problems for which the model can obtain all the correct answer steps every time it solves them. The full wrong group includes sample math problems for which the model cannot obtain any correct answer steps every time it solves them. Sample math problems in the full wrong group indicate that the current model's problem-solving ability is still far from being able to solve the sample math problems. The partial error group includes sample math problems other than those in the full correct group and the full wrong group among the sample math problems used for training. The sample math problems in the partial error group represent sample math problems with difficulty within the ability range of the current model and with learning potential.
[0109] Sample math problems from the all-correct and all-wrong groups were filtered out. A pre-set reinforcement learning algorithm was then used to train a math problem-solving model based on the sample math problems from the partial-wrong group. Filtering out sample math problems from the all-correct group prevents wasting computing resources on problems that the model is already capable of solving. Filtering out sample math problems from the all-wrong group prevents ineffective exploration and training of the model on problems beyond its current capabilities, thereby preventing training instability and potentially harming the model's performance on other problems.
[0110] The sample math problems in the error group are divided into different difficulty levels, and courses are constructed. Specifically, the sample math problems in the error group are divided into different difficulty levels according to the difficulty labels of the sample math problems or the model's answering effect on the sample math problems. These sample math problems can be divided into multiple difficulty levels, such as simple, medium, and difficult. Then, a course learning sequence is constructed to specify the proportion of sampling from sample math problems of different difficulty levels at different stages of training. For example, in the initial stage of model training, you can mainly sample from the simple level, and then use the sampled sample math problems to train the model. As the training progresses, the model's problem-solving ability gradually improves, and the proportion of sampling from the medium difficulty level and the difficult difficulty level can be gradually increased.
[0111] The embodiments of this specification can sample efficiently. Specifically, first, by filtering out questions that are too simple or too difficult, we ensure that training resources are concentrated on the questions that are most beneficial to improving model capabilities. Secondly, based on the difficulty level division of sample math problems, training is made more targeted. Finally, following the idea of course learning, the model capabilities are gradually increased, avoiding instability in the early stages of training, helping the model to explore and learn more stably, and ultimately achieving a higher performance level, so that the model can better inherit and exert its basic capabilities in the fine-tuning stage.
[0112] The embodiments of this specification can use a preset reinforcement learning algorithm to train the model. Specifically, based on a preset reward data calculation mechanism, standard reinforcement learning algorithms such as the policy gradient method and PPO (Proximal Policy Optimization) can be used to update model parameters.
[0113] To address the problem that existing sparse reward signals make it difficult for models to effectively learn correct problem-solving paths and strategies, resulting in low training efficiency and unstable convergence, embodiments of this specification provide a reward data calculation mechanism. This mechanism fully utilizes the correct solution steps and the weights of these steps in sample math problems to analyze and evaluate the model's output, thereby providing richer and more instructive feedback signals than binary rewards.
[0114] Specifically, Figure 2This is a flowchart of calculating reward data provided by an embodiment of this specification, such as Figure 2 As shown, the reward data calculation process of the embodiment of this specification is as follows.
[0115] First, the correct solution steps for a sample math problem can be obtained. These correct solution steps can be based on the experience of experts, such as math teachers, who have solved the sample math problem. The sample math problem can then be input into the model to obtain the model solution steps. Specifically, the model solution steps output by the model can be structured or semi-structured using pre-set grammatical rules, keyword matching, pattern recognition, and other methods to identify the model solution steps that correspond to the correct solution steps. Furthermore, the model solution steps can be compared with the correct solution steps to determine the correctness of the model solution steps and obtain the correct solution steps included in the model solution steps. For example, this can determine whether the model correctly performs a certain important algebraic transformation, correctly applies a certain geometric theorem, or correctly executes a certain calculation step. A percentage reward can then be calculated. The percentage reward can be calculated by the proportion of the correct solution steps included in the model solution steps to the total number or score of all correct solution steps. The percentage reward can be expressed as a percentage. For example, if a sample math problem has N correct solution steps, and the model solution steps include M of these correct solution steps, the percentage reward is calculated as M / N × 100%. This percentage reward reflects the model's progress and the correct answers in solving the sample math problems. Even if the final answer is incorrect, the model still receives positive feedback for the correct steps in solving the sample math problems. Furthermore, the difficulty coefficient of the sample math problems solved by the model can be obtained. The final reward data can then be calculated based on the calculated weights and the obtained difficulty coefficient. The final reward data is calculated as K × M / N × 100%, where K is the difficulty coefficient.
[0116] The embodiments of this specification have a fine-grained reward mechanism. This fine-grained reward mechanism has significant advantages. First, the percentage reward overcomes the sparsity of the binary reward, provides the model with a more intensive feedback signal, and helps the model better understand the problem-solving process. Secondly, the introduction of the difficulty coefficient allows the model to obtain higher rewards when it successfully solves difficult problems, which encourages the model to explore and learn more challenging problems, and helps to improve the upper limit of the model's capabilities. Finally, even if the model fails to completely solve the problem, as long as it demonstrates partially correct reasoning or calculations in the problem-solving process, it can obtain a corresponding percentage reward, which avoids the risk of the model's capabilities collapsing due to a long period of lack of positive feedback when facing difficult problems.
[0117] The embodiments of this specification aim to train the model by introducing more detailed and instructive reward data. This reward data not only focuses on the model's final problem-solving results, but also evaluates or utilizes information from the model's problem-solving process to a certain extent, thereby providing a more intensive feedback signal than a simple binary reward, better guiding the model's learning process. This helps the model more quickly understand which behavioral sequences are effective and which are ineffective. Even if the final problem-solving result is incorrect, the model can learn useful information from the feedback of the intermediate process, thereby accelerating convergence.
[0118] The embodiments of this specification also solve the instability problem in the reinforcement learning training process in the related art. In the related art, sparse rewards often lead to drastic changes in the reward signal during the training process, unclear direction of policy update, and prone to drastic fluctuations in model performance, reward oscillations, and even training crashes. The embodiments of this specification can jointly reduce the variance in the training process and stabilize the update direction of the strategy through an improved reward data calculation mechanism and an innovative efficient sampling strategy. Among them, the efficient sampling strategy can more effectively explore the state-action space associated with high rewards, reduce dependence on inefficient or irrelevant samples, and thus enable the model to learn the optimal strategy more stably. This helps to avoid reward oscillations and non-convergence risks during training, ensure that the reinforcement learning process can proceed smoothly, and ultimately achieve continuous and reliable improvement in model performance.
[0119] The embodiments of this specification build a stable and efficient reinforcement learning training framework by combining a fine-grained reward calculation mechanism and an efficient sampling strategy based on performance grouping and curriculum learning. Fine-grained rewards provide richer learning signals and solve the reward sparsity problem in related technologies. The efficient sampling strategy optimizes the distribution of training data, solves the problem of training instability and achieves a gradual increase in model capabilities. The synergy between the fine-grained reward calculation mechanism and the efficient sampling strategy based on performance grouping and curriculum learning enables the embodiments of this specification to overcome the limitations of related technologies, effectively improve the performance of large language models in complex tasks such as mathematical problem solving, and especially provides a feasible solution for efficient and stable reinforcement learning optimization based on fine-tuned models.
[0120] Corresponding to the above-mentioned embodiment of the training method for a model for solving mathematical problems, this specification also provides an embodiment of a training device for a model for solving mathematical problems. Figure 3 This is a schematic diagram of a structure of a training device for a model for solving mathematical problems provided by one embodiment of this specification. Figure 3 As shown, the device may include: The model answering module 302 is configured to use the model to be trained to answer multiple sample math problems and obtain a first answer result; the first answer result includes each model answering step for each of the sample math problems; the model to be trained is a fine-tuned model.
[0121] The incorrect math problem determining module 304 is configured to determine, from the plurality of sample math problems, sample math problems that the model to be trained has answered incorrectly based on the first answer result.
[0122] The sample math problem determination module 306 is configured to determine a first sample math problem and a second sample math problem from the sample math problems that the model to be trained incorrectly answers; the number of correct answer steps included in the model answer steps of the first sample math problem is greater than the number of correct answer steps included in the model answer steps of the second sample math problem.
[0123] The first training module 308 is configured to train the model to be trained based on the first sample math problem using a preset reinforcement learning algorithm to obtain a pre-trained model; the reward data of the preset reinforcement learning algorithm is calculated based on the weights of the correct answer steps included in each model answer step.
[0124] The second training module 310 is configured to train the pre-trained model based on the second sample math problem using the preset reinforcement learning algorithm to obtain a trained model.
[0125] Optionally, the model solution module 302 may be specifically configured to: Constructing prompt words for inputting the model to be trained; the prompt words include the multiple sample math problems and task instruction information, and the task instruction information is used to instruct the model to be trained to solve the multiple sample math problems.
[0126] The prompt words are input into the model to be trained to obtain a first answer result of the model to be trained for each of the sample math problems.
[0127] Optionally, the incorrect math question determining module 304 may be specifically configured to: Each model solution step of any sample math problem among the plurality of sample math problems is obtained from the first solution result.
[0128] The model solution steps of any sample math problem are compared with the correct solution steps of any sample math problem to obtain a comparison result.
[0129] If the comparison result indicates that the various model solution steps of the any sample math problem do not include the various correct solution steps of the any sample math problem, then the any sample math problem is determined to be a sample math problem that the model to be trained solves incorrectly.
[0130] Optionally, the sample math problems that the model to be trained incorrectly answers include a third sample math problem and a fourth sample math problem, the model answering steps of the third sample math problem include at least one correct answering step of the sample math problem, and the model answering steps of the fourth sample math problem do not include any correct answering step of the sample math problem.
[0131] The sample math problem determination module 306 may be specifically configured to: A first sample math problem and a second sample math problem are determined from the third sample math problem.
[0132] Optionally, the step of training the pre-trained model based on the second sample math problem using the preset reinforcement learning algorithm to obtain a trained model may specifically include: The second sample math problem and the fourth sample math problem are solved using the pre-trained model to obtain a second solution result; the second solution result includes each model solution step for the second sample math problem and the fourth sample math problem.
[0133] According to the second answer result, sample math problems that are incorrectly answered by the pre-trained model are determined from the second sample math problems and the fourth sample math problems.
[0134] A fifth sample math problem is determined from the sample math problems that are incorrectly answered by the pre-trained model; the number of correct answer steps included in the model answer steps of the fifth sample math problem is greater than a preset number threshold.
[0135] Based on the fifth sample math problem, the pre-trained model is trained using a preset reinforcement learning algorithm to obtain a trained model.
[0136] Optionally, the device may further include: The math problem determination module is configured to determine a sixth sample math problem from the sample math problems that are incorrectly answered by the pre-trained model; the number of correct answer steps included in the model answer steps of the sixth sample math problem is not greater than the preset number threshold.
[0137] The model training module is configured to train the trained model based on the sixth sample math problem using a preset reinforcement learning algorithm.
[0138] Optionally, the first training module 308 may be specifically configured to: Inputting the first sample math problem into the model to be trained, and obtaining a first model problem-solving step output by the model to be trained; Calculating a first weight of a correct solution step of the first sample math problem included in the problem-solving steps of the first model; Determining reward data of the preset reinforcement learning algorithm based on the first weight; Based on the reward data of the preset reinforcement learning algorithm, the model to be trained is trained to obtain a pre-trained model.
[0139] Optionally, calculating a first weight of a correct solution step of the first sample math problem included in the first model problem-solving step may specifically include: The first weight is calculated based on at least one of a first calculation method and a second calculation method; the first calculation method is to calculate the first weight based on a first proportion of the number of correct solution steps of the first sample math problem contained in the first model problem-solving steps in the number of correct solution steps of the first sample math problem; the second calculation method is to calculate the first weight based on a second proportion of the score of the correct solution steps of the first sample math problem contained in the first model problem-solving steps in the total score of the first sample math problem.
[0140] Optionally, determining the reward data of the preset reinforcement learning algorithm based on the first weight may specifically include: Obtain a first difficulty coefficient of the first sample math problem.
[0141] Based on the first weight and the first difficulty coefficient, the reward data of the preset reinforcement learning algorithm is determined; if the first sample math problem includes a seventh sample math problem and an eighth sample math problem with the same first weight, and the first difficulty coefficient of the seventh sample math problem is higher than the first difficulty coefficient of the eighth sample math problem, then the reward represented by the reward data corresponding to the seventh sample math problem is higher than the reward represented by the reward data corresponding to the eighth sample math problem.
[0142] Optionally, the second training module 310 may be specifically configured to: The second sample math problem is input into the pre-trained model to obtain a second model problem-solving step output by the pre-trained model.
[0143] Calculate a second weight of the correct solution steps of the second sample math problem included in the second model problem-solving steps.
[0144] Based on the second weight, reward data of the preset reinforcement learning algorithm is determined.
[0145] Based on the reward data of the preset reinforcement learning algorithm, the pre-trained model is trained to obtain a trained model.
[0146] Optionally, calculating the second weight of the correct solution steps of the second sample math problem included in the second model problem-solving step may specifically include: The second weight is calculated based on at least one of a third calculation method and a fourth calculation method; the third calculation method is to calculate the second weight based on a third proportion of the number of correct solution steps of the second sample math problem contained in the second model problem-solving steps in the number of correct solution steps of the second sample math problem; the fourth calculation method is to calculate the second weight based on a fourth proportion of the score of the correct solution steps of the second sample math problem contained in the second model problem-solving steps in the total score of the second sample math problem.
[0147] Optionally, determining the reward data of the preset reinforcement learning algorithm based on the second weight may specifically include: A second difficulty coefficient of the second sample math problem is obtained.
[0148] Based on the second weight and the second difficulty coefficient, the reward data of the preset reinforcement learning algorithm is determined; if the second sample math problems include a ninth sample math problem and a tenth sample math problem with the same second weight, and the second difficulty coefficient of the ninth sample math problem is higher than the second difficulty coefficient of the tenth sample math problem, then the reward represented by the reward data corresponding to the ninth sample math problem is higher than the reward represented by the reward data corresponding to the tenth sample math problem.
[0149] The above is a schematic diagram of a training device for a model for solving mathematical problems according to this embodiment. It should be noted that the technical solution for this training device for solving mathematical problems and the technical solution for the training method for solving mathematical problems are based on the same concept. For details not described in detail in the technical solution for the training device for solving mathematical problems, please refer to the description of the technical solution for the training method for solving mathematical problems.
[0150] Figure 4 The block diagram of a computing device 400 according to one embodiment of the present disclosure is shown. Components of the computing device 800 include, but are not limited to, a memory 410 and a processor 420. The processor 420 is connected to the memory 410 via a bus 430, and a database 450 is used to store data.
[0151] Computing device 400 also includes an access device 440 that enables computing device 400 to communicate via one or more networks 460. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. Access device 440 may include one or more of any type of network interface (e.g., a network interface controller (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.
[0152] In one embodiment of the present specification, the above components of the computing device 400 and Figure 4 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 4 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0153] Computing device 400 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 400 can also be a mobile or stationary server.
[0154] The processor 420 is configured to execute the following computer-executable instructions, which implement the steps of the above method when executed by the processor.
[0155] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the aforementioned method for training a model for solving mathematical problems are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned method for training a model for solving mathematical problems.
[0156] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which implement the steps of the above method when executed by a processor.
[0157] The above is an illustrative embodiment of a computer-readable storage medium. It should be noted that the technical solution of this storage medium and the technical solution of the aforementioned method for training a model for solving mathematical problems share the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the aforementioned method for training a model for solving mathematical problems.
[0158] An embodiment of the present specification further provides a computer program product, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above method.
[0159] The above is an illustrative embodiment of a computer program. It should be noted that the technical solution of this computer program and the technical solution of the aforementioned method for training a model for solving mathematical problems are based on the same concept. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solution of the aforementioned method for training a model for solving mathematical problems.
[0160] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0161] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0162] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0163] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0164] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A method for training a model for solving mathematical problems, characterized in that: include: Use the model to be trained to solve multiple sample math problems and obtain the first solution result; The first solution result includes each model solution step for each sample math problem; The model to be trained is a fine-tuned model; Determining, from the plurality of sample math problems, sample math problems that the model to be trained has answered incorrectly based on the first answer result; Determining a first sample math problem and a second sample math problem from sample math problems that are incorrectly answered by the model to be trained; the number of correct answer steps included in the model answering steps of the first sample math problem is greater than the number of correct answer steps included in the model answering steps of the second sample math problem; Based on the first sample math problem, the model to be trained is trained using a preset reinforcement learning algorithm to obtain a pre-trained model; reward data of the preset reinforcement learning algorithm is calculated based on the weights of correct answer steps included in each model solution step; Based on the second sample math problem, the pre-trained model is trained using the preset reinforcement learning algorithm to obtain a trained model.
2. The method according to claim 1, characterized in that The method of solving a plurality of sample math problems using the model to be trained to obtain a first solution result specifically includes: Constructing a prompt word for inputting the model to be trained; the prompt word includes the multiple sample math problems and task instruction information, and the task instruction information is used to instruct the model to be trained to solve the multiple sample math problems; The prompt words are input into the model to be trained to obtain a first answer result of the model to be trained for each of the sample math problems.
3. The method according to claim 1, characterized in that The step of determining, based on the first answer result, from the plurality of sample math problems, sample math problems that the model to be trained has answered incorrectly, specifically includes: Obtaining, from the first solution result, each model solution step of any one of the plurality of sample math problems; Comparing each model solution step of any sample math problem with each correct solution step of any sample math problem to obtain a comparison result; If the comparison result indicates that the various model solution steps of the any sample math problem do not include the various correct solution steps of the any sample math problem, then the any sample math problem is determined to be a sample math problem that the model to be trained solves incorrectly.
4. The method according to claim 1, wherein The sample math problems that the model to be trained incorrectly solves include a third sample math problem and a fourth sample math problem, the model solving steps of the third sample math problem include at least one correct solving step of the sample math problem, and the model solving steps of the fourth sample math problem do not include any correct solving step of the sample math problem; Determining the first sample math problem and the second sample math problem from the sample math problems incorrectly answered by the model to be trained specifically includes: A first sample math problem and a second sample math problem are determined from the third sample math problems.
5. The method according to claim 4, characterized in that The step of training the pre-trained model based on the second sample math problem using the preset reinforcement learning algorithm to obtain a trained model specifically includes: Solving the second sample math problem and the fourth sample math problem using the pre-trained model to obtain a second solution result; the second solution result includes each model solution step for the second sample math problem and the fourth sample math problem; determining, based on the second answer result, sample math problems that the pre-trained model incorrectly answers from the second sample math problems and the fourth sample math problems; Determining a fifth sample math problem from the sample math problems that the pre-trained model incorrectly answers; wherein the number of correct answer steps included in the model answering steps of the fifth sample math problem is greater than a preset number threshold; Based on the fifth sample math problem, the pre-trained model is trained using a preset reinforcement learning algorithm to obtain a trained model.
6. The method according to claim 5, characterized in that After the pre-trained model is trained based on the fifth sample math problem using a preset reinforcement learning algorithm to obtain a trained model, the method further includes: Determining a sixth sample math problem from the sample math problems incorrectly answered by the pre-trained model; wherein the number of correct answer steps included in the model answer steps of the sixth sample math problem is not greater than the preset number threshold; Based on the sixth sample math problem, the trained model is trained using a preset reinforcement learning algorithm.
7. The method according to claim 1, characterized in that The method of training the model to be trained based on the first sample math problem using a preset reinforcement learning algorithm to obtain a pre-trained model specifically includes: Inputting the first sample math problem into the model to be trained, and obtaining a first model problem-solving step output by the model to be trained; Calculating a first weight of a correct solution step of the first sample math problem included in the problem-solving steps of the first model; Determining reward data of the preset reinforcement learning algorithm based on the first weight; Based on the reward data of the preset reinforcement learning algorithm, the model to be trained is trained to obtain a pre-trained model.
8. The method according to claim 7, characterized in that The calculating of the first weight of the correct solution steps of the first sample math problem included in the first model problem-solving step specifically includes: The first weight is calculated based on at least one of a first calculation method and a second calculation method; the first calculation method is to calculate the first weight based on a first proportion of the number of correct solution steps of the first sample math problem contained in the first model problem-solving steps in the number of correct solution steps of the first sample math problem; the second calculation method is to calculate the first weight based on a second proportion of the score of the correct solution steps of the first sample math problem contained in the first model problem-solving steps in the total score of the first sample math problem.
9. The method according to claim 7, characterized in that Determining the reward data of the preset reinforcement learning algorithm based on the first weight specifically includes: Obtaining a first difficulty coefficient of the first sample math problem; Based on the first weight and the first difficulty coefficient, the reward data of the preset reinforcement learning algorithm is determined; if the first sample math problem includes a seventh sample math problem and an eighth sample math problem with the same first weight, and the first difficulty coefficient of the seventh sample math problem is higher than the first difficulty coefficient of the eighth sample math problem, then the reward represented by the reward data corresponding to the seventh sample math problem is higher than the reward represented by the reward data corresponding to the eighth sample math problem.
10. The method according to claim 1, characterized in that The step of training the pre-trained model based on the second sample math problem using the preset reinforcement learning algorithm to obtain a trained model specifically includes: Inputting the second sample math problem into the pre-trained model to obtain a second model problem-solving step output by the pre-trained model; Calculating a second weight of a correct solution step of the second sample math problem included in the second model problem-solving step; Determining reward data of the preset reinforcement learning algorithm based on the second weight; Based on the reward data of the preset reinforcement learning algorithm, the pre-trained model is trained to obtain a trained model.
11. The method according to claim 10, characterized in that The calculating of the second weight of the correct solution steps of the second sample math problem included in the second model problem-solving step specifically includes: The second weight is calculated based on at least one of a third calculation method and a fourth calculation method; the third calculation method is to calculate the second weight based on a third proportion of the number of correct solution steps of the second sample math problem contained in the second model problem-solving steps in the number of correct solution steps of the second sample math problem; the fourth calculation method is to calculate the second weight based on a fourth proportion of the score of the correct solution steps of the second sample math problem contained in the second model problem-solving steps in the total score of the second sample math problem.
12. The method according to claim 10, characterized in that The step of determining the reward data of the preset reinforcement learning algorithm based on the second weight specifically includes: Obtaining a second difficulty coefficient of the second sample math problem; Based on the second weight and the second difficulty coefficient, the reward data of the preset reinforcement learning algorithm is determined; if the second sample math problems include a ninth sample math problem and a tenth sample math problem with the same second weight, and the second difficulty coefficient of the ninth sample math problem is higher than the second difficulty coefficient of the tenth sample math problem, then the reward represented by the reward data corresponding to the ninth sample math problem is higher than the reward represented by the reward data corresponding to the tenth sample math problem.
13. A training device for a model for solving mathematical problems, characterized in that: include: a model solving module configured to solve a plurality of sample math problems using the model to be trained to obtain a first solving result; The first solution result includes each model solution step for each sample math problem; the model to be trained is a fine-tuned model; an incorrect math problem determining module, configured to determine, from the plurality of sample math problems, sample math problems that the model to be trained has answered incorrectly, based on the first answer result; a sample math problem determining module configured to determine a first sample math problem and a second sample math problem from the sample math problems that the model to be trained incorrectly answers; wherein the number of correct answer steps included in the model answering steps of the first sample math problem is greater than the number of correct answer steps included in the model answering steps of the second sample math problem; A first training module is configured to train the model to be trained using a preset reinforcement learning algorithm based on the first sample math problem to obtain a pre-trained model; reward data of the preset reinforcement learning algorithm is calculated based on the weights of correct answer steps included in each model solution step; The second training module is configured to train the pre-trained model based on the second sample math problem using the preset reinforcement learning algorithm to obtain a trained model.
14. A computing device, characterized in that include: memory and processor; The memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions. When the computer programs or instructions are executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer-readable storage medium, characterized in that It stores a computer program or instruction, which implements the steps of the method according to any one of claims 1 to 12 when executed by a processor.
16. A computer program product, characterized in that The method comprises a computer program or instructions, which implements the steps of the method according to any one of claims 1 to 12 when executed by a processor.
Citation Information
Patent Citations
Mathematical problem answering model training method and device
CN116595159A
Cited By
Common sense error correction method and system, storage medium and terminal
CN121543767A