Mathematical question and answer method and device, storage medium and equipment
Through reinforcement learning methods of data augmentation and near-end strategy optimization, combined with process reward model, a more efficient and accurate mathematical question-and-answer model is built, solving the problem of poor performance of large language models in mathematical problem solving, and achieving better solution effects and user experience.
Patent Information
- Application Number
- CN202510124304.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-30
AI Technical Summary
The existing large language models perform poorly in answering mathematical problems, lack in-depth understanding of mathematical language and logic, resulting in unsatisfactory answers, high cost of relying on manual labeling of data, and poor generalization ability.
Through reinforcement learning methods and process reward models that combine data augmentation with proximity strategy optimization, a more accurate, efficient and better generalization ability mathematical question-and-answer model is built. This model obtains the initial mathematical problem text for diversification of the problem and solution process, generates a diverse sample mathematical question-and-answer pair, and uses the process reward model to supervise and fine-tune the solution process.
It improves the response efficiency and accuracy of mathematical question-and-answer models in mathematical question-and-answer models, reduces dependence on manual annotation data, enhances the generalization ability of the model, and thus improves the user's mathematical question-and-answer experience.
Smart Images

Figure CN120067253A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of natural language processing, and in particular, to a mathematical question-answering method, apparatus, storage medium, and device. Background Art
[0002] With the rapid development of information technologies such as artificial intelligence and the Internet of Things, the application scenarios of human-computer interaction are becoming more and more extensive. Various large language models (LLMs) have emerged in people's life and work, providing intelligent interaction functions for many application scenarios such as knowledge Q&A to assist users in completing various behavioral intents.
[0003] Currently, when traditional LLMs answer mathematical questions, they usually need to be trained through supervised learning, relying on a large amount of labeled data to learn the answering strategies for mathematical questions. However, the performance of the trained models in answering mathematical questions is not ideal because they lack a deep understanding of mathematical language and logic. In this regard, there are usually three existing ways to improve the answering performance of LLMs in mathematical questions: One is to enrich the training set through data augmentation, but this method depends on manually designed rules and is difficult to cover all possible types of mathematical questions, with poor generalization. The second is to improve the answering quality of the model through process verification, but this method highly depends on a large amount of manually labeled data for training implementation, with a high cost. The third is to research and attempt to improve the mathematical question-answering ability of LLMs through fine-tuning, but this method depends on specific data sets and fine-tuning strategies and may be difficult to adapt to different types of mathematical questions, which will affect the reply efficiency and accuracy of the model. It can be seen that the above-mentioned existing ways to improve the answering performance of LLMs in mathematical questions have not achieved an ideal answering effect, and there are problems such as strong dependence on manually labeled data, limited reasoning ability, and low accuracy, which will further reduce the user's mathematical Q&A experience. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to provide a mathematical question-answering method, apparatus, storage medium, and device. By combining data augmentation with process rewards, a more accurate, efficient, and better-generalized mathematical question-answering model is constructed, thereby improving the reply efficiency and accuracy for the mathematical questions raised by users, and further improving the user's mathematical Q&A experience.
[0005] The embodiments of this application provide a mathematical question-answering method, including:
[0006] Obtain the target mathematical question text to be replied;
[0007] Input the target mathematical problem text into a pre-constructed mathematical question-answering model to predict the target answer text for answering the target mathematical problem; the target answer text contains the solution process for the target mathematical problem.
[0008] Among them, the mathematical question-answering model is obtained by training an initial large language model using the proximal policy optimization-based reinforcement learning method and a process reward model, with the use of diverse sample mathematical question-answer pairs obtained after data augmentation; the process reward model is obtained by supervised fine-tuning training using the loss function based on the sample mathematical question-answer pairs and the process reward scores of the solution processes in the sample mathematical question-answer pairs.
[0009] In a possible implementation manner, the training method of the mathematical question-answering model is as follows:
[0010] Obtain the initial mathematical problem text; and perform data augmentation processing on the initial mathematical problem text for question diversification and solution process diversification to obtain diverse sample mathematical question-answer pairs.
[0011] Use the process reward model to assign correctness probability scores to each step of the solution process in the sample mathematical question-answer pairs.
[0012] Based on the proximal policy optimization-based reinforcement learning method, use the sample mathematical question-answer pairs and the correctness probability scores assigned to each step of the solution process in the sample mathematical question-answer pairs to train the initial large language model to obtain the mathematical question-answering model.
[0013] In a possible implementation manner, the data augmentation processing for question diversification and solution process diversification of the initial mathematical problem text includes:
[0014] Based on the initial mathematical problem text, construct positive mathematical problem texts and reverse mathematical problem texts to form diverse mathematical problem texts.
[0015] Generate the solution processes for the diverse mathematical problems; and use a preset correction engine to correct the solution processes, so as to generate diverse sample mathematical question-answer pairs according to the correction results; and / or, use a preset mathematical scoring model to score the solution processes, so as to generate diverse sample mathematical question-answer pairs according to the scoring results.
[0016] In a possible implementation manner, the construction method of the positive mathematical problem text is as follows:
[0017] Input the initial mathematical problem text combined with the first prompt instruction "prompt" into the large language model, and obtain the positive mathematical problem text output by the large language model through at least one of the following processing methods: changing the problem statement, changing specific numbers, introducing fractions or percentages, and introducing or changing the application scenario.
[0018] In a possible implementation, the reverse problem text is constructed as follows:
[0019] Construct the reverse problem text based on the initial mathematical problem text and the preset construction rules;
[0020] And / or, input the initial mathematical problem text combined with the second prompt instruction "prompt" into the large language model to obtain the reverse mathematical problem text output by the large language model.
[0021] In a possible implementation, the preset mathematical scoring model is the Reward model.
[0022] In a possible implementation, the loss function is the mean square error (MSE) loss function.
[0023] The embodiment of the present application also provides a mathematical question and answer device, including:
[0024] A first acquisition unit, configured to acquire the target mathematical problem text to be answered;
[0025] A prediction unit, configured to input the target mathematical problem text into a pre-constructed mathematical question and answer model, and predict the target answer text for the target mathematical problem; the target answer text includes the solution process for the target mathematical problem.
[0026] Wherein, the mathematical question and answer model is trained by using the proximal policy optimization-based reinforcement learning method and the process reward model, and using the diversified sample mathematical question and answer pairs obtained after data augmentation to train the initial large language model; the process reward model is obtained by using the diversified sample mathematical question and answer pairs obtained after data augmentation and the process reward scores of the solution processes in the sample mathematical question and answer pairs, and performing supervised fine-tuning training by using the loss function.
[0027] In a possible implementation, the device further includes:
[0028] A second acquisition unit, configured to acquire the initial mathematical problem text; and perform data augmentation processing on the initial mathematical problem text for problem diversification and solution process diversification to obtain diversified sample mathematical question and answer pairs;
[0029] An allocation unit for allocating a correctness probability score to each step of the solution process in the sample math question-answer pair by using a process reward model;
[0030] A training unit for training an initial large language model based on the proximal policy optimization reinforcement learning method, using the sample math question-answer pair and the correctness probability score allocated to each step of the solution process in the sample math question-answer pair, to obtain a math question-answer model.
[0031] In a possible implementation, the second acquisition unit includes:
[0032] A construction unit for constructing a forward math question text and a reverse math question text based on the initial math question text to form a diversified math question text;
[0033] A generation unit for generating the solution process of the diversified math question; and using a preset correction engine to correct the solution process, so as to generate diversified sample math question-answer pairs according to the correction result; and / or, using a preset math scoring model to score the solution process, so as to generate diversified sample math question-answer pairs according to the scoring result.
[0034] In a possible implementation, the construction unit is specifically used for:
[0035] Combining the initial math question text with a first prompt instruction prompt and inputting it into the large language model, and obtaining the forward math question text output by the large language model by performing at least one of the following processing methods on the initial math question text: changing the question expression, changing specific numbers, introducing fractions or percentages, and introducing or changing the application scenario.
[0036] In a possible implementation, the construction unit is specifically used for:
[0037] Constructing a reverse question text based on the initial math question text and a preset construction rule;
[0038] And / or, combining the initial math question text with a second prompt instruction prompt and inputting it into the large language model to obtain the reverse math question text output by the large language model.
[0039] In a possible implementation, the preset math scoring model is a Reward model.
[0040] In a possible implementation, the loss function is a mean squared error MSE loss function.
[0041] An embodiment of the present application further provides a math question-answer device, including: a processor, a memory, and a system bus;
[0042] The processor and the memory are connected through the system bus;
[0043] The memory is used to store one or more programs, and the one or more programs include instructions, and when the instructions are executed by the processor, the processor executes any one of the implementation manners of the above-mentioned mathematical question-answering method.
[0044] An embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored, and when the instructions run on a terminal device, the terminal device executes any one of the implementation manners of the above-mentioned mathematical question-answering method.
[0045] An embodiment of the present application further provides a computer program product, and when the computer program product runs on a terminal device, the terminal device executes any one of the implementation manners of the above-mentioned mathematical question-answering method.
[0046] A mathematical question-answering method, device, storage medium and device provided by an embodiment of the present application first obtain a target mathematical question text to be answered; then input the target mathematical question text into a pre-constructed mathematical question-answering model to predict a target answer text for answering the target mathematical question; wherein, the target answer text contains a solution process for the target mathematical question, and the mathematical question-answering model is a reinforcement learning method and a process reward model based on proximal policy optimization, and uses diversified sample mathematical question-answer pairs obtained after data augmentation to train an initial large language model; the process reward model is obtained by supervised fine-tuning training using a loss function according to the sample mathematical question-answer pairs and the process reward scores of the solution processes in the sample mathematical question-answer pairs.
[0047] It can be seen that since the present application first obtains diversified sample mathematical question-answer pairs with higher quality and wider coverage through data augmentation, and then trains an initial large language model using the sample mathematical question-answer pairs based on the reinforcement learning method and the process reward model of proximal policy optimization to generate a mathematical question-answering model, the answering accuracy and efficiency of the mathematical question-answering model are effectively improved. Therefore, when using this model to answer a target mathematical question, the answering efficiency and accuracy of the target mathematical question can be improved, and thus the user's mathematical question-answering experience is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0049] Figure 1 Schematic flowchart of a mathematical question - answering method provided by an embodiment of the present application;
[0050] Figure 2 Schematic diagram of the process of performing data augmentation processing on the initial mathematical problem text to diversify the questions and the answering process, obtaining diversified sample mathematical question - answering pairs provided by an embodiment of the present application;
[0051] Figure 3 Schematic diagram of the composition of a mathematical question - answering device provided by an embodiment of the present application. Detailed implementation manners
[0052] With the rapid development of artificial intelligence technology, large language models (LLMs) have achieved remarkable results in processing natural language. They perform well in tasks such as language understanding, text generation, and dialogue systems. However, when it comes to mathematical problems that require complex reasoning and calculation steps, the performance of these models is often unsatisfactory. Especially in the field of primary school mathematics education, students need to solve various mathematical problems to cultivate logical thinking and problem - solving abilities. When traditional LLMs answer mathematical problems, they usually need to be trained through supervised learning, relying on a large amount of labeled data to learn the answering strategies of mathematical problems. For example, GPT - series models have obtained the basic ability to process natural language through pre - training on a large amount of text data. However, the performance of these trained models in answering mathematical problems is not ideal because they lack an in - depth understanding of mathematical language and logic.
[0053] In response to this, there are usually three existing ways to improve the answering performance of LLMs on mathematical problems: The first is to enrich the training set through data augmentation. Although this method can improve the generalization ability of the model to a certain extent, it relies on manually designed rules and is difficult to cover all possible types of mathematical problems. The second is to improve the answering quality of the model through process verification. This method provides feedback by evaluating each step in the process of answering mathematical problems, thereby being able to more accurately locate errors. However, it highly depends on a large amount of manually labeled data for training implementation, which is very time - consuming and expensive in practice, with a high cost. The third is to try to improve the mathematical problem - answering ability of LLMs through fine - tuning. Although this method can improve the performance of the model to a certain extent, it depends on specific data sets and fine - tuning strategies and may be difficult to adapt to different types of mathematical problems, which will affect the response efficiency and accuracy of the model.
[0054] It can be seen that there are several main disadvantages in the above - mentioned existing ways to improve the answering performance of LLMs on mathematical problems:
[0055] 1. Strong dependence on manually labeled data and poor model generalization ability: LLMs in the existing technologies usually require a large amount of labeled data for training, which is not only costly, but also the scale of the existing labeled data limits the model's generalization ability on unseen problem types. In addition, the existing data augmentation methods mainly rely on manually designed rules, which may not cover all possible variants of mathematical problems, resulting in poor performance of the model when facing novel or complex mathematical problems.
[0056] 2. Insufficient accuracy and logic in the reasoning process: Many existing models lack step-by-step verification of the reasoning process when answering mathematical problems, which may lead to logical errors or incompleteness in the reasoning steps although the final answer is correct, and the errors in the answering process cannot be accurately located. This results in unsatisfactory accuracy and logic in the reasoning process of the model.
[0057] Therefore, how to improve the answering effect of LLMs on mathematical problems to solve the existing problems such as strong dependence on manually labeled data, limited reasoning ability, and low accuracy, and thus improve the user's mathematical Q&A experience is an urgent problem to be solved currently.
[0058] To address the above deficiencies, this application provides a mathematical Q&A method. First, obtain the target mathematical problem text to be answered; then input the target mathematical problem text into a pre-constructed mathematical Q&A model to predict the target answer text for answering the target mathematical problem; wherein, the target answer text contains the solution process for the target mathematical problem, and the mathematical Q&A model is trained from an initial large language model using a proximal policy optimization-based reinforcement learning method and a process reward model with diverse sample mathematical Q&A pairs obtained after data augmentation; the process reward model is obtained by supervised fine-tuning training using the loss function based on the sample mathematical Q&A pairs and the process reward scores of the solution processes in the sample mathematical Q&A pairs.
[0059] It can be seen that since this application first obtains higher-quality and wider-coverage diverse sample mathematical Q&A pairs through data augmentation, and then trains the initial large language model using the proximal policy optimization-based reinforcement learning method and the process reward model to generate a mathematical Q&A model, effectively improving the answering accuracy and efficiency of the mathematical Q&A model. Therefore, when using this model to answer the target mathematical problem, the answering efficiency and accuracy for the target mathematical problem can be improved, thereby enhancing the user's mathematical Q&A experience.
[0060] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0061] First Embodiment
[0062] See Figure 1 , which is a schematic flowchart of a mathematical question-answering method provided in this embodiment. The method includes the following steps:
[0063] S101: Obtain the target mathematical question text to be answered.
[0064] In this embodiment, any mathematical question text proposed by the user and to be answered using this embodiment is defined as the target mathematical question text to be answered. Moreover, this embodiment does not limit the language type of the target mathematical question text. For example, the target mathematical question text can be a Chinese text or an English text, etc. This embodiment also does not limit the length of the target mathematical question text, that is, the target mathematical question text can be a sentence text (i.e., a set of words) or a passage text (i.e., a set of sentences), etc. This embodiment also does not limit the content of the target mathematical question text. For example, the target mathematical question text can be any primary school mathematics question, etc.
[0065] Illustrative example: The content of the target mathematical question text to be answered can be: "Two cube blocks of the same size are combined into a cuboid, and the total edge length is reduced by 24 cm. What is the total edge length of these two cube blocks originally?" etc.
[0066] S102: Input the target mathematical question text into a pre-constructed mathematical question-answering model to predict the target answer text for answering the target mathematical question; wherein, the target answer text contains the solution process for the target mathematical question; the mathematical question-answering model is obtained by training an initial large language model using a reinforcement learning method based on proximal policy optimization and a process reward model, and using diversified sample mathematical question-answer pairs obtained after data augmentation; the process reward model is obtained by performing supervised fine-tuning training using a loss function based on the sample mathematical question-answer pairs and the process reward scores of the solution processes in the sample mathematical question-answer pairs.
[0067] In this embodiment, after obtaining the target math problem text to be replied through step S101, further, the target math problem text (or integrated with the prompt instruction prompt) can be input into a pre-constructed math Q&A model, so as to utilize the rich domain knowledge and sufficient semantic analysis and understanding capabilities of the label prediction model to answer the target math problem, generate an answer text for replying to the target math problem, and define it as the target answer text. The target answer text contains a detailed and accurate solution process for the target math problem, which is convenient for users to understand, so as to further improve the user's math Q&A experience.
[0068] For example: Suppose the content of the target math problem text to be replied is: "Two cube blocks of the same size are combined into a cuboid, and the total edge length is reduced by 24 cm. What is the total edge length of these two cube blocks originally?" After using the math Q&A model to answer it, the target answer text containing the solution process is:
[0069] "This question examines the properties of cubes and the splicing problem of cuboids. According to the question, after two cube blocks of the same size are combined into a cuboid, the total edge length is reduced by 24 cm. This is because during the splicing process, 8 edges overlap and are no longer included in the total edge length. Therefore, we can know that the total length of these 8 edges is the reduced 24 cm, and then we can find the length of one edge. Then, using this edge length, we calculate the total edge length of one cube, and then multiply by 2 to get the total edge length of the two cube blocks originally.
[0070] Solution: First, since after two cube blocks of the same size are combined into a cuboid, 8 edges overlap and are no longer included in the total edge length, the total length of the 8 edges is 24 / 8 = 3 (cm); then the total edge length of one cube is 3*12 = 36 (cm); so, the total edge length of the two cube blocks originally is 36*2 = 72 (cm).
[0071] Answer: The total edge length of these two cube blocks originally is 72 cm."
[0072] Among them, it should be noted that in order to improve the efficiency and accuracy of answering mathematical questions in this application, a mathematical question-answering model is pre-constructed. This model is based on the reinforcement learning method of Proximal Policy Optimization (PPO) and the Process Reward Model (PRM), and is trained on the initial large language model using diverse sample mathematical question-answer pairs obtained after data augmentation. This training method can not only enhance the logic and correctness of the model's mathematical solution process through the feedback of the Process Reward Model (PRM), but also reduce the dependence on manually annotated data, improving the efficiency and economy of model training. Thus, when applied to fields such as primary school mathematics education, it can help improve students' learning efficiency and teachers' teaching effectiveness.
[0073] Next, this embodiment will introduce the training process of the mathematical question-answering model. Among them, an optional implementation method is that the mathematical question-answering model can specifically include the following steps A-C:
[0074] Step A: Obtain the initial mathematical question text; and perform data augmentation processing on the initial mathematical question text for question diversification and answer process diversification to obtain diverse sample mathematical question-answer pairs.
[0075] In this implementation method, in order to construct the mathematical question-answering model, a large amount of preparatory work needs to be done in advance. First, a large amount of existing mathematical question text data needs to be collected as the initial mathematical question text, and then data augmentation processing for question diversification and answer process diversification is performed on the initial mathematical question text to obtain diverse sample mathematical question-answer pairs, which are used to form the model training data. For example, a large amount of primary school mathematics question text data can be collected in advance from the primary school mathematics question bank as the initial mathematical question text, and then data augmentation processing for question diversification and answer process diversification is performed on these initial mathematical question texts to obtain diverse sample mathematical question-answer pairs, forming a data-augmented question bank as the model training data, as Figure 2 shown.
[0076] Among them, an optional implementation method is that the specific implementation process of performing data augmentation processing on the initial mathematical question text for question diversification and answer process diversification can include: First, based on the initial mathematical question text, construct positive mathematical question texts and negative mathematical question texts to form diverse mathematical question texts. Then generate the answer processes for these diverse mathematical questions; and use a preset correction engine to correct the answer processes to generate diverse sample mathematical question-answer pairs according to the correction results; and / or use a preset mathematical scoring model to score the answer processes to generate diverse sample mathematical question-answer pairs according to the scoring results.
[0077] In this implementation, different from traditional problem enhancement strategies, this application maximizes the diversity of problems from both positive and negative aspects. For the construction of positive math problem texts, this application greatly enriches the initial math problems through the idea of drawing inferences from one instance. The specific construction method is as follows: Combine the initial math problem text with the first prompt and input it into a large language model LLM (such as GPT-3.5 and GPT-4, etc.). By performing at least one of the following processing methods on the initial math problem text, namely, "changing the problem statement", "changing specific numbers", "introducing fractions or percentages", "introducing or changing application scenarios", obtain the positive math problem text output by the large language model LLM.
[0078] Among them, the large language model LLM can be a language model based on deep learning, which can generate new language expressions according to the input text content, such as texts, sentences, paragraphs, or even articles, etc. The large language model LLM is trained using a large-scale language dataset through an autoregressive generation method to obtain language rules and patterns, and can simulate human instructions to generate language expressions (such as text data). Specifically, when the large language model LLM generates new text data, it predicts the probability of the next language unit based on the previously generated content until the complete text data is generated.
[0079] Illustrative example: The example of the first prompt input to the large language model LLM is as follows:
[0080] "You are an AI assistant that helps me rewrite problems. Please operate according to the examples given below:
[0081] Initial math problem: Xiaohong has 15 chocolates, and her sister has 30. If they ate 35, how many chocolates do they have left in total?
[0082] Changing the problem statement: Xiaohong and her sister have 15 and 30 chocolates respectively. After they ate 35 chocolates together, how many are left?
[0083] Changing specific numbers: Xiaohong has 20 chocolates, and her sister has 40. If they ate 25, how many chocolates do they have left in total?
[0084] Introducing fractions or percentages: Xiaohong has 15 chocolates, and her sister has 30. If they ate 7 / 9 of the total, how many chocolates do they have left in total?
[0085] Introducing or changing the application scenario: Xiaohong has 15 books, and her sister has 30. If they lent out 35, how many books do they have left in total?
[0086] Initial math problem: A has 20 lollipops. He gives some lollipops to B. Now A has 5 lollipops left. How many lollipops did A give to B?
[0087] Change the problem statement: If A initially had 20 lollipops and now has 5 left after giving some to B, how many did he give to B?
[0088] Change specific numbers: A has 12 lollipops. He gives some lollipops to B. Now A has 5 lollipops left. How many lollipops did A give to B?
[0089] Introduce fractions or percentages: A has 20 lollipops. He gives some lollipops to B. Now A has 25% of the total left. How many lollipops did A give to B?
[0090] Introduce or change the application scenario: A has 20 basketballs. He sells some basketballs to B. Now A has 5 basketballs left. How many basketballs did A sell to B?
[0091] Initial math problem: {question}”.
[0092] Among them, when the initial math problem text is “A set of 'Encyclopedia' costs 30 yuan, and two sets cost 54 yuan. I have 300 yuan. How many sets can I buy at most?”, the positive math problem text example output by the large language model is as follows:
[0093] “Change the problem statement: The price of a set of 'Encyclopedia' is 30 yuan, and the price of two sets is 54 yuan. If I have 300 yuan, how many sets can I buy at most?
[0094] Change specific numbers: A set of 'Encyclopedia' costs 36 yuan, and two sets cost 65 yuan. I have 350 yuan. How many sets can I buy at most?
[0095] Introduce fractions or percentages: A set of 'Encyclopedia' costs 30 yuan, and there is a 10% discount for buying two sets. I have 300 yuan. How many sets can I buy at most?
[0096] Introduce or change the application scenario: A set of stationery costs 30 yuan, and two sets cost 54 yuan. Xiaoming has 300 yuan. How many sets can he buy at most?”.
[0097] For the construction of reverse math problem texts, this application is implemented in two ways: construction based on rules and generation based on models. Among them, the method of construction based on rules is: based on the initial math problem text and preset construction rules, construct the reverse problem text. For example, use the identifier "x" to randomly mask a known variable in the initial math problem, and provide the answer in the original initial math problem, and require the masked variable to be solved, so as to form a new reverse reasoning problem that infers the known variable from the original answer. However, there is a phenomenon of a single form of asking questions in the reverse problems constructed based on rules. Therefore, in order to diversify the reverse problems, the reverse problem text can also be generated based on models. The specific construction method is: combine the initial math problem text with the second prompt instruction and input it into the large language model to obtain the reverse math problem text output by the large language model.
[0098] Illustrative example: The examples of the second prompt instruction prompt input to the large language model LLM are as follows:
[0099] "You are an AI assistant that helps me rewrite math problems. Rewriting rules: Given a problem and its answer, add the answer to the problem as a known quantity, and then randomly change one of the known quantities in the original problem to the unknown quantity to be solved, form a new problem and output the new answer. The following examples can be referred to:
[0100] Initial math problem: 20 young pioneers harvested 160 kilograms of apples. If each basket can hold 20 kilograms, 2 baskets are still lacking. How many baskets were there originally?
[0101] Answer: There were 6 baskets originally.
[0102] New problem: 20 young pioneers harvested some apples. If they are packed in 6 baskets, with 20 kilograms in each basket, 2 baskets are still lacking. How many kilograms of apples are there in total?
[0103] New answer: 160 kilograms.
[0104] Initial math problem: There are 100 monks eating 100 buns. One big monk eats 4 buns, and 4 little monks eat 1 bun. Find out how many big and little monks there are.
[0105] Answer: There are 20 big monks and 80 little monks.
[0106] New problem: There are 20 big monks and 80 little monks eating buns together. One big monk eats 4 buns, and 4 little monks eat 1 bun. Find out how many buns the big and little monks eat in total.
[0107] New answer: 100 buns.
[0108] Initial math problem: {question}”.
[0109] Among them, when the initial math problem is "20 young pioneers harvested 160 kilograms of apples. If each basket can hold 20 kilograms, 2 baskets are still lacking. How many baskets were there originally?", and the corresponding answer is "There were originally 6 baskets", the reverse math problem text example output by the large language model can be: "New problem: 20 young pioneers harvested 160 kilograms of apples. If they are packed in 6 baskets, 2 baskets are still lacking. How many kilograms can each basket hold?", and the corresponding new answer is "Each basket can hold 20 kilograms".
[0110] On this basis, in order to further enhance the training dataset, this embodiment also generates the solution processes of diverse math problems; and uses a preset correction engine to correct the solution processes, so as to generate diverse sample math question-answer pairs according to the correction results; and / or, uses a preset math scoring model (such as the Reward model, etc.) to score the solution processes, so as to generate diverse sample math question-answer pairs for performing the subsequent step B.
[0111] Specifically, since generating more reasoning paths is a simple but effective method to enhance the training set. Therefore, after forming diverse math problem texts using the forward math problem texts and reverse math problem texts, further, a Spark large model trained by supervised fine-tuning (SFT) based on the original primary school math question bank, etc., can be used to generate the solution processes of each diverse math problem. And the generated solution processes also need to go through a data screening strategy to screen out correct data as much as possible. For example, for questions with original answers, a preset correction engine (such as the correction engine of iFlytek) can be used to correct the solution processes and screen out the solution processes that are consistent with the original answers; for questions with unknown answers, a preset math scoring model (such as the Reward model) can be used to score the solution processes and screen out the data with higher scores to generate diverse sample math question-answer pairs, such as Figure 2 shown.
[0112] Among them, the function of the preset math scoring model (such as the Reward model) is to give a math problem and the solution process of this problem, and the model scores the correctness of this process. The score range is [-5, 5]. The larger the score, the greater the probability of being correct. The specific construction will not be elaborated here.
[0113] It can be seen that this application uses an automated data augmentation strategy. Based on the initial math problem text, from three aspects: forward math problem text construction, reverse math problem text construction, and diversification of solution processes, it automatically generates diverse math problems and solution processes, reduces the dependence on manually labeled data, and at the same time improves the coverage and diversity of the training data.
[0114] Step B: Use the process reward model to assign a correctness probability score to each step in the solution process of the sample math question-answer pair.
[0115] In this implementation, after performing data augmentation processing on the initial math problem text for question diversification and solution process diversification through Step A to obtain diversified sample math question-answer pairs, further, the process reward model can be used to assign a correctness probability score to each step in the solution process of the sample math question-answer pair, so as to perform subsequent Step C, train the math question-answer model, and ensure the accuracy and logic of the model-generated solution process.
[0116] It should be noted that in a complex math reasoning process, if a certain step in the problem-solving process contains an error, the final result is more likely to be incorrect. The process reward model construction strategy aims to distribute reward scores to each step in the math problem-solving process. The main benefit is that it can provide precise feedback to the model by identifying the specific location where any possible error may occur, which is a valuable signal for reinforcement learning and automatic correction in subsequent Step C. However, currently, collecting the reward scores for each step in the solution process is a difficult process that heavily relies on manual annotation. Therefore, this application proposes an automatic process reward score construction strategy, constructs a process reward model, and breaks the bottleneck of relying heavily on manual annotation in existing methods.
[0117] Specifically, first, when the problem to be solved is the probability of a certain event occurring or the expected value of a certain random variable, they can be obtained by a certain "trial" method to get the frequency of this event occurring or the average value of this random variable, and use them as the solution to the problem. This application defines the quality of each step in the solution process as the potential to infer the correct answer from this step, that is, if an inference step can infer more correct inference paths than other inference steps, then it will be assigned a higher correctness score.
[0118] Then, to quantify the potential of each step s i in the solution process, this application defines a process reward score for each step s i and represents it as SE. Then, use the thought chain prompting idea to let the large language model generate N (where N is a positive integer greater than 1) inference processes from step s i and represent them as: where, a j and Ki respectively represent the final answer of the jth inference process and the total number of steps in the jth inference process. Then, construct the process reward score according to the correctness of all answers of these N inference processes. Among them, the specific calculation formula of the process reward score SE is as follows:
[0119]
[0120] Among them, a* represents the standard answer (a question may have answers to multiple sub-questions, and a* represents the set of standard answers to these sub-questions); Π(a j = a*) means comparing the final answer with the standard answers to multiple sub-questions. If each sub-question is correct, it is 1, and if it is wrong, it is 0. Then multiply the comparison results of multiple sub-questions. Only when all sub-questions are correct, the result is 1.
[0121] The above formula indicates that step s i The process reward score SE refers to the probability of successfully obtaining the correct reasoning path starting from step s i That is, the potential of this step to infer the correct answer.
[0122] In this way, after constructing the process reward score for each step in the solution process of the sample math question and answer through the above method, the process reward model PRM can be supervised and fine-tuned during the training process according to the sample math question and answer pair and the process reward score of the solution process in the sample math question and answer pair. Among them, the specific content of the loss function in this application is not limited, and it can be selected according to the actual situation and experience. For example, it can be set as the most commonly used mean square error (MSE) loss function in the regression loss function. The specific calculation formula is as follows:
[0123]
[0124] Among them, represents the process reward score of step s i ; represents the score assigned by the process reward model to step s i ; K represents the number of steps in this reasoning process. By training the process reward model, precise feedback can be provided for specific positions where errors occur in the mathematical reasoning process, providing a reference for the reinforcement learning algorithm and model automatic correction in subsequent steps for proximal policy optimization.
[0125] It should be noted that the specific composition structure and training process of the process reward model in this application are not restricted. An optional implementation method is to use the Spark large model base as the basic model. From the questions and solution processes in the diversified sample math question and answer pairs obtained after data augmentation, use the calculation formula of the process reward score SE to assign scores to each step in the solution process, and then form SFT data in the following format to perform SFT training on the Spark large model base. Among them, an example of SFT data is:
[0126] {
[0127] 'input': 'Suppose you are a professional math teacher and grade each step of the solution process according to the given problem.
[0128]
Problem
[0129]
Solution process
[0130] Solution: The total length of 8 edges is 24÷8 = 3 (cm); then the total edge length of one cube is 3×12 = 36 (cm); so, the total edge length of the two original cube blocks is 36×2 = 72 (cm).',
[0131] 'target': ['0.6', '0.8', '0.8', '0.9']
[0132] }
[0133] The above SFT data example can be used as the input data example, and the output result example of the model can be: 'predict': ['0.7', '0.6', '0.8', '0.8'], which is used to compare with the target result, and the model is trained according to the comparison result until the preset conditions are met, then the update of the model parameters is stopped to obtain the process reward model.
[0134] It can be seen that in the construction strategy of the process reward model, this application proposes an automated process reward model construction strategy, defines a novel process reward score calculation method, calculates the reward score for each step in the solution process of math problems, provides more accurate error localization and feedback for reinforcement learning training, so as to improve the correctness and logic of the trained model in the math solution process.
[0135] Step C: Based on the proximal policy optimization reinforcement learning method, use the sample math Q&A pairs and assign a correctness probability score to each step in the solution process of the sample math Q&A pairs to train the initial large language model to obtain a math Q&A model.
[0136] In this implementation, after implementing the process reward model through step B, further, based on the proximal policy optimization reinforcement learning method, use the sample math Q&A pairs and assign a correctness probability score to each step in the solution process of the sample math Q&A pairs to train the initial large language model to obtain a math Q&A model.
[0137] Among them, the Proximal Policy Optimization (PPO) method is different from the traditional training strategy that uses the PRM and the Outcome Reward Model (ORM). The ORM result reward model strategy only provides rewards at the end of the inference. The Process Reward Model (PRM) can provide rewards at the end of each inference step. The PPO training process can be roughly divided into: (1) Initialization: First, it is necessary to initialize the policy network (Actor) and the value network (Critic). The Actor is the model that generates the solution steps, and the Critic is used to estimate the advantage of the current solution step. (2) Collect scores: Input a question, the Actor generates the solution process, and the PRM generates scores, that is, rewards, for each solution step. (3) Calculate the advantage function: For the scores of the collected solution steps, combine the Critic to calculate the advantage corresponding to a certain solution step generated by the model, which is used to evaluate the quality of the model generating this step. (4) Update the policy: According to the advantage function, judge the quality of the model generating the next step, so as to suppress the generation probability of bad steps and increase the generation probability of good steps. When training the model, the input of the large model optimized by PPO is the question text, and the output is the text containing the solution process of the question.
[0138] When the Process Reward Model (PRM) is combined with the PPO algorithm for training, the objective function can be expressed as a combination of the expected reward and the policy probability. The form of the PPO objective function is as follows:
[0139]
[0140] Among them, represents the action probability ratio of the old and new policies in the same state. Among them, the old and new policies refer to the answer models before and after the update of the reinforcement learning algorithm PPO, and the action probability ratio refers to the probability ratio of generating a certain solution step; Ε t represents the mathematical expectation, which refers to the expected value of the future reward for generating the next solution step starting from a certain solution step s; represents the estimated value of the advantage function, which is an estimated value calculated by combining the Process Reward Model (PRM) and the value network (Critic); clip(r t (θ), 1 - ε, 1 + ε) represents the clipping function that limits the probability ratio within the range of [1 - ε, 1 + ε]. The Process Reward Model (PRM) plays a key role in the PPO algorithm, helping the model better evaluate and optimize its policy during the training process, thereby improving the accuracy of answering questions.
[0141] In summary, for the mathematical question-answering method provided in this embodiment, first, the target mathematical question text to be answered is obtained; then, the target mathematical question text is input into a pre-constructed mathematical question-answering model, and the target answer text for answering the target mathematical question is predicted; wherein, the target answer text contains the solution process for the target mathematical question, and the mathematical question-answering model is obtained by training an initial large language model using a diverse set of sample mathematical question-answer pairs obtained through data augmentation based on the proximal policy optimization-based reinforcement learning method and the process reward model; the process reward model is obtained by performing supervised fine-tuning training using the loss function based on the sample mathematical question-answer pairs and the process reward scores of the solution processes in the sample mathematical question-answer pairs.
[0142] It can be seen that since in this application, first, a diverse set of sample mathematical question-answer pairs with higher quality and wider coverage are obtained through data augmentation, and then based on the proximal policy optimization-based reinforcement learning method and the process reward model, the initial large language model is trained using the sample mathematical question-answer pairs to generate a mathematical question-answering model, effectively improving the response accuracy and efficiency of the mathematical question-answering model. Therefore, when using this model to answer the target mathematical question, the response efficiency and accuracy for the target mathematical question can be improved, thereby enhancing the user's mathematical question-answering experience.
[0143] Second Embodiment
[0144] This embodiment will introduce a mathematical question-answering device. For related content, please refer to the above method embodiment.
[0145] See Figure 3 , which is a schematic diagram of the composition of a mathematical question-answering device provided in this embodiment. The device 300 includes:
[0146] A first acquisition unit 301, configured to acquire the target mathematical question text to be answered;
[0147] A prediction unit 302, configured to input the target mathematical question text into a pre-constructed mathematical question-answering model to predict the target answer text for the target mathematical question; the target answer text contains the solution process for the target mathematical question;
[0148] Wherein, the mathematical question-answering model is obtained by training an initial large language model using a diverse set of sample mathematical question-answer pairs obtained through data augmentation based on the proximal policy optimization-based reinforcement learning method and the process reward model; the process reward model is obtained by performing supervised fine-tuning training using the loss function based on the sample mathematical question-answer pairs obtained through data augmentation and the process reward scores of the solution processes in the sample mathematical question-answer pairs.
[0149] In one implementation manner of this embodiment, the device further includes:
[0150] A second acquisition unit, configured to acquire an initial mathematical problem text; and perform data augmentation processing on the initial mathematical problem text for problem diversification and answer process diversification, to obtain diversified sample mathematical Q&A pairs.
[0151] An allocation unit, configured to use a process reward model to allocate a correctness probability score to each step of the answer process in the sample mathematical Q&A pairs.
[0152] A training unit, configured to perform training on an initial large language model based on the proximal policy optimization reinforcement learning method, using the sample mathematical Q&A pairs and the correctness probability scores allocated to each step of the answer process in the sample mathematical Q&A pairs, to obtain a mathematical Q&A model.
[0153] In an implementation manner of this embodiment, the second acquisition unit includes:
[0154] A construction unit, configured to construct a forward mathematical problem text and a reverse mathematical problem text based on the initial mathematical problem text, to form diversified mathematical problem texts.
[0155] A generation unit, configured to generate an answer process for the diversified mathematical problems; and use a preset correction engine to correct the answer process, so as to generate diversified sample mathematical Q&A pairs according to the correction results; and / or use a preset mathematical scoring model to score the answer process, so as to generate diversified sample mathematical Q&A pairs according to the scoring results.
[0156] In an implementation manner of this embodiment, the construction unit is specifically configured to:
[0157] Combine the initial mathematical problem text with a first prompt instruction prompt, and input it into a large language model, and obtain a forward mathematical problem text output by the large language model through at least one of the following processing methods: changing the problem statement, changing specific numbers, introducing fractions or percentages, and introducing or changing the application scenario for the initial mathematical problem text.
[0158] In an implementation manner of this embodiment, the construction unit is specifically configured to:
[0159] Construct a reverse problem text based on the initial mathematical problem text and a preset construction rule.
[0160] And / or, combine the initial mathematical problem text with a second prompt instruction prompt, and input it into a large language model, to obtain a reverse mathematical problem text output by the large language model.
[0161] In an implementation manner of this embodiment, the preset mathematical scoring model is a Reward model.
[0162] In one implementation of this embodiment, the loss function is the mean squared error (MSE) loss function.
[0163] Furthermore, an embodiment of the present application also provides a mathematical question-answering device, including: a processor, a memory, and a system bus;
[0164] The processor and the memory are connected through the system bus;
[0165] The memory is used to store one or more programs, and the one or more programs include instructions that, when executed by the processor, cause the processor to execute any implementation method of the above-mentioned mathematical question-answering method.
[0166] Furthermore, an embodiment of the present application also provides a computer-readable storage medium, in which instructions are stored. When the instructions are run on a terminal device, the terminal device is caused to execute any implementation method of the above-mentioned mathematical question-answering method.
[0167] Furthermore, an embodiment of the present application also provides a computer program product that, when run on a terminal device, causes the terminal device to execute any implementation method of the above-mentioned mathematical question-answering method.
[0168] From the description of the above embodiments, those skilled in the art can clearly understand that all or part of the steps in the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present application.
[0169] It should be noted that the various embodiments in this specification are described in a progressive manner, and the key point of each embodiment is to describe the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0170] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0171] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A mathematical question-answering method, characterized in that: include: Obtain the target math problem text to be answered; Inputting the target math question text into a pre-built math question-answering model to predict a target answer text that answers the target math question; The target answer text includes a solution process for the target math problem; The mathematical question-answering model is obtained by training the initial large language model using a reinforcement learning method and a process reward model based on proximal strategy optimization and using a diverse sample mathematical question-answering pair obtained after data enhancement; The process reward model is obtained by performing supervised fine-tuning training using a loss function based on sample math question-answer pairs and process reward scores of the solution process in the sample math question-answer pairs.
2. The method according to claim 1, characterized in that: The training method of the mathematical question answering model is as follows: Acquire an initial math question text; and perform data enhancement processing on the initial math question text to diversify the questions and the answer process, so as to obtain diversified sample math question-answer pairs; Using a process reward model, assigning a correctness probability score to each step of the answering process in the sample math question-answer pair; A reinforcement learning method based on proximal strategy optimization uses the sample math question-answer pairs and each step of the answer process in the sample math question-answer pairs to assign a correctness probability score, trains the initial large language model, and obtains a math question-answering model.
3. The method according to claim 2, characterized in that The data enhancement processing of diversifying the questions and the answering process of the initial math problem text includes: Based on the initial math problem text, construct a forward math problem text and a reverse math problem text to form a diversified math problem text; Generate a solution process for the diversified mathematical problems; and use a preset grading engine to grade the solution process to generate diversified sample mathematical question-answer pairs based on the grading results; and / or, use a preset mathematical scoring model to score the solution process to generate diversified sample mathematical question-answer pairs based on the scoring results.
4. The method according to claim 3, characterized in that The positive math problem text is constructed as follows: The initial math problem text is combined with a first prompt instruction prompt and input into a large language model, and a positive math problem text output by the large language model is obtained by changing the problem statement, changing specific numbers, introducing fractions or percentages, and introducing or changing at least one processing method in an application scenario.
5. The method according to claim 3, characterized in that: The reverse question text is constructed as follows: Based on the initial math problem text and preset construction rules, construct a reverse problem text; And / or, the initial math problem text is combined with a second prompt instruction prompt, and input into a large language model to obtain a reverse math problem text output by the large language model.
6. The method according to claim 3, characterized in that The preset mathematical scoring model is the Reward model.
7. The method according to any one of claims 1 to 6, characterized in that: The loss function is the mean square error (MSE) loss function.
8. A mathematical question-answering device, characterized in that: include: A first acquisition unit is used to acquire the target math question text to be answered; A prediction unit, configured to input the target math question text into a pre-built math question-answering model, and predict a target answer text for the target math question; The target answer text includes a solution process for the target math problem; The mathematical question-answering model is obtained by training the initial large language model using a reinforcement learning method and a process reward model based on proximal strategy optimization and using a diverse sample mathematical question-answering pair obtained after data enhancement; The process reward model is obtained by supervised fine-tuning training using a loss function based on the diversified sample math question-answer pairs obtained after data enhancement and the process reward scores of the solution process in the sample math question-answer pairs.
9. A mathematical question-answering device, characterized in that: include: Processor, memory, system bus; The processor and the memory are connected via the system bus; The memory is used to store one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by the processor, the processor executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device executes the method according to any one of claims 1 to 7.
Citation Information
Cited By
Process reward model training method and system
CN120430424A
A process reward model training method and system
CN120430424B