Step correction large model training method, homework correction method, device and system
By using the large model trained by the field training data and the whole question scoring label, the probability label of the step correction results is gradually sampled and estimated, and reinforcement learning is carried out, the problem of step-level correction cannot be achieved in the existing technology, and the efficiency of homework correction is improved.
Patent Information
- Application Number
- CN202510032098.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-09
AI Technical Summary
Existing intelligent education correction products cannot achieve step-level correction, and high-quality step-level correction data is scarce, which limits the efficiency of large-scale development and operation correction.
By obtaining the large model trained by the field training data and using the whole question scoring label for training, the probability label of the step correction results is gradually sampled and estimated, and reinforcement learning is carried out to improve the step correction ability.
In the case of scarce step-level correction labeling data, efficient training of step-level correction large models is achieved, which improves the efficiency of homework correction and can assist in manual correction.
Smart Images

Figure CN119416858B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing technology, and more specifically, to a step-correction large model training method, homework correction method, device and system. Background Art
[0002] Currently, common intelligent education grading products on the market, such as learning machines, usually have a certain ability to grade entire questions, but do not have the ability to correct specific error details of students' answers, that is, they cannot achieve step-level correction of user answers.
[0003] With the development of big model technology, it has become possible to use the understanding ability of big models for intelligent grading. However, the step-level grading ability requires the use of high-quality step-level annotated data to fine-tune the big model. In reality, the whole question grading and scoring data can usually be collected in many educational grading scenarios (such as the college entrance examination, monthly exams of various schools, etc.), but high-quality step-level grading data is often scarce and the acquisition cost is high. This also limits the development of big models that can achieve step-level grading capabilities, and then it is necessary to rely on manual grading to grade student assignments, which is inefficient. Summary of the invention
[0004] In view of the above problems, this application is proposed to provide a step correction large model training method, job correction method, device and system, so as to complete the training of the step correction large model in the case of scarcity of step-level correction annotation data, ensure that it has a certain step correction capability, and then assist manual job correction to improve work efficiency. The specific plan is as follows:
[0005] In a first aspect, a step-by-step correction large model training method is provided, comprising:
[0006] Obtaining a large model trained with domain training data, wherein the domain training data at least includes question answering data, and the question answering data includes questions, standard answers, and user answers;
[0007] Acquire first training data, wherein the first training data at least includes the question answering data and the whole question scoring label of the user's answer;
[0008] Using the large model as the initial step correction large model, sampling the output of the step correction large model step by step for the user's answer in the first training data, and estimating the probability label of each step correction result being accurate based at least on the sampling results and the whole question score label of the user's answer;
[0009] The step correction large model is trained using the first training data and the estimated accurate probability label of each step correction result to obtain a trained step correction large model.
[0010] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the process of sampling the output of the step correction large model step by step, and estimating the accurate probability label of each step correction result based on at least the sampling result and the whole question score label answered by the user, includes:
[0011] Based on the Monte Carlo tree search method, the output of the step correction model is sampled step by step, and based on at least the sampling results and the whole question score label answered by the user, the accurate probability label of each step correction result is estimated.
[0012] In a possible design, in another implementation of the first aspect of the embodiment of the present application, for the user answer in the first training data, based on the Monte Carlo tree search method, sampling the output of the step correction large model step by step, and estimating the accurate probability label of each step correction result based on at least the sampling result and the whole question score label of the user answer, the process includes:
[0013] Traverse each step in the user's answer, and for the current t-th step traversed:
[0014] Sampling is performed according to the configured sampling times. The grading results of the first t steps are fixed each time sampling is performed, and the step-grading large model is used to continue predicting the grading results of the subsequent steps and the score of the entire question to obtain the sampling results;
[0015] Based at least on the sampling result and the whole question score label answered by the user, the Monte Carlo score corresponding to each step correction result is calculated, and the Monte Carlo score is used as the accurate probability label of each step correction result.
[0016] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the process of calculating the Monte Carlo score corresponding to each step correction result based on at least the sampling result and the whole question scoring label answered by the user includes:
[0017] For the current t-th step, based on the whole question score in the sampled result and the whole question score label answered by the user, a result reward score is calculated;
[0018] A Monte Carlo score corresponding to the correction result of the current step t is determined based at least on the result reward score.
[0019] In a possible design, in another implementation of the first aspect of the embodiment of the present application, at least part of the user's answers in the first training data are also annotated with step correction results;
[0020] The process of calculating the Monte Carlo score corresponding to each step correction result based at least on the sampling result and the whole question scoring label answered by the user also includes:
[0021] For the current t-th step, based on the correction results of each step after the current t-th step in the sampling result and the step correction results marked by each step, the process reward score is calculated;
[0022] A total reward score is determined based on the result reward score and the process reward score, and the total reward score is used as the Monte Carlo score corresponding to the correction result of the current t-th step.
[0023] In a possible design, in another implementation of the first aspect of the embodiment of the present application, for the current t-th step, based on the correction results of each step after the current t-th step in the sampling result and the step correction results of each step annotation, the process of calculating the process reward score includes:
[0024] Determine the proportion of step correction results that are consistent with the step correction results marked by the corresponding step in the correction results of each step after the current t-th step in the sampling result, and use the proportion as the process reward score.
[0025] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the process of training the step correction big model by using the first training data and the estimated accurate probability label of each step correction result to obtain the trained step correction big model includes:
[0026] By using the first training data and the estimated accurate probability label of each step correction result, reinforcement learning is used to constrain the step correction process of the step correction large model to obtain the step correction large model after reinforcement learning.
[0027] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the step correction process of the step correction big model is constrained by reinforcement learning to obtain the step correction big model after reinforcement learning, including:
[0028] The initial step correction big model is used as the policy network Actor, and is combined with the evaluation network Critic to perform reinforcement learning to obtain the step correction big model after reinforcement learning. The policy network is used to predict the correction result of each step in the user's answer and the score of the whole question according to the input question answer data, and the evaluation network is used to evaluate the probability of the accuracy of each step correction result output by the policy network;
[0029] In the reinforcement learning process, the evaluation network is trained according to a first goal, and the policy network is trained according to a second goal. The first goal includes minimizing the distance between the probability that the correction result of each step output by the evaluation network is accurate and the probability label of the accurate correction result of each step, and the second goal includes maximizing the probability that the correction result of each step is accurate.
[0030] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the reinforcement learning process includes two training stages. The first training stage trains the evaluation network alone according to the first objective, and the second training stage jointly trains the evaluation network and the policy network. The joint training process trains the evaluation network according to the first objective and trains the policy network according to the second objective.
[0031] In a possible design, in another implementation of the first aspect of the embodiments of the present application, the second objective further includes:
[0032] Maximize the total reward score, which includes the reward for correctly grading the entire question and the reward for accurately grading each step.
[0033] In a possible design, in another implementation of the first aspect of the embodiments of the present application, the second objective further includes:
[0034] The distance between the output of the policy network and the output of the initial step-corrected large model is constrained by the KL divergence.
[0035] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the domain training data includes: second training data and third training data, the second training data includes a plurality of question answer data, the third training data includes the question answer data, the step correction results corresponding to the marked user answers, and the whole question score label;
[0036] The process of obtaining a large model trained with domain training data includes:
[0037] Pre-training the base large model using the second training data to obtain a pre-trained large model;
[0038] The pre-trained large model is fine-tuned in a supervised manner using the third training data to obtain a fine-tuned large model.
[0039] In a second aspect, a method for marking homework is provided, comprising:
[0040] Acquire question answer data, wherein the question answer data includes user answers to be corrected, questions and standard answers;
[0041] The question answer data is sent to the step correction model to obtain the correction result of each step in the user's answer and the whole question score of the user's answer output by the model;
[0042] Among them, the step correction large model is a model trained using the step correction large model training method described in any one of the first aspects above.
[0043] In a third aspect, a step-correction large model training device is provided, comprising:
[0044] A large model acquisition unit, used to acquire a large model trained with domain training data, wherein the domain training data at least includes question answering data, and the question answering data includes questions, standard answers, and user answers;
[0045] A first training data acquisition unit, configured to acquire first training data, wherein the first training data at least includes the question answering data and a whole question scoring label of the user's answer;
[0046] a label estimation unit, configured to use the large model as an initial step correction large model, sample the output of the step correction large model step by step for the user's answer in the first training data, and estimate the probability label of each step correction result based on at least the sampling result and the whole question score label of the user's answer;
[0047] The training unit is used to train the step correction large model by using the first training data and the estimated accurate probability label of each step correction result to obtain the trained step correction large model.
[0048] In a fourth aspect, an electronic device is provided, comprising: a memory and a processor;
[0049] The memory is used to store programs;
[0050] The processor is used to execute the program to implement the various steps of the step-correcting large model training method described in any one of the first aspects of the present application, or to implement the various steps of the homework correction method described in any one of the second aspects of the present application.
[0051] In a fifth aspect, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the step-grading large model training method described in any one of the first aspects of the present application are implemented, or the steps of the homework grading method described in any one of the second aspects of the present application are implemented.
[0052] In a sixth aspect, a computer program product is provided, comprising a computer program. When the computer program is executed by a processor, the computer program implements the various steps of the step-grading large model training method described in any one of the first aspects of the present application, or implements the various steps of the homework grading method described in any one of the second aspects of the present application.
[0053] By means of the above technical solution, the present application first obtains a large model trained with domain training data as the initial step correction large model, so that the large model has certain basic subject knowledge and problem-solving ability. On this basis, the first training data for the next training can be obtained. The first training data at least includes question answering data and the whole-question scoring labels of the annotated user answers. The first training data is easier to obtain than a large number of step-level annotated data. In order to train the step correction large model, the present application samples the output of the step correction large model step by step for the user answers in the training data, and estimates the accurate probability label of each step correction result based on at least the sampling results and the whole-question scoring labels of the user answers. In this way, there is no need to manually annotate the step-level correction results in large quantities, reducing the cost of obtaining the annotated data. On this basis, the step correction large model is trained using the first training data and the accurate probability labels of each step correction result obtained. The method of the present application allows for efficient use of all training data to train the large step correction model when the training data is unbalanced (i.e., the step correction results may not be marked in the first training data, or some users may have answers marked with step correction results, while some users may not have answers marked with step correction results), thereby achieving the effect of taking into account both step correction capabilities and scoring capabilities. The trained large step correction model can then be used to assist in correcting student assignments, thereby improving the efficiency of homework correction. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0055] Figure 1 A flowchart of a method for correcting a large model training step provided in an embodiment of the present application;
[0056] Figure 2 A flowchart of a homework correction method disclosed in an embodiment of the present application;
[0057] Figure 3 A schematic diagram of the structure of a large model training device for step correction provided in an embodiment of the present application;
[0058] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] Before introducing this application solution, the relevant concepts involved in this article are first explained:
[0060] Large models: In the field of artificial intelligence, large models usually refer to large-scale pre-trained models. Such models are called "large" because they can be pre-trained on a large amount of data and can be transferred to a variety of downstream tasks. Their full English name is Large Pre-Trained Models or Large-Scale Pre-Training Models. Large models are characterized by their large scale and contain billions or even more parameters, which help them learn complex patterns in the data. The emerging capabilities of large models include but are not limited to: contextual learning, instruction following, code generation, step-by-step reasoning capabilities, etc. Large models can include large language models (LLMs) and multimodal large models. Large language models are mainly used to process text modal data. Multimodal large models further integrate multimodal capabilities on the basis of large language models and can process information in multiple modalities, such as images, text, audio, etc.
[0061] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0062] The present application provides a method for training a large model for step correction and a method for using the trained large model for step correction to assist in homework correction. The method provided in the present application can be divided into a training phase and a reasoning phase. The training phase is the phase of training the large model for step correction, and the reasoning phase is the process of using the trained large model for step correction to assist in homework correction. Among them, the training phase and the reasoning phase can be deployed in the same device or in different devices. For example, the training phase can be deployed in the cloud or in a server, and the reasoning phase can be deployed in an intelligent terminal, such as a mobile phone, a tablet, a learning machine, etc., or the reasoning phase can be completed by the cooperation of an intelligent terminal and a server.
[0063] The step correction big model of the present application can perform step-level correction on user answers and score the whole question. It is suitable for correcting user answers in a variety of scenarios, such as the correction scenario of student examination papers, especially questions in which students' answers contain one or more steps, such as answers to math questions, answers to physics questions, etc. By providing questions, standard answers and user answers, the step correction big model can be used to automatically correct user answers, and output step-level correction results and whole-question scores.
[0064] For ease of understanding, this application introduces the processes of the training phase and the reasoning phase respectively.
[0065] 1. Training Phase
[0066] Reference Figure 1 The step-by-step correction large model training method provided in this embodiment specifically includes the following steps:
[0067] Step S100: obtaining a large model trained with domain training data, wherein the domain training data at least includes question answering data, and the question answering data includes questions, standard answers and user answers.
[0068] In this embodiment, when correcting the big model in the training step, it is necessary to train on the basis of an initial big model. The initial big model can adopt a base big model (such as a general big model for example). Of course, in this embodiment, a big model trained with domain training data is further used as the initial big model.
[0069] According to this step, the domain correction task to be applied by the large model is corrected. The domain training data can be the training data under the corresponding domain correction task. In order to enable the large model to have basic domain knowledge and problem-solving ability, the domain training data can at least include question answering data. , where x represents the question, the standard answer, and the user's answer.
[0070] Domain training data can be obtained through the Internet or domain corpus. By using domain training data to train the big model, the big model can acquire domain knowledge such as grading and problem solving, and at the same time be familiar with the data format of questions and standard answers, and have certain logical reasoning ability in the field of education grading.
[0071] Taking mathematics as an example, the following is an example of a question answer data:
[0072]
Question
[0073] [Standard answer] 10-5=5 (pieces) Answer: Xiao Ming has 5 more apples than Xiao Hong.
[0074]
User answer
[0075] In a possible implementation, the domain training data may include second training data, which is unlabeled data and includes a plurality of question answer data. .
[0076] On this basis, the second training data can be used to pre-train the base large model in an unsupervised training manner to obtain a pre-trained large model.
[0077] The goal of the above process of pre-training the base large model using the second training data may be to maximize the following objective function:
[0078]
[0079] Among them, u represents token, represents the model parameters, P represents the probability of the next token, and the training goal is to maximize the log-likelihood probability of the next token.
[0080] In another possible implementation, the domain training data may further include third training data, which is labeled data, including question answer data, step correction results corresponding to the labeled user answers, and the whole question score label:
[0081] .
[0082] Among them, x represents the question answer data, y represents the step correction result, and z represents the whole question score label.
[0083] The step correction result is to mark the correction result of each step in the user's answer, including correct, wrong, omitted, redundant, etc. The whole question scoring label is the correction result of the process score of the user's answer to the whole question. It not only needs to consider whether the user's final answer is correct, but also needs to refer to the step correction and scoring points to give an appropriate process score when the user's answer is not completely correct.
[0084] Taking mathematics as an example, an example of the third training data is provided as follows:
[0085]
Question
[0086] [Standard answer] 10-5=5 (pieces) Answer: Xiao Ming has 5 more apples than Xiao Hong.
[0087]
Student answer
[0088] [Step Correction] The first step is wrong. The score for the whole question is 0.
[0089] In a possible implementation, in order to adapt to questions with different scores, the whole question score label annotated in this application can be a score rate value between 0 and 1 (including the endpoints). The actual score of the user's answer can be obtained by multiplying the whole question score by the question score.
[0090] Based on the large model obtained by unsupervised training using the second training data, the third training data can be further used to perform supervised fine-tuning on the pre-trained large model to obtain a fine-tuned large model.
[0091] In the process of using the third training data to fine-tune the pre-trained large model in a supervised manner, the objective function used can be the same as the objective function of the pre-training stage, but in the supervised fine-tuning process, the question answer data does not calculate the loss, and only the step correction results corresponding to the annotated user answers and the whole question score label are calculated. The training goal of the fine-tuning process is to maximize the log-likelihood probability of the generated step correction results and the whole question score given the question answer data.
[0092] After the above-mentioned supervised fine-tuning process, a small amount of high-quality step correction results and whole-question scoring data that have been manually marked are used to supervise the large model after pre-training in the previous stage. This can activate the instruction following ability of the large model and learn to master the step correction task. The fine-tuned large model already has the ability to perform the step correction task, but due to the small amount of third-party training data, the performance of the large model on the step correction task has not yet reached the optimal effect, so it still needs to undergo subsequent steps of reinforcement learning training.
[0093] Step S110: Obtain first training data, wherein the first training data at least includes question answering data and a whole-question scoring label of the user's answer.
[0094] Specifically, in order to facilitate reinforcement learning in the subsequent steps, in this step, the first training data required for the reinforcement learning stage is first obtained. . Among them, x represents the question answer data, and z represents the whole question score label.
[0095] The first training data may include a large amount of question answering data and the overall scoring labels of the user's answers. Compared with the step correction result annotation data, the whole question scoring labels are easier to obtain in large quantities. For example, question answering data and whole question scoring labels can be collected in many educational correction scenarios (such as the college entrance examination, monthly exams of various schools, etc.).
[0096] Step S120: Using the large model as the initial step correction large model, sampling the output of the step correction large model step by step for the user's answers in the first training data, and estimating the accurate probability label of each step correction result based at least on the sampling results and the whole question score label of the user's answer.
[0097] In a possible implementation, this step can use a Monte Carlo tree search method to perform a large number of step correction searches in the step correction problem space, and use the Monte Carlo score to estimate the probability label of the step correction result based on the accuracy of the whole question score. That is, based on the Monte Carlo tree search method, the output of the step correction large model is sampled step by step, and the probability label of each step correction result is estimated based on at least the sampled results and the whole question score label answered by the user.
[0098] Monte Carlo Tree Search (MCTS) is a heuristic search algorithm used in decision-making processes. MCTS guides current decisions by simulating possible future moves and combining random sampling to estimate the pros and cons of each move.
[0099] In this step, the Monte Carlo tree search method is used to search in the step correction solution space. Specifically, the question answer data in the first training data is used as the input of the step correction model to obtain the output of the step correction model (including the step correction results of each step in the user's answer and the score of the whole question). The above process of reasoning using the step correction model can be executed for a set number of sampling times, and finally obtain the sampling results of multiple samplings.
[0100] After obtaining the sampling results, the probability of the correctness of the grading results of each step can be estimated based on the sampling results and the whole question score label answered by the user, which serves as the probability label of the correctness of the grading results of each step.
[0101] By adopting the Monte Carlo tree search method and combining the overall score labels of the user's answers to estimate the accurate probability labels of each step correction result, there is no need for large-scale manual labeling of step-level correction results, saving labeling costs.
[0102] It is understandable that, in addition to the Monte Carlo tree search method, other random sampling algorithms may also be used.
[0103] Step S130: train the step correction large model by using the first training data and the estimated accurate probability label of each step correction result to obtain the trained step correction large model.
[0104] After the above steps, the user answers in the first training data have accurate probability labels of step correction results and whole question score labels. On this basis, in this step, a variety of optional training methods can be used to train the step correction big model to constrain the step correction process of the step correction big model and obtain the final trained step correction big model.
[0105] In a possible implementation, in this embodiment, a reinforcement learning training method can be used to constrain the step correction process of the step correction large model to obtain the step correction large model after reinforcement learning.
[0106] Through the reinforcement learning process, the step correction ability and the whole question scoring and correction ability of the step correction model can be further improved.
[0107] The training method of the step correction large model provided in the embodiment of the present application can adopt a reinforcement learning training strategy. The Monte Carlo tree search method can be used in the training strategy to search in the step correction solution space, and the step correction ability can be reversely improved with the help of the large model's accurate whole-question scoring and correction ability, thereby achieving a training effect similar to self-evolution. Ultimately, a large amount of whole-question scoring training data (a small amount of step correction training data can also be added) can be used to train a step correction large model.
[0108] In some embodiments of the present application, taking the Monte Carlo tree search method used in the aforementioned step S120 as an example, an optional implementation process of step S120 is introduced, which may specifically include the following processing process:
[0109] Traverse each step in the user's answer, and for the current t-th step traversed:
[0110] Sampling is performed according to the configured number of sampling times. The correction results of the first t steps are fixed each time sampling is performed, and the step correction model is used to continue to predict the correction results of subsequent steps and the score of the entire question to obtain the sampling results.
[0111] Specifically, assume that the step correction model has predicted the correction results of the first t steps as y1,…y t When sampling, the correction results of the first t steps can be fixed, and the step correction model can continue to predict the correction results of the t+1th and subsequent steps, as well as the score z of the entire question. The i-th sampling process can be expressed using the following formula:
[0112] .
[0113] On this basis, at least the Monte Carlo score corresponding to each step of the correction result can be calculated based on the sampling results and the whole question score label answered by the user, and the Monte Carlo score can be used as the accurate probability label of the correction result of each step.
[0114] In one possible implementation, for the current t-th step, the result reward score can be calculated based on the whole question score in the sampling result and the whole question score label of the user's answer.
[0115] The Monte Carlo score corresponding to the correction result of the current step t is determined at least based on the result reward score.
[0116] That is, in this embodiment, the accuracy of the whole question scoring can be used as the result reward score, and the result reward score can be used to estimate the Monte Carlo score of the grading result of the tth step, and the Monte Carlo score can be used to estimate the probability label of the accuracy of the grading result of the tth step.
[0117] The formula is as follows:
[0118]
[0119] in, represents the Monte Carlo score of the tth step, that is, the probability label of the correctness of the tth step correction result. k represents the total number of samplings, and r_o represents the result reward score calculated for one sampling. For one sampling result, if the whole question score and the whole question score label answered by the user are consistent, then r_o=1, otherwise r_o=-1.
[0120] In another possible implementation, the first training data may also include step correction result annotations corresponding to some user answers. That is, at least some user answers in the first training data may also be annotated with step correction results.
[0121] On this basis, the process of calculating the Monte Carlo score corresponding to each step correction result based on at least the sampling result and the whole question score label answered by the user may also include:
[0122] For the current t-th step, the process reward score r_p is calculated based on the correction results of each step after the current t-th step in the sampling results and the step correction results marked by each step.
[0123] In this embodiment, an additional process bonus score is introduced to address the situation where a step is marked incorrectly but the overall question is scored correctly.
[0124] In a possible implementation, the process reward score r_p can be calculated as follows:
[0125] Determine the proportion of step correction results that are consistent with the step correction results marked by the corresponding step in the correction results of each step after the current t-th step in the sampling results, and use the proportion as the process reward score r_p.
[0126] The calculation formula can be expressed as follows:
[0127] r_p=m / (nt)
[0128] Among them, n represents the number of steps answered by the user, and m represents the number of step correction results of the t+1th and subsequent steps in the sampling results that are consistent with the actual step correction results (that is, the marked step correction results).
[0129] The total reward score r is determined based on the result reward score r_o and the process reward score r_p, and the total reward score r is used as the Monte Carlo score corresponding to the correction result of the current step t:
[0130]
[0131] r=r_o+r_p
[0132] In actual usage scenarios, a small amount of training data with manually labeled step correction results can be used to calculate the process reward score r_p. A large amount of data that is only labeled with the entire question score label cannot calculate r_p and can be regarded as r_p=0.
[0133] As more training data with manually labeled step-by-step correction results are obtained, the total reward score will take the process reward score into consideration more. The more accurate the Monte Carlo score of the correction result of the t-th step is, the more accurate the probability label of the correction result of the t-th step is.
[0134] In some embodiments of the present application, taking the correction process of a math problem as an example, an example of calculating the result reward r_o and the process reward r_p is illustrated:
[0135]
Question
[0136] [Standard answer] 10+5=15 (apples), 15+3=18 (apples), Answer: They have 18 apples in total.
[0137]
User’s answer
[0138] [Real grading results] The first step is correct. The second step is wrong. The third step is wrong. The score is 0.33 points.
[0139] Taking the process reward r_p and result reward r_o of the first step "y1=correct" in the evaluation of the user's answer as an example, each time sampling, the question answer data is sent to the step correction model to obtain the correction results of each step and the whole question scoring result of the model output. Assuming that the sampling is performed 3 times, the results of each sampling are:
[0140]
Sampling result 1
[0141] [Sampling result 2] The first step is correct. The second step is wrong. The third step is wrong. The score is 0.33 points.
[0142] [Sampling result 3] The first step is correct. The second step is correct. The third step is wrong. The score is 0.66 points.
[0143] Based on these three sampling results, we can calculate:
[0144] r_o = ((-1)+1+(-1)) / 3 = -1 / 3;
[0145] r_p = (2 / 2+2 / 2+1 / 2) / 3 = 5 / 6;
[0146] r(y1=correct) = r_o+r_p = 1 / 2.
[0147] In some embodiments of the present application, the process of correcting the large model in the training step in the aforementioned step S130 is described.
[0148] In this embodiment, the reinforcement learning training method is taken as an example to introduce the reinforcement learning training process of the step-by-step correction large model.
[0149] x is still used to represent the question answering data (question, standard answer and user answer), y t The reinforcement learning process includes two network models to be trained, namely the policy network Actor and the evaluation network Critic.
[0150] The policy network is a step correction model to be trained by reinforcement learning, which can be initialized using the initial step correction model described in step S120. It means taking the input question answer data x and the grading results of the first t steps, predicting the grading results of the t+1th and subsequent steps, and the score of the whole question.
[0151] The evaluation network is a model used to evaluate the correctness of step corrections in reinforcement learning training. Specifically, it can evaluate the probability of each step correction result output by the Actor being accurate. It indicates the probability of predicting the correct result of the t-th step.
[0152] The evaluation network can adopt a variety of neural network structures. In this embodiment, the text generation layer at the end of the initial step correction model can be replaced by a probability prediction layer.
[0153] In the reinforcement learning process, the first training data and the estimated accurate probability labels of each step correction result can be used to train the policy network and the evaluation network. After the training is completed, the policy network is used as the large model for step correction after reinforcement learning.
[0154] In the reinforcement learning process, the evaluation network is trained according to the first goal, and the policy network is trained according to the second goal.
[0155] The first goal includes minimizing the distance between the probability of the correct result of each step of the evaluation network output and the probability label of the correct result of each step. The loss function can use MSE loss, that is:
[0156]
[0157] The second goal consists in maximizing the probability that the correction result is accurate at each step, namely: .
[0158] Optionally, the second goal may also include:
[0159] Maximize total reward score ,The total reward score includes the reward for correctly grading the entire question, and the reward for the accurate grading results of each step, r=r_o+r_p.
[0160] Optionally, the second goal may also include:
[0161] The distance between the output of the policy network and the output of the initial step correction model is constrained by KL divergence:
[0162] .
[0163] in, represents the policy network, Represents the initial steps to modify the large model.
[0164] By adding the KL divergence constraint, the output of the strategy network cannot be too far away from the initial step correction model, which can improve the stability of model training.
[0165] Based on the above embodiments, the second goal may include:
[0166] .
[0167] In a possible implementation, the reinforcement learning process of this application can adopt a two-stage training strategy:
[0168] In the first training phase, the evaluation network is trained separately according to the first objective.
[0169] In the second training stage, the evaluation network and the policy network are jointly trained. The joint training process trains the evaluation network according to the first goal and trains the policy network according to the second goal.
[0170] Through the above two-stage training strategy, the performance of the policy network can be gradually improved, and finally the policy network after reinforcement learning is obtained as the final step correction model, which improves the step correction ability and the whole question scoring correction ability of the step correction model.
[0171] 2. Reasoning Stage
[0172] Reference Figure 2 The homework marking method provided in this embodiment specifically includes the following steps:
[0173] Step S200: Acquire question answer data, wherein the question answer data includes user answers to be corrected, questions and standard answers.
[0174] Step S210: Send the question answer data to the step correction model to obtain the correction result of each step in the user's answer output by the model and the overall score of the user's answer.
[0175] Among them, the step correction large model is a model trained using the step correction large model training method introduced in the above embodiment.
[0176] By adopting the model training method of the aforementioned embodiment, the existing training data can be efficiently used to train the step correction model, achieving the effect of taking into account both the step correction ability and the scoring ability. Furthermore, the trained step correction model can be used to assist in correcting students' homework, thereby improving the efficiency of homework correction.
[0177] The following is a description of the large model training device for step correction provided in an embodiment of the present application. The large model training device for step correction described below and the large model training method for step correction described above can refer to each other.
[0178] See also Figure 3 , Figure 3 This is a schematic diagram of the structure of a large model training device for step correction disclosed in an embodiment of the present application.
[0179] like Figure 3 As shown, the device may include:
[0180] A large model acquisition unit 11 is used to acquire a large model trained with domain training data, wherein the domain training data at least includes question answering data, and the question answering data includes questions, standard answers, and user answers;
[0181] A first training data acquisition unit 12 is used to acquire first training data, wherein the first training data at least includes the question answering data and the whole question scoring label of the user's answer;
[0182] The label estimation unit 13 is used to use the large model as the initial step correction large model, sample the output of the step correction large model step by step for the user's answer in the first training data, and estimate the probability label of each step correction result based on at least the sampling result and the whole question score label of the user's answer;
[0183] The training unit 14 is used to train the step correction large model by using the first training data and the estimated accurate probability label of each step correction result to obtain the trained step correction large model.
[0184] In a possible implementation, the label estimation unit samples the output of the step correction large model step by step, and estimates the accurate probability label of each step correction result based on at least the sampling results and the whole question score label answered by the user, including:
[0185] Based on the Monte Carlo tree search method, the output of the step correction model is sampled step by step, and based on at least the sampling results and the whole question score label answered by the user, the accurate probability label of each step correction result is estimated.
[0186] In a possible implementation, the label estimation unit samples the output of the step correction large model step by step based on the Monte Carlo tree search method for the user's answer in the first training data, and estimates the accurate probability label of each step correction result based on at least the sampling result and the whole question score label of the user's answer, including:
[0187] Traverse each step in the user's answer, and for the current t-th step traversed:
[0188] Sampling is performed according to the configured sampling times. The grading results of the first t steps are fixed each time sampling is performed, and the step-grading large model is used to continue predicting the grading results of the subsequent steps and the score of the entire question to obtain the sampling results;
[0189] Based at least on the sampling result and the whole question score label answered by the user, the Monte Carlo score corresponding to each step correction result is calculated, and the Monte Carlo score is used as the accurate probability label of each step correction result.
[0190] In a possible implementation, the label estimation unit calculates the Monte Carlo score corresponding to each step correction result based on at least the sampling result and the whole question score label answered by the user, including:
[0191] For the current t-th step, based on the whole question score in the sampled result and the whole question score label answered by the user, a result reward score is calculated;
[0192] A Monte Carlo score corresponding to the correction result of the current step t is determined based at least on the result reward score.
[0193] In a possible implementation, at least some of the user answers in the first training data are also annotated with step correction results. On this basis, the label estimation unit calculates the Monte Carlo score corresponding to each step correction result based on at least the sampling result and the whole question score label of the user's answer, and the process also includes:
[0194] For the current t-th step, based on the correction results of each step after the current t-th step in the sampling result and the step correction results marked by each step, the process reward score is calculated;
[0195] A total reward score is determined based on the result reward score and the process reward score, and the total reward score is used as the Monte Carlo score corresponding to the correction result of the current t-th step.
[0196] In a possible implementation, the label estimation unit calculates the process reward score for the current t-th step based on the correction results of each step after the current t-th step in the sampling result and the step correction results of each step annotation, including:
[0197] Determine the proportion of step correction results that are consistent with the step correction results marked by the corresponding step in the correction results of each step after the current t-th step in the sampling result, and use the proportion as the process reward score.
[0198] In a possible implementation, the training unit trains the step correction big model using the first training data and the estimated accurate probability label of each step correction result to obtain the trained step correction big model, including:
[0199] By using the first training data and the estimated accurate probability label of each step correction result, reinforcement learning is used to constrain the step correction process of the step correction large model to obtain the step correction large model after reinforcement learning.
[0200] In a possible implementation, the training unit uses reinforcement learning to constrain the step correction process of the step correction large model to obtain the step correction large model after reinforcement learning, including:
[0201] The initial step correction big model is used as the policy network Actor, and is combined with the evaluation network Critic to perform reinforcement learning to obtain the step correction big model after reinforcement learning. The policy network is used to predict the correction result of each step in the user's answer and the score of the whole question according to the input question answer data, and the evaluation network is used to evaluate the probability of the accuracy of each step correction result output by the policy network;
[0202] In the reinforcement learning process, the evaluation network is trained according to a first goal, and the policy network is trained according to a second goal. The first goal includes minimizing the distance between the probability that the correction result of each step output by the evaluation network is accurate and the probability label of the accurate correction result of each step, and the second goal includes maximizing the probability that the correction result of each step is accurate.
[0203] In one possible implementation, the reinforcement learning process includes two training stages. In the first training stage, the evaluation network is trained separately according to the first objective. In the second training stage, the evaluation network and the policy network are jointly trained. In the joint training process, the evaluation network is trained according to the first objective, and the policy network is trained according to the second objective.
[0204] In a possible implementation, the second objective further includes:
[0205] Maximize the total reward score, which includes the reward for correctly grading the entire question and the reward for accurately grading each step.
[0206] In a possible implementation, the second goal also includes: constraining the distance between the output of the policy network and the output of the initial step-corrected large model through a KL divergence.
[0207] In a possible implementation, the domain training data includes: second training data and third training data, the second training data includes a plurality of question answer data, and the third training data includes the question answer data, the step correction results corresponding to the user answer and the whole question score label. Then the process of the large model acquisition unit acquiring the large model trained with the domain training data includes:
[0208] Pre-training the base large model using the second training data to obtain a pre-trained large model;
[0209] The pre-trained large model is fine-tuned in a supervised manner using the third training data to obtain a fine-tuned large model.
[0210] The present application also provides an electronic device in an embodiment. Figure 4 As shown, it shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to fixed terminals such as servers, mobile phones, tablet computers, teaching large screens, learning machines, etc. Figure 4 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0211] like Figure 4 As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603, so as to implement the step-corrected large model training method or the homework correction method of the aforementioned embodiment of the present application. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in RAM 603. The processing device 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0212] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 4 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0213] Also provided in an embodiment of the present application is a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any one of the step-grading large model training methods or homework grading methods provided in the embodiments of the present application.
[0214] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any one of the step-correcting large model training methods or homework correction methods provided in the embodiments of the present application.
[0215] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.
[0216] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0217] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0218] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.
[0219] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can refer to each other.
Claims
1. A step-by-step correction large model training method, characterized in that: include: Obtaining a large model trained with domain training data, wherein the domain training data at least includes question answering data, and the question answering data includes questions, standard answers, and user answers; Acquire first training data, wherein the first training data at least includes the question answering data and the whole question scoring label of the user's answer; Using the large model as the initial step correction large model, for the user's answer in the first training data, traverse each step in the user's answer, and for the current t-th step traversed: sampling is performed according to the configured sampling times, and the correction results of the first t steps are fixed each time sampling, and the step correction large model is used to continue to predict the correction results of subsequent steps and the whole question score to obtain the sampling results; based on at least the sampling results and the whole question score label of the user's answer, the Monte Carlo score corresponding to each step correction result is calculated, and the Monte Carlo score is used as the accurate probability label of each step correction result; The step correction large model is trained using the first training data and the estimated accurate probability label of each step correction result to obtain a trained step correction large model.
2. The method according to claim 1, characterized in that The process of calculating the Monte Carlo score corresponding to each step correction result based on at least the sampling result and the whole question score label answered by the user includes: For the current t-th step, based on the whole question score in the sampled result and the whole question score label answered by the user, a result reward score is calculated; A Monte Carlo score corresponding to the correction result of the current step t is determined based at least on the result reward score.
3. The method according to claim 2, characterized in that At least some of the user's answers in the first training data are also annotated with step correction results; The process of calculating the Monte Carlo score corresponding to each step correction result based at least on the sampling result and the whole question scoring label answered by the user also includes: For the current t-th step, based on the correction results of each step after the current t-th step in the sampling result and the step correction results marked by each step, the process reward score is calculated; A total reward score is determined based on the result reward score and the process reward score, and the total reward score is used as the Monte Carlo score corresponding to the correction result of the current t-th step.
4. The method according to claim 3, characterized in that For the current t-th step, based on the correction results of each step after the current t-th step in the sampling result and the step correction results of each step annotation, the process of calculating the process reward score includes: Determine the proportion of step correction results that are consistent with the step correction results marked by the corresponding step in the correction results of each step after the current t-th step in the sampling result, and use the proportion as the process reward score.
5. The method according to claim 1, characterized in that The process of training the step correction big model by using the first training data and the estimated accurate probability label of each step correction result to obtain the trained step correction big model includes: By using the first training data and the estimated accurate probability label of each step correction result, reinforcement learning is used to constrain the step correction process of the step correction large model to obtain the step correction large model after reinforcement learning.
6. The method according to claim 5, characterized in that The step correction process of the step correction big model is constrained by using reinforcement learning to obtain the step correction big model after reinforcement learning, including: The initial step correction big model is used as the policy network Actor, and is combined with the evaluation network Critic to perform reinforcement learning to obtain the step correction big model after reinforcement learning. The policy network is used to predict the correction result of each step in the user's answer and the score of the whole question according to the input question answer data, and the evaluation network is used to evaluate the probability of the accuracy of each step correction result output by the policy network; In the reinforcement learning process, the evaluation network is trained according to a first goal, and the policy network is trained according to a second goal. The first goal includes minimizing the distance between the probability that the correction result of each step output by the evaluation network is accurate and the probability label of the accurate correction result of each step, and the second goal includes maximizing the probability that the correction result of each step is accurate.
7. The method according to claim 6, characterized in that The reinforcement learning process includes two training stages. In the first training stage, the evaluation network is trained separately according to the first goal. In the second training stage, the evaluation network and the policy network are jointly trained. In the joint training process, the evaluation network is trained according to the first goal, and the policy network is trained according to the second goal.
8. The method according to claim 6 or 7, characterized in that: The second objective also includes: Maximize the total reward score, which includes the reward for correctly grading the entire question and the reward for accurately grading each step.
9. The method according to claim 6 or 7, characterized in that: The second objective also includes: The distance between the output of the policy network and the output of the initial step-corrected large model is constrained by the KL divergence.
10. The method according to claim 1, characterized in that The domain training data includes: second training data and third training data, wherein the second training data includes a plurality of question answer data, and the third training data includes the question answer data, the step correction results corresponding to the user's answer and the whole question score label; The process of obtaining a large model trained with domain training data includes: Pre-training the base large model using the second training data to obtain a pre-trained large model; The pre-trained large model is fine-tuned in a supervised manner using the third training data to obtain a fine-tuned large model.
11. A homework marking method, characterized in that: include: Acquire question answer data, wherein the question answer data includes user answers to be corrected, questions and standard answers; The question answer data is sent to the step correction model to obtain the correction result of each step in the user's answer and the whole question score of the user's answer output by the model; Among them, the step correction large model is a model trained using the step correction large model training method of any one of claims 1 to 10.
12. A step-correction large model training device, characterized in that: include: A large model acquisition unit, used to acquire a large model trained with domain training data, wherein the domain training data at least includes question answering data, and the question answering data includes questions, standard answers, and user answers; A first training data acquisition unit, configured to acquire first training data, wherein the first training data at least includes the question answering data and a whole question scoring label of the user's answer; a label estimation unit, configured to use the large model as the initial step correction large model, for the user answer in the first training data, traverse each step in the user answer, and for the current t-th step traversed: sampling according to the configured sampling times, fixing the correction results of the first t steps at each sampling, and using the step correction large model to continue predicting the correction results of subsequent steps and the whole question score to obtain the sampling results; at least based on the sampling results and the whole question score label of the user answer, calculate the Monte Carlo score corresponding to each step correction result, and use the Monte Carlo score as the accurate probability label of each step correction result; The training unit is used to train the step correction large model by using the first training data and the estimated accurate probability label of each step correction result to obtain the trained step correction large model.
13. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the steps of the step-grading large model training method as described in any one of claims 1 to 10, or to implement the steps of the homework grading method as described in claim 11.
14. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the step-grading large model training method as described in any one of claims 1 to 10, or implements the steps of the homework grading method as described in claim 11.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the steps of the step-grading large model training method as described in any one of claims 1 to 10, or implements the steps of the homework grading method as described in claim 11.
Citation Information
Patent Citations
Test question answering scoring method, related device, equipment and storage medium
CN118394879A
Error detection model training method, error detection method, device and equipment
CN118410341A