Model training method, device, storage medium and program product
By filtering responses with different dominant values as positive and negative sample pairs during the training process of the question-answer model, and updating parameters based on the current and reference probability, the problem of mismatch between the training data and the dynamic error distribution is solved, and the model training efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510774535.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-06-11
AI Technical Summary
During the model training process, the existing technology relies on manual judgment or simple random sampling to select training data, resulting in the training data not matching the dynamic error distribution of the model, resulting in low model training efficiency.
By generating a question response group during the training of the question-answer model, filtering out responses with different dominant values as positive and negative sample pairs, and updating the model parameters in combination with the current and reference probability to ensure that the training data matches the dynamic error distribution of the model.
It improves the efficiency of model training, shortens the training time, ensures the correct parameter update direction, and improves the accuracy and efficiency of model training.
Smart Images

Figure CN120278285B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to model training methods, devices, storage media, and program products. Background Art
[0002] Currently, in the training process of some models, a portion of data from a sample dataset is typically selected as training data, and the model is trained based on this training data. However, when selecting training data from a sample dataset, relying on manual judgment or simple random sampling, the selected training data may not match the dynamic error distribution of the model, resulting in relatively low model training efficiency. Summary of the Invention
[0003] The present application provides a model training method, a model training device, an electronic device, a computer-readable storage medium, and a computer program product to at least solve the problem of relatively low model training efficiency in related technologies.
[0004] This application provides a model training method, which includes:
[0005] In a current training iteration, questions in the first sample data set are input into a question-answering model, and the question-answering model is used to generate a question-response group for each question, each question-response group including multiple responses to the same question and a current probability of the question-answering model generating each response.
[0006] determining a dominance value for each of the responses, wherein the dominance value represents the accuracy of each of the responses relative to other responses in the set of responses to the question;
[0007] In each of the question response groups, a first response having a first advantage value and a second response having a second advantage value are respectively screened, wherein the first advantage value is higher than the second advantage value;
[0008] updating parameters of the question-answering model based on current probabilities of the first response and the second response and reference probabilities, where the reference probability refers to the probability that the question-answering model generated the first response and the second response before the current training iteration;
[0009] In the next training iteration, the question-answering model with updated parameters obtained in the current training iteration is used to regenerate the question-response group of the question, and based on the regenerated question-response group, the first response and the second response are rescreened, and the parameters of the question-answering model are updated based on the current probability and reference probability of the rescreened first response and the second response.
[0010] The present application also provides a model training device, comprising:
[0011] a response generation module, configured to input, in a current training iteration, questions in the first sample dataset into a question-answering model, the question-answering model being configured to generate a question-response group for each question, each question-response group comprising multiple responses to the same question and a current probability of the question-answering model generating each response;
[0012] a dominance value determination module, configured to determine a dominance value for each of the responses, wherein the dominance value represents the accuracy of each of the responses relative to other responses in the question response group;
[0013] a response screening module, configured to screen, in each of the question response groups, a first response having a first advantage value and a second response having a second advantage value, wherein the first advantage value is higher than the second advantage value;
[0014] a first parameter updating module, configured to update parameters of the question-answering model based on current probabilities and reference probabilities of the first response and the second response, where the reference probability refers to the probability of the question-answering model generating the first response and the second response before the current training iteration;
[0015] The second parameter updating module is used to, in the next training iteration, use the question-answering model with updated parameters obtained in the current training iteration to regenerate the question-response group of the question, and re-screen the first response and the second response based on the regenerated question-response group, and update the parameters of the question-answering model based on the current probability and reference probability of the first response and the second response obtained by re-screening.
[0016] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned model training methods when executing the computer program.
[0017] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned marking methods are implemented.
[0018] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned marking methods when executed by a processor.
[0019] In the technical solutions of some embodiments of the present application, during the training process of the question-answering model, after the question-answering model generates a question response group for each question, the first response and the second response with different advantage values can be screened from each question response group according to the advantage value of each response. Since the advantage value of the first response is higher than the advantage value of the second response, the first response can be regarded as the response of the positive sample pair, and the second response can be regarded as the response of the negative sample pair. Since the positive and negative sample pairs are screened from the output responses of the question-answering model, the positive and negative sample pairs can be matched with the dynamic error distribution of the question-answering model. Furthermore, combined with the current probability of the question-answering model generating the first response and the second response in the current training iteration, and the reference probability of the question-answering model generating the first response and the second response before the current training iteration, the parameter update direction and parameter update amplitude of the question-answering model can be determined more accurately, thereby shortening the training time of the question-answering model, greatly improving the model training efficiency, and solving the problem of low model training efficiency in some technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 A flowchart of a model training method provided for some embodiments of the present application;
[0022] Figure 2 A flowchart of a model training method provided for other embodiments of the present application;
[0023] Figure 3 A schematic diagram of a module of a model training device provided in some embodiments of the present application;
[0024] Figure 4 A schematic diagram of a module of an electronic device provided for some embodiments of the present application. DETAILED DESCRIPTION
[0025] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0026] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0027] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0028] A model's dynamic error distribution refers to the distribution of its prediction errors for input data at different stages of model training. Specifically, the dynamic error distribution can include the distribution of the types and number of model prediction errors. As model training progresses, model performance gradually improves, and its dynamic error distribution changes accordingly. For example, suppose a sentiment prediction model is trained to identify the sentiment (i.e., positive or negative) in input text. In the early stages of training, the sentiment prediction model may inaccurately predict the sentiment of all input text. However, by the mid-stage of training, the sentiment prediction model may learn to recognize some simple emotional expressions and output correct predictions. For example, it may correctly predict the sentiment for 60% of the input text and incorrectly predict the sentiment for the remaining 40%. Here, the distribution of the number of prediction errors of the sentiment prediction model changes from the early to mid-stage of training. For example, in the mid-stage of training, the sentiment prediction model may correctly predict the sentiment for input text with a strong sentiment tendency and incorrectly predict the sentiment for input text with a weak sentiment tendency. However, in the later stages of training, the sentiment prediction model can correctly predict the sentiment for both strong and weak sentiment tendencies. Here, the distribution of the types of prediction errors of the sentiment prediction model changes from the early stage to the middle stage of training.
[0029] Furthermore, during the model training process, if the distribution of training data matches the dynamic error distribution of the model, the model training efficiency can be greatly improved. Conversely, if the distribution of training data does not match the dynamic error distribution of the model, the model training efficiency will be greatly reduced. For example, take the above-mentioned sentiment prediction model as an example. In the middle of training, since the sentiment prediction model cannot correctly predict the sentiment tendency of input texts with unclear sentiment tendencies, if the training data includes as many input texts with unclear sentiment tendencies as possible, the sentiment prediction model can conduct targeted learning on this type of input text, thereby improving the convergence speed of model training and thus greatly improving the model training efficiency. Conversely, if the training data includes relatively few input texts with unclear sentiment tendencies, the sentiment prediction model cannot conduct targeted learning on this type of input text, which may result in a relatively slow convergence speed of model training and thus greatly reduce the model training efficiency.
[0030] Currently, in some technologies, model training data is screened from some publicly available sample data sets. When screening training data, manual judgment or simple random sampling is often relied upon. As a result, the distribution of the training data ultimately screened out may not match the model's dynamic error distribution, leading to low model training efficiency. Furthermore, in these technologies, the training data is static and does not change with the model's dynamic error distribution. Therefore, even if the distribution of the training data initially screened from the sample data set matches the model's dynamic error distribution, as model training progresses, the mismatch between the training data and the model's dynamic error distribution may occur, leading to low model training efficiency.
[0031] In view of this, the present application provides a model training method, which can dynamically screen the positive and negative sample pairs in the training data during the model training process, so that the distribution of training data and the dynamic error distribution of the model can be kept in real time to match each other, thereby improving the efficiency of model training. The model training method can be applied to electronic devices. Electronic devices can include but are not limited to tablets, laptops, desktop computers, servers, etc. Figure 1 , which is a flow chart of the model training method provided in some embodiments of the present application. Figure 1 In [1], the model training method includes the following steps:
[0032] Step S101: In the current training iteration, the questions in the first sample data set are input into the question-answering model. The question-answering model is used to generate a question response group for each question. Each question response group includes multiple responses to the same question and the current probability of each response generated by the question-answering model.
[0033] In this embodiment, the questions in the first sample dataset may include programming questions and non-programming questions. Programming questions are used to instruct the question-answering model to generate program code, while non-programming questions are used to instruct the question-answering model to generate content other than program code. Specifically, programming questions can be used to describe program code features, which may include but are not limited to the logical functions to be implemented by the program code and the rules for writing the program code. Based on the programming questions, the question-answering model outputs program code that matches the programming questions. For example, a programming question might be, "Please output a program code in Java that performs user login authentication." Based on this programming question, the question-answering model might output a program code written in Java. Running this program code can authenticate the user. Non-programming questions refer to questions unrelated to program code, including but not limited to math, geography, history, language, economics, philosophy, and everyday life. For example, a non-programming question might be, "A rectangle is 10 centimeters long and 5 centimeters wide. What is its area in square centimeters?" Based on this non-programming question, the question-answering model might output the answer 50.
[0034] The first sample data set can be a publicly available sample data set, such as Eurus-2-RL-Data. For any piece of sample data in the first sample data set, the sample data can include a question, a reasoning process annotated for the question, and an answer to the question. When training the question-answering model, the sample data in the first sample data set can be divided into multiple batches (i.e., patches), and the question-answering model can be iteratively trained by batch. For example, in the first sample data set, every 256 pieces of sample data can be divided into a batch. By running the question-answering model multiple times or running the question-answering model in parallel, the questions and question prompts in the same batch of sample data can be input into the question-answering model to obtain a question response group generated by the question-answering model for each question.
[0035] In this embodiment, the question prompt may be similar to the following: A conversation between User and Assistant. The user asks a question, and the Assistant solves it. The assistant first thinks about the reasoning process in the mind and then provides the user with the answer. The reasoning process and answer are enclosed within <think>< / think> and <answer>< / answer>tags, respectively, ie, <think> reasoning process here< / think> <answer> answer here< / answer> . User:prompt. Assistant:.
[0036] In this embodiment, each question response group may include 16 responses to the same question, and the responses in the same question response group are not exactly the same. The response may include the reasoning process of the question-answering model to answer the question, the answer obtained based on the reasoning process, and the current probability of the question-answering model generating the corresponding response. For example, 16 question-answering models are run in parallel, and question Q1 in sample data A1 is input into the 16 question-answering models, and the response generated by each question-answering model for question Q1 can be obtained. Due to the randomness of the question-answering model reasoning, the responses generated by different question-answering models for question Q1 may not be exactly the same. The responses generated by these 16 question-answering models for question Q1 can be integrated into a question response group, as the question response group generated by the question-answering model for question Q1. For example, the question response group for question Q1 may be similar to the following:
[0037] <think> Reasoning Process 1< / think> <answer> Answer 1: First current probability< / answer> ;
[0038] <think> Reasoning Process 2< / think> <answer> Answer 2: Second current probability< / answer> ;
[0039] ...;
[0040] <think> Reasoning Process 16< / think> <answer> Answer 16: Sixteenth Current Probability< / answer> ;
[0041] It's important to note that the so-called current probability refers to the probability of the question-answering model generating each response based on the model's current parameters. As you can see, since the model's parameters change dynamically during training, the probability of the model generating the same response changes dynamically across different training iterations.
[0042] Step S102: determining a dominance value of each response, where the dominance value represents the accuracy of each response relative to other responses in the question response group.
[0043] In this embodiment, the question can be marked as a reference response. For the reference response of any question and the various responses generated by the question-answering model for the question, the dominance value of each response can be determined based on the degree of proximity of each response to the reference response. Specifically, if one of the responses is relatively close to the reference response, and the other responses are relatively far from the reference response, the dominance value of the response can be relatively large. If one of the responses is relatively far from the reference response, and the other responses are relatively close to the reference response, the dominance value of the response can be relatively small. If the degree of proximity of multiple responses to the reference response is relatively similar, the dominance values of the multiple responses can be relatively close. For ease of understanding, the following examples are used to illustrate.
[0044] For example, assume that the reference response for question Q1 is reference response R, and the question response group for question Q1 includes responses 1 to 16. If response 1 is closest to reference response R, responses 2 to 10 are closer to reference response R than response 1, and the closeness of each of responses 2 to 10 to reference response R is relatively similar, and responses 11 to 16 are closer to reference response R than responses 2 to 10, then the dominance value of response 1 may be greater than the dominance values of responses 2 to 10, the dominance values of responses 2 to 10 may be greater than the dominance values of responses 11 to 16, and the dominance values of each of responses 2 to 10 may be the same.
[0045] It is understandable that, in a group of question responses to the same question, responses with higher dominance values are more accurate, while responses with lower dominance values are relatively less accurate.
[0046] Step S103 : In each question response group, first responses having a first advantage value and second responses having a second advantage value are respectively screened, wherein the first advantage value is higher than the second advantage value.
[0047] Specifically, in this embodiment, the response with the highest dominance value can be selected from each question response group as the first response, and the response with the lowest dominance value can be selected as the second response. In this way, multiple groups of first responses and second responses can be obtained. For example, the first and second responses of question response group 1 can be selected; the first and second responses of question response group 2 can be selected; and so on.
[0048] It should be noted that other screening methods can also be used to select responses that meet the order of dominance values as the first response and the second response. For example, in each question response group, the response with the highest dominance value can be selected as the first response, and the response with a dominance value slightly higher than the lowest dominance value can be selected as the second response. For another example, in each question response group, the response with a dominance value slightly lower than the highest dominance value can be selected as the first response, and the response with a dominance value slightly higher than the lowest dominance value can be selected as the second response.
[0049] Step S104: Update the parameters of the question-answering model based on the current probabilities and reference probabilities of the first response and the second response. The reference probability refers to the probability of the question-answering model generating the first response and the second response before the current training iteration.
[0050] Specifically, for any question response group, the first response obtained by screening the question response group and the question corresponding to the question response group can be regarded as a positive sample pair, and the second response obtained by screening the question response group and the question corresponding to the question response group can be regarded as a negative sample pair. It can be understood that the purpose of model training is to increase the probability of the question-answering model generating the first response and to reduce the probability of the question-answering model generating the second response. Since the reference probability is the probability that the question-answering model generates the first response and the second response before the current training iteration, the parameters of the question-answering model can be updated in the direction that the current probability of the first response is greater than the reference probability and the current probability of the second response is lower than the reference probability. In this way, it can be ensured that the parameters of the question-answering model are updated in the correct direction, thereby reducing the problem of increased model training time due to parameter update errors, and thus greatly improving the training efficiency of the question-answering model.
[0051] Step S105: In the next training iteration, the question-answering model with updated parameters obtained in the current training iteration is used to regenerate the question response group of the question, and based on the regenerated question response group, the first response and the second response are rescreened, and the parameters of the question-answering model are updated based on the current probability and reference probability of the rescreened first response and the second response.
[0052] In short, steps S101 to S104 can be executed in multiple iterative loops. For example, steps S101 to S104 are executed for the first time to perform a first parameter update on the question-answering model. Based on the question-answering model after the first parameter update, steps S101 to S104 can be executed for a second time to perform a second parameter update on the question-answering model. And so on. The loop can be executed 50 times, 100 times, etc.
[0053] It can be understood that since the first response and the second response are re-screened based on the current dynamic error distribution of the question-answering model in each iterative training, the first response and the second response can reflect the current dynamic error distribution of the question-answering model. Furthermore, based on the current probability and reference probability of the first response and the second response in each iterative training, the model training efficiency can be greatly improved.
[0054] In summary, in the technical solutions of some embodiments of the present application, during the training process of the question-answering model, after the question-answering model generates a question response group for each question, the first response and the second response with different advantage values can be screened from each question response group according to the advantage value of each response. Since the advantage value of the first response is higher than the advantage value of the second response, the first response can be regarded as the response of the positive sample pair, and the second response can be regarded as the response of the negative sample pair. Since the positive and negative sample pairs are screened from the output responses of the question-answering model, the positive and negative sample pairs can be matched with the dynamic error distribution of the question-answering model, and then combined with the current probability of the question-answering model generating the first response and the second response in the current training iteration, and the reference probability of the question-answering model generating the first response and the second response before the current training iteration, the parameter update direction of the question-answering model can be determined more accurately, thereby shortening the training time of the question-answering model, greatly improving the model training efficiency, and solving the problem of low model training efficiency in some technologies.
[0055] The above steps S102 and S104 are further described below.
[0056] In some embodiments, determining the advantage value of each response in step S102 may include:
[0057] Generate reward values for each response, which represent the accuracy of the response;
[0058] In the same question response group, the dominance value of each response is determined based on the reward value distribution between the responses. For any response, the dominance value of the response is related to the discreteness of the reward value distribution and the deviation degree of the reward value of the response. The deviation degree of the reward value refers to the distance of the reward value of the response from the center position of the reward value distribution.
[0059] Specifically, the more accurate the response, the higher the reward value. Conversely, the less accurate the response, the lower the reward value.
[0060] For the same set of responses to a question, if the reward distribution has a low degree of dispersion (i.e., the rewards are relatively concentrated), and the reward of a response is greater than the center of the reward distribution, then the dominance of that response can be high. Conversely, if the reward of a response is less than the center of the reward distribution, then the dominance of that response can be low.
[0061] For the same set of responses to a question, if the reward distribution is highly dispersed (i.e., the reward distribution is wide), if there are multiple responses with reward values greater than the center of the reward distribution, then the dominance values of these responses do not need to be high (because the presence of multiple responses with reward values greater than the center does not make these responses particularly dominant). Similarly, if there are multiple responses with reward values less than the center of the reward distribution, then the dominance values of these responses do not need to be low.
[0062] Specifically, within the same set of responses to a question, if the reward value distribution has a high degree of dispersion and there are relatively few responses with reward values greater than the center of the reward value distribution, then a response with a reward value greater than the center of the reward value distribution may have a higher dominance value. Similarly, within the same set of responses to a question, if the reward value distribution has a high degree of dispersion and there are relatively few responses with reward values less than the center of the reward value distribution, then a response with a reward value less than the center of the reward value distribution may have a lower dominance value.
[0063] In the above embodiment, the advantage value of the response is determined based on the reward value distribution in the same question response group and the deviation degree of the reward value of the response. The advantage value can be quantified to ensure the calculation accuracy of the advantage value.
[0064] Specifically, in some embodiments, the reward value mean and reward value variance of the responses in the same question response group can be determined, and the reward value distribution of the question response group can be determined based on the reward value mean and reward value variance. The reward value mean can represent the center position of the reward value distribution, and the reward value variance can represent the degree of dispersion of the reward value distribution. The larger the reward value variance, the higher the degree of dispersion of the reward value distribution. Conversely, the smaller the reward value variance, the lower the degree of dispersion of the reward value distribution. In this way, the reward value distribution of each question response group is determined by mathematical statistics, with relatively high accuracy and credibility.
[0065] In some embodiments, for any question response group, after obtaining the reward value variance and reward value mean corresponding to the question response group, the advantage value of each response in the question response group can be determined based on expression (1):
[0066]
[0067] in, represents the advantage value of the i-th response in the response group of the question, represents the reward value of the i-th response in the response group of the question, represents the mean reward value of the response group to this question, Represents the variance of the reward value of the response group for this question.
[0068] The above expression (1) can also be regarded as normalizing the reward values of each response. In this way, the dimensional differences of the reward values of different responses can be eliminated and the comparability between the reward values can be improved.
[0069] In some embodiments, when the first sample data set includes programming questions, generating reward values for respective responses may include:
[0070] For any response, if the response is a response generated by the question-answering model for a target programming question in the first sample data set, and the response includes target program code, running at least one test case for detecting the correctness of the target program code;
[0071] Based on the test case execution results, the reward value of the response is determined.
[0072] Specifically, in the first sample data set, at least one test case for each programming question can be marked. After the question-answering model generates a response for the target programming question, the test case marked for the target programming question can be run to detect whether the target program code generated by the question-answering model is correct. For example, if the test case runs successfully, it means that the target program code is correct, and the reward value of the response can be the third reward value. Conversely, if the test case fails, it means that the target program code is wrong, and the reward value of the response can be the fourth reward value. The third reward value can be greater than the fourth reward value. For example, the third reward value can be 1 and the fourth reward value can be 0. By detecting the correctness of the target program code by running the test case, the logical accuracy of the target program code can be verified more accurately.
[0073] Furthermore, when the target programming problem has multiple annotated test cases, some test cases may pass while others fail. In this case, it is not possible to determine the reward value of the response based on whether the test case passes. In view of this, in some embodiments, determining the reward value of the response based on the test case execution results may include:
[0074] Get the total number of test cases for the target program code and the number of test cases that have passed the run;
[0075] The reward value of the response is determined based on the proportion of the number of test cases that have passed the run to the total number of test cases, where the proportion is proportional to the reward value of the response.
[0076] In this way, when only some test cases pass the test, the corresponding reward value can still be determined, thereby improving the applicability of the solution.
[0077] Furthermore, in some embodiments, different test cases may have different levels of importance. For example, suppose the target programming problem has two labeled test cases, test case A and test case B. If test case A fails, it indicates that the overall logic of the target program code is incorrect. If test case B fails, it indicates that the value of parameter A in the target program code is out of range. Obviously, the importance of test case A will definitely be higher than that of test case B. If test case A fails, it can indicate that the target program code is incorrect, and the corresponding reward value can be lower. However, if test case B fails, it can indicate that only part of the target program code is incorrect, and the corresponding reward value can be relatively high. In view of this, each test case can also have its own corresponding weight, wherein test cases with higher importance can have higher weights, and test cases with lower importance can have lower weights. When calculating the reward value, the weights of each test case can be referenced so that the corresponding reward value is proportional to the importance of each test case, thereby ensuring the accuracy of the reward value.
[0078] In some embodiments, when the first sample data set includes non-programming questions, generating reward values for respective responses may include:
[0079] For any response, if the response is a response generated by the question-answering model for a target non-programming question in the first sample dataset, obtaining a reference response for the target non-programming question;
[0080] If the response matches the reference response, a first reward value is generated for the response, and if the response does not match the reference response, a second reward value is generated for the response, wherein the first reward value is greater than the second reward value.
[0081] Specifically, the reference response is the reasoning process and question result labeled for the target non-programming question. The response generated by the question-answering model is compared with the reference response, and a reward value is generated based on the comparison result, in line with the normal model training process.
[0082] In some embodiments, a reward model can be pre-trained. After the question-answering model generates responses to questions, the responses can be fed into the trained reward model to obtain reward values for each response. This can reduce the reward calculation process and improve model training efficiency.
[0083] In some embodiments, updating the parameters of the question-answering model based on the current probabilities and reference probabilities of the first and second responses in step S104 may include:
[0084] determining a first log-likelihood ratio between the current probability of the first response and the reference probability, and determining a second log-likelihood ratio between the current probability of the second response and the reference probability;
[0085] Determine the changing trend of the probability difference between the first response and the second response output by the question-answering model based on the first log-likelihood ratio and the second log-likelihood ratio;
[0086] Update the parameters of the question-answering model based on the changing trend of the probability difference.
[0087] Specifically, the first loss function shown in Expression (2) can be constructed based on the Direct Preference Optimization (DPO) framework.
[0088]
[0089] in, represents the question-answering model of the current training iteration, represents the question-answering model before the current training iteration, X represents the question, Indicates the first response, Indicates the second response, represents the current probability of the question-answering model generating the first response in the current training iteration, represents the reference probability of the first response generated by the question-answering model before the current training iteration, represents the first log-likelihood ratio between the current probability of the first response and the reference probability, represents the current probability of the question-answering model generating the second response in the current training iteration, represents the reference probability of the second response generated by the question-answering model before the current training iteration, represents the second log-likelihood ratio between the current probability and the reference probability of the second response, Indicates the changing trend of the probability difference between the first and second responses output by the question-answering model, is the temperature coefficient (for example, 0.1), represents the base of the logarithmic function, Represents the first loss value.
[0090] From expression (2), we can see that the first log-likelihood ratio can represent the relative change in the probability of the question-answering model outputting the first response and the second response at the current training iteration, and the second log-likelihood ratio can represent the relative change in the probability of the question-answering model outputting the first response and the second response before the current training iteration. For example, before the current training iteration, the probability of the question-answering model outputting the first response is 0.6, and the probability of outputting the second response is 0.4, with a probability change of 0.2. In the current training iteration, the probability of the question-answering model outputting the first response is 0.7, and the probability of outputting the second response is 0.3, with a probability change of 0.4. The trend of change between the probability difference before the current training iteration and the probability difference at the current training iteration is the probability difference change trend, such as from 0.4 to 0.2. It can be understood that the purpose of model training is to increase the probability difference between the first response and the second response output by the question-answering model as much as possible, and at the same time, the probability difference change trend can also be increased as much as possible. Since there is a negative sign in front of the first loss function, the meaning of expression (2) is that in the first loss value When the value of is less than the first threshold, it indicates that the training of the question-answering model is completed.
[0091] In the above embodiment, the probability difference change trend of the first response and the second response output by the question-answering model is determined based on the first log-likelihood ratio and the second log-likelihood ratio, and the parameters of the question-answering model are updated based on the probability difference change trend. This can ensure that the parameters of the question-answering model are updated in the correct direction, thereby improving the efficiency of model training.
[0092] See also Figure 2 In some embodiments, the model training method of the present application may specifically include the following steps:
[0093] 1) Construct a second sample dataset. Specifically, the second sample dataset can be a multi-source heterogeneous open source reasoning dataset that includes multiple training data. The multiple training data can cover the three core areas of mathematics, code generation, and logical reasoning. The training data related to mathematics is centered around the sample dataset Open-R1-Math-220k, and the training data related to code generation and logical reasoning can come from the sample dataset OpenThoughts-114k. After obtaining sample data from the sample datasets Open-R1-Math-220k and OpenThoughts-114k, the sample data in the second sample dataset can be organized according to the format of problem description, reasoning process, solution, standard answer, data source, data type, and test case. That is, each sample data in the second sample dataset includes content such as problem description, reasoning process, solution, standard answer, data source, data type, and test case.
[0094] In some embodiments, the sample data in the second sample data set may be deduplicated and quality filtered. For example, if the problem descriptions of multiple sample data are highly similar, only one of the sample data may be retained and the other similar sample data may be deleted.
[0095] 2) Perform initial training on the question-answering model based on the second sample dataset. Initial training refers to supervised fine-tuning of the question-answering model based on the second sample dataset.
[0096] Specifically, 8 graphics processors can be used to build a model training environment. The model of the graphics processor can be NVIDIA A100, and each graphics processor can have 80GB of video memory. During model training, a data parallel optimization strategy + ZeRO-3 can be adopted, and efficient use of video memory can be achieved through the DeepSpeed acceleration framework. The maximum sequence length that a single graphics processor can carry is 32k tokens. When processing training data with ultra-long sequences (>16k), gradient checkpoint technology can be enabled. Video memory usage can be controlled within 65GB / card. Mixed precision training can use BF16 mode, which can save about 40% of video memory overhead compared to FP32 while maintaining numerical stability.
[0097] In some embodiments, each piece of training data in the second sample data set has a corresponding sequence length. The so-called sequence length refers to the number of tokens included in the question in the training data. The initial training of the question-answering model based on the second sample data set may include:
[0098] According to the sequence length, multiple training data are divided into multiple groups, and the sequence lengths of training data in different groups are in different length ranges;
[0099] In order of length intervals from small to large, each group of training data is used in turn to perform initial training on the question-answering model.
[0100] For example, based on sequence length, multiple training data can be divided into three groups: 4k-8k, 8k-16k, and 16k-32k. That is, the training data in the first group ranges from 4k to 8k, the training data in the second group ranges from 8k to 16k, and the training data in the third group ranges from 16k to 32k. In the first training phase, the question-answering model can be trained using the training data in the first group. The single-card batch size can be set to 4, the global batch size can be expanded to 256 through gradient accumulation (gradient accumulation steps = 8), the initial learning rate can be set to 1e-5, and a linear warmup of 2000 steps can be used. In the second training phase, the question-answering model can be trained using the training data in the second group. The single-card batch size can be reduced to 2, the global batch size can be maintained at 128 (gradient accumulation steps = 8), the learning rate can be reduced to 5e-6, and a cosine annealing schedule can be used. In the third training phase, the question-answering model is trained using the training data from the third group. The single-card batch size can be compressed to 1, the global batch size is maintained at 64 by accumulating 8 steps, and the learning rate is maintained at 2e-6. This can avoid gradient oscillation during training.
[0101] Furthermore, to balance video memory and training efficiency, selective parameter updates can be used. For example, if the question-answering model consists of 28 layers, only the last 14 layers of the question-answering model can be opened to training, and the first 14 layers and position encoding parameters can be completely frozen.
[0102] 3) Execute the above steps S101 and S102.
[0103] 4) Update the parameters of the question-answering model based on the advantage value, current probability, reference probability of each response, and the current probability distribution and reference probability distribution between responses.
[0104] Specifically, the update direction of the first parameter of the question-answering model can be determined based on the advantage value, current probability and reference probability of each response, and the update amplitude of the first parameter of the question-answering model can be determined based on the current probability distribution and reference probability distribution between the responses;
[0105] Update the parameters of the question-answering model according to the first parameter update direction and the first parameter update amplitude.
[0106] Specifically, based on the advantage value, current probability and reference probability of each response, a second loss function as shown in Expression (3) can be constructed, and the update direction of the first parameter of the question-answering model can be determined based on the second loss function.
[0107]
[0108] In expression (3), each response output by the question-answering model is divided into multiple time steps according to the number of tokens. For example, if response A includes three tokens, the first token is the first time step, the second token is the second time step, and the third token is the third time step. The dominance value of each time step is the same and is the dominance value of the response. For example, if the dominance value of response A is 20, then the dominance values of the first time step, the second time step, and the third time step are all 20.
[0109] Based on the above description, in expression (3), represents the tth time step of the ith response, represents the model parameter state of the question-answering model at the tth time step, represents the advantage value of the t-th time step of the i-th response, T represents the number of time steps for each question-response group, N represents the number of question-response groups, represents the question-answering model of the current training iteration, represents the question-answering model before the current training iteration, Indicates the question answering model in the current training iteration Generated in state The probability of Indicates the question-answering model before the current training iteration Generated in state The probability of Indicates that the question answering model was generated before and after the current training iteration. The log-odds ratio of the probability of Indicates Generated in state The quality of the question-answering model can be used to determine the first parameter update direction of the question-answering model. and The function is used to ensure that the parameter update is not too radical, that is, to ensure the stable update of the model parameters. Represents the second loss value.
[0110] Furthermore, a third loss function as shown in Expression (4) can be constructed based on the current probability distribution and the reference probability distribution between the responses, and the update amplitude of the first parameter of the question-answering model can be determined based on the third loss function.
[0111]
[0112] Specifically, represents the third loss value, Indicates the question answering model in the current training iteration Output the probability distribution of all responses in the state, Indicates the question-answering model before the current training iteration The probability distribution of all responses output in the state. KL is also known as Kullback-Leibler divergence, which is used to measure the difference between two probability distributions. Used to quantify the difference between two probability distributions. During model training, if you want If the value is less than the first threshold, the sum of the differences between the two probability distributions of the question response group needs to be less than the second threshold. In this way, the update amplitude of the first parameter can be limited to avoid the problem of unstable model training caused by excessive parameter changes.
[0113] In some embodiments, the first parameter update direction has a first weight, and the first parameter update magnitude has a second weight;
[0114] Update the parameters of the question-answering model according to the first parameter update direction and the first parameter update amplitude, including:
[0115] Determining a second parameter update direction and a second parameter update amplitude of the question-answering model according to the first weight, the second weight, the first parameter update direction, and the first parameter update amplitude;
[0116] Update the parameters of the question-answering model according to the second parameter update direction and the second parameter update amplitude.
[0117] Specifically, a fourth loss function as shown in Expression (5) can be constructed according to the first weight, the second weight, the first parameter update direction, and the first parameter update amplitude.
[0118]
[0119] in, Update amplitude for the first parameter The second weight, the first parameter update direction The second weight of is 1. When the question-answering model is trained based on the loss function shown in Expression (5), the accuracy of the model training can be guaranteed.
[0120] 5) Execute the above steps S103 and S104.
[0121] During the model training process, the question-answering model can be iteratively trained by repeatedly executing steps 3) to 5) above. During the iterative training process, the dynamic error distribution of the question-answering model can change dynamically. Since the first response and the second response used for updating the model parameters in this application are obtained by screening the output results of the question-answering model, the first response and the second response will also change with the change of the dynamic error distribution, that is, the training data distribution of the question-answering model and the dynamic error distribution of the question-answering model can be kept in a matching state in real time, thereby greatly improving the efficiency of model training.
[0122] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0123] The embodiment of the present application also provides a model training device, in conjunction with reference to Figure 3 , which is a module diagram of the model training device provided in some embodiments of the present application. Figure 3 In [1], the model training device includes:
[0124] Response generation module 301, in the current training iteration, inputs questions in the first sample dataset into the question-answering model. The question-answering model is used to generate a question-response group for each question. Each question-response group includes multiple responses to the same question and the current probability of each response generated by the question-answering model.
[0125] A dominance value determination module 302 determines a dominance value for each response, where the dominance value represents the accuracy of each response relative to other responses in the question response group.
[0126] The response screening module 303 screens, in each question response group, a first response having a first advantage value and a second response having a second advantage value, wherein the first advantage value is higher than the second advantage value;
[0127] A parameter updating module 304 updates the parameters of the question-answering model based on the current probabilities of the first and second responses and the reference probabilities, where the reference probabilities refer to the probabilities of the first and second responses generated by the question-answering model before the current training iteration.
[0128] The second parameter updating module 305 is used to, in the next training iteration, use the question-answering model with updated parameters obtained in the current training iteration to regenerate the question-response group of the question, and re-screen the first response and the second response based on the regenerated question-response group, and update the parameters of the question-answering model based on the current probability and reference probability of the re-screened first response and the second response.
[0129] In some embodiments, the advantage value determination module 302 is configured to:
[0130] Generate reward values for each response, which represent the accuracy of the response;
[0131] In the same question response group, the dominance value of each response is determined based on the reward value distribution between the responses. For any response, the dominance value of the response is related to the discreteness of the reward value distribution and the deviation degree of the reward value of the response. The deviation degree of the reward value refers to the distance of the reward value of the response from the center position of the reward value distribution.
[0132] In some embodiments, before determining the dominance value of each response, the dominance value determination module 302 is configured to:
[0133] In the same question response group, the reward value mean and reward value variance of the responses are determined, and the reward value distribution of the question response group is determined based on the reward value mean and reward value variance.
[0134] In some embodiments, the questions in the first sample dataset include programming questions, which are used to instruct the question-answering model to generate program code; the advantage value determination module 302 is used to:
[0135] For any response, if the response is a response generated by the question-answering model for a target programming question in the first sample data set, and the response includes target program code, running at least one test case for detecting the correctness of the target program code;
[0136] Based on the test case execution results, the reward value of the response is determined.
[0137] In some embodiments, the advantage value determination module 302 is configured to:
[0138] Get the total number of test cases for the target program code and the number of test cases that have passed the run;
[0139] The reward value of the response is determined based on the proportion of the number of test cases that have passed the run to the total number of test cases, where the proportion is proportional to the reward value of the response.
[0140] In some embodiments, the questions in the first sample dataset include non-programming questions, which are used to instruct the question-answering model to generate content other than program code; the advantage value determination module 302 is used to:
[0141] For any response, if the response is a response generated by the question-answering model for a target non-programming question in the first sample dataset, obtaining a reference response for the target non-programming question;
[0142] If the response matches the reference response, a first reward value is generated for the response, and if the response does not match the reference response, a second reward value is generated for the response, wherein the first reward value is greater than the second reward value.
[0143] In some embodiments, the parameter updating module 304 is configured to:
[0144] determining a first log-likelihood ratio between the current probability of the first response and the reference probability, and determining a second log-likelihood ratio between the current probability of the second response and the reference probability;
[0145] Determine the changing trend of the probability difference between the first response and the second response output by the question-answering model based on the first log-likelihood ratio and the second log-likelihood ratio;
[0146] Update the parameters of the question-answering model based on the changing trend of the probability difference.
[0147] In some embodiments, the parameter updating module 304 is configured to:
[0148] Before updating the parameters of the question-answering model based on the current probability and reference probability of the first response and the second response, the parameters of the question-answering model are updated based on the advantage value, current probability, reference probability of each response and the current probability distribution and reference probability distribution between the responses.
[0149] In some embodiments, the parameter updating module 304 is configured to:
[0150] Determining a direction for updating a first parameter of the question-answering model based on the advantage value, current probability, and reference probability of each response, and determining an update amplitude of the first parameter of the question-answering model based on the current probability distribution and reference probability distribution between the responses;
[0151] Update the parameters of the question-answering model according to the first parameter update direction and the first parameter update amplitude.
[0152] In some embodiments, the first parameter update direction has a first weight, and the first parameter update magnitude has a second weight; the parameter update module 304 is configured to:
[0153] Determining a second parameter update direction and a second parameter update amplitude of the question-answering model according to the first weight, the second weight, the first parameter update direction, and the first parameter update amplitude;
[0154] Update the parameters of the question-answering model according to the second parameter update direction and the second parameter update amplitude.
[0155] In some embodiments, the parameter updating module 304 is configured to:
[0156] Before updating the parameters of the question-answering model based on the advantage value, current probability, reference probability of each response, and the current probability distribution and reference probability distribution between the responses, the question-answering model is initially trained based on the second sample data set.
[0157] In some embodiments, the second sample data set includes multiple training data, and each training data has a corresponding sequence length; the parameter updating module 304 is used to:
[0158] According to the sequence length, multiple training data are divided into multiple groups, and the sequence lengths of training data in different groups are in different length ranges;
[0159] In order of length intervals from small to large, each group of training data is used in turn to perform initial training on the question-answering model.
[0160] For the description of the features in the embodiment corresponding to the model training device, please refer to the relevant description of the embodiment corresponding to the model training method, and no further details will be given here.
[0161] See also Figure 4 An embodiment of the present application also provides an electronic device, including a memory 10 and a processor 20, wherein the memory 10 stores a computer program, and the processor 20 is configured to run the computer program to execute the steps in any one of the above-mentioned model training method embodiments.
[0162] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned model training method embodiments when running.
[0163] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0164] An embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned model training method embodiments.
[0165] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned model training method embodiments.
[0166] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0167] The above is a detailed introduction to a model training method, device, equipment and storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A model training method, characterized in that: The method comprises: In a current training iteration, questions in the first sample data set are input into a question-answering model, and the question-answering model is used to generate a question-response group for each question, each question-response group including multiple responses to the same question and a current probability of the question-answering model generating each response. determining a dominance value for each of the responses, wherein the dominance value represents the accuracy of each of the responses relative to other responses in the set of responses to the question; In each of the question response groups, a first response having a first advantage value and a second response having a second advantage value are respectively screened, wherein the first advantage value is higher than the second advantage value; updating parameters of the question-answering model based on current probabilities of the first response and the second response and reference probabilities, where the reference probability refers to the probability that the question-answering model generated the first response and the second response before the current training iteration; In the next training iteration, the question-answering model whose parameters have been updated in the current training iteration is used to regenerate a question-response set for the question, and the first response and the second response are rescreened based on the regenerated question-response set, and the parameters of the question-answering model are updated based on the current probabilities and reference probabilities of the rescreened first response and the second response; The updating of the parameters of the question-answering model according to the current probabilities and reference probabilities of the first response and the second response includes: determining a first log-likelihood ratio between a current probability of the first response and a reference probability, and determining a second log-likelihood ratio between the current probability of the second response and the reference probability; Determining, based on the first log-likelihood ratio and the second log-likelihood ratio, a trend of a change in the probability difference between the first response and the second response output by the question-answering model; The parameters of the question-answering model are updated according to the changing trend of the probability difference.
2. The method according to claim 1, characterized in that Determining the advantage value of each of the responses includes: generating a reward value for each of the responses, wherein the reward value represents the accuracy of the response; In the same question response group, the advantage value of each response is determined based on the reward value distribution between the responses, wherein, for any response, the advantage value of the response is related to the discreteness of the reward value distribution and the reward value deviation degree of the response, and the reward value deviation degree refers to the distance of the reward value of the response from the center position of the reward value distribution.
3. The method according to claim 2, characterized in that Before determining the advantage value of each of the responses, the method further comprises: In the same question response group, the reward value mean and reward value variance of the responses are determined, and the reward value distribution of the question response group is determined based on the reward value mean and the reward value variance.
4. The method according to claim 2 or 3, characterized in that The questions in the first sample data set include programming questions, and the programming questions are used to instruct the question-answering model to generate program code; Generating a reward value for each of the responses respectively includes: For any of the responses, if the response is a response generated by the question-answering model to a target programming question in the first sample data set, and the response includes target program code, running at least one test case, wherein the test case is used to detect the correctness of the target program code; Based on the running result of the test case, a reward value of the response is determined.
5. The method according to claim 4, characterized in that The determining of the reward value of the response based on the running result of the test case includes: Obtaining the total number of test cases for the target program code and the number of test cases that have passed the test; The reward value of the response is determined according to the proportion of the number of test cases that have passed the run in the total number of test cases, wherein the proportion is proportional to the reward value of the response.
6. The method according to claim 2 or 3, characterized in that The questions in the first sample data set include non-programming questions, and the non-programming questions are used to instruct the question-answering model to generate content other than program code; Generating a reward value for each of the responses respectively includes: For any of the responses, if the response is a response generated by the question-answering model for a target non-programming question in the first sample data set, obtaining a reference response to the target non-programming question; If the response matches the reference response, a first reward value is generated for the response, and if the response does not match the reference response, a second reward value is generated for the response, wherein the first reward value is greater than the second reward value.
7. The method according to claim 1, characterized in that The method further comprises: Before updating the parameters of the question-answering model based on the current probability and reference probability of the first response and the second response, the parameters of the question-answering model are updated based on the advantage value, current probability, reference probability of each response and the current probability distribution and reference probability distribution between the responses.
8. The method according to claim 7, characterized in that The updating of the parameters of the question-answering model according to the advantage value, current probability, reference probability of each response, and the current probability distribution and reference probability distribution between responses includes: Determining a direction for updating a first parameter of the question-answering model based on the advantage value, current probability, and reference probability of each of the responses, and determining an amplitude for updating the first parameter of the question-answering model based on the current probability distribution and reference probability distribution between the responses; Update the parameters of the question-answering model according to the first parameter update direction and the first parameter update amplitude.
9. The method according to claim 8, characterized in that The first parameter update direction has a first weight, and the first parameter update magnitude has a second weight; Updating the parameters of the question-answering model according to the first parameter update direction and the first parameter update amplitude includes: Determining a second parameter update direction and a second parameter update amplitude of the question-answering model according to the first weight, the second weight, the first parameter update direction, and the first parameter update amplitude; Update the parameters of the question-answering model according to the second parameter update direction and the second parameter update amplitude.
10. The method according to claim 7, characterized in that The method further comprises: Before updating the parameters of the question-answering model based on the advantage value, current probability, reference probability of each response and the current probability distribution and reference probability distribution between responses, the question-answering model is initially trained based on the second sample data set.
11. The method according to claim 10, characterized in that The second sample data set includes a plurality of training data, and each training data has a corresponding sequence length; The initial training of the question-answering model based on the second sample data set includes: Dividing the plurality of training data into a plurality of groups according to the sequence lengths, wherein the sequence lengths of the training data in different groups are within different length intervals; The question-answering model is initially trained using each set of training data in order of length intervals from small to large.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.
13. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 11 is implemented.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Data processing method, device and equipment and computer readable storage medium
CN118981520A
Optimization method and device for improving near-end strategy based on language model, and electronic equipment
CN120068993A