Model training method and device, storage medium and program product
By dynamically filtering the positive and negative sample pairs in the training data and updating the model parameters, the problem of mismatch between the training data and the model error distribution is solved, and the model training efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510774535.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-11
AI Technical Summary
During the training process of existing models, the training data does not match the dynamic error distribution of the model, resulting in low training efficiency.
By dynamically filtering the positive and negative sample pairs in the training data during the model training process, ensuring that the training data matches the dynamic error distribution of the model, using the question-and-answer model to generate a question response group, and filtering out the first and second responses based on the advantageous value of the response, and updating the model parameters based on the current probability and reference probability.
It improves the efficiency of model training, shortens the training time, ensures the accuracy of parameter update direction and amplitude, and improves the training effect of the model.
Smart Images

Figure CN120278285A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a model training method, device, storage medium, and program product. Background Art
[0002] Currently, during the training process of some models, usually part of the data in the sample dataset is selected as training data, and the model is trained based on the training data. However, when selecting training data from the sample dataset, it depends on manual judgment or simple random sampling, resulting in the possibility that the selected training data does not match the dynamic error distribution of the model, thereby leading to relatively low model training efficiency. Summary of the Invention
[0003] This application provides a model training method, a model training device, an electronic device, a computer-readable storage medium, and a computer program product to at least solve the problem of relatively low model training efficiency in the related art.
[0004] This application provides a model training method, the method including: In the current training iteration, input the questions in the first sample dataset into the question-and-answer model, where the question-and-answer model is used to generate a question response group for each of the questions, and each question response group includes multiple responses to the same question and the current probability of the question-and-answer model generating each of the responses; Determine the advantage value of each of the responses, where the advantage value characterizes the accuracy of each of the responses relative to other responses in the question response group where it is located; In each question response group, respectively screen a first response with a first advantage value and a second response with a second advantage value, where the first advantage value is higher than the second advantage value; Update the parameters of the question-and-answer model according to the current probability and reference probability of the first response and the second response, where the reference probability refers to the probability that the question-and-answer model generates the first response and the second response before the current training iteration; In the next training iteration, use the question-and-answer model with updated parameters obtained in the current training iteration to regenerate the question response group of the questions, and based on the regenerated question response group, re-screen the first response and the second response, and update the parameters of the question-and-answer model according to the current probability and reference probability of the re-screened first response and second response.
[0005] This application also provides a model training device, the device including: A response generation module, configured to input the questions in the first sample dataset into a question-answering model in the current training iteration. The question-answering model is used to generate a question response group for each of the questions. Each question response group includes multiple responses to the same question and the current probabilities of the question-answering model generating each of the responses; An advantage value determination module, configured to determine the advantage value of each of the responses. The advantage value characterizes the accuracy of each of the responses relative to other responses in the question response group where it is located; A response screening module, configured to respectively screen a first response with a first advantage value and a second response with a second advantage value in each of the question response groups, where the first advantage value is higher than the second advantage value; A first parameter update module, configured to update the parameters of the question-answering model according to the current probabilities and reference probabilities of the first response and the second response. The reference probability refers to the probability that the question-answering model generates the first response and the second response before the current training iteration; A second parameter update module, configured to, in the next training iteration, use the question-answering model with updated parameters obtained in the current training iteration to regenerate the question response group of the question, and respectively screen the first response and the second response according to the regenerated question response group, and update the parameters of the question-answering model according to the current probabilities and reference probabilities of the first response and the second response obtained by the re-screening;
[0006] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above model training methods when executing the computer program.
[0007] This application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above annotation methods are implemented.
[0008] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any of the above annotation methods are implemented.
[0009] In the technical solutions of some embodiments of the present application, during the training process of the question-and-answer model, after generating a question response group for each question by the question-and-answer model, a first response and a second response with different advantage values can be respectively selected from each question response group according to the advantage values of each response. Since the advantage value of the first response is higher than that of the second response, the first response can be regarded as the response of the positive sample pair, and the second response can be regarded as the response of the negative sample pair. Also, since the positive and negative sample pairs are selected from the output responses of the question-and-answer model, the positive and negative sample pairs can match the dynamic error distribution of the question-and-answer model. Furthermore, by combining the current probabilities of the question-and-answer model generating the first response and the second response in the current training iteration, and the reference probabilities of the question-and-answer model generating the first response and the second response before the current training iteration, the parameter update direction and parameter update amplitude of the question-and-answer model can be determined more accurately, thereby shortening the training duration of the question-and-answer model and greatly improving the model training efficiency, and solving the problem of low model training efficiency in some technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more clearly illustrate the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 The flowchart of the model training method provided by some embodiments of the present application; Figure 2 The flowchart of the model training method provided by other embodiments of the present application; Figure 3 The module diagram of the model training device provided by some embodiments of the present application; Figure 4 The module diagram of the electronic device provided by some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.
[0013] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0014] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.
[0015] The dynamic error distribution of a model refers to the distribution state presented by the model's prediction errors for input data at different stages of model training. Specifically, the dynamic error distribution can include the type and quantity distribution of the model's prediction errors. As the model training progresses, the performance of the model will gradually improve, and its dynamic error distribution will also change accordingly. For example, assume that an emotion prediction model needs to be trained to identify the emotion tendency (i.e., positive or negative) in the input text. In the initial stage of training, this emotion prediction model may predict the emotion tendency of all input texts inaccurately. However, by the mid-stage of training, this emotion prediction model can learn to identify some simple emotion expressions and output correct prediction results. For example, it can correctly predict the emotion tendency of 60% of the input texts and incorrectly predict the emotion tendency of the other 40% of the input texts. Here, from the initial stage to the mid-stage of training, the quantity distribution of the prediction errors of this emotion prediction model has changed. Another example is that in the mid-stage of training, this emotion prediction model can correctly predict the emotion tendency of input texts with obvious emotion tendencies and incorrectly predict the emotion tendency of input texts with unobvious emotion tendencies. However, in the late stage of training, this emotion prediction model can correctly predict the emotion tendency of both input texts with obvious and unobvious emotion tendencies. Here, from the initial stage to the mid-stage of training, the type distribution of the prediction errors of this emotion prediction model has changed.
[0016] Furthermore, during the model training process, if the training data distribution matches the dynamic error distribution of the model, the model training efficiency can be greatly improved. Conversely, if the training data distribution does not match the dynamic error distribution of the model, the model training efficiency will be greatly reduced. For example, taking the above-mentioned sentiment prediction model as an example. In the middle of training, since the sentiment prediction model cannot correctly predict the sentiment tendency of input texts with unclear sentiment tendencies, if the training data includes as many input texts with unclear sentiment tendencies as possible, the sentiment prediction model can conduct targeted learning on this type of input text, thereby improving the convergence speed of model training and further greatly improving the model training efficiency. Conversely, if the training data includes fewer input texts with unclear sentiment tendencies, the sentiment prediction model cannot conduct targeted learning on this type of input text, which may lead to a slower convergence speed of model training and may further greatly reduce the model training efficiency.
[0017] Currently, in some technologies, the training data of the model is selected from some publicly available sample datasets. When screening the training data, it often relies on manual judgment or simple random sampling. Therefore, the finally selected training data distribution may not match the dynamic error distribution of the model, resulting in the problem of low model training efficiency. In addition, in these technologies, the training data is static and does not change with the change of the dynamic error distribution of the model. Therefore, even if the training data distribution initially selected from the sample dataset matches the dynamic error distribution of the model, as the model training progresses, there will also be a problem that the training data does not match the dynamic error distribution of the model, thereby resulting in the problem of low model training efficiency.
[0018] In view of this, the present application provides a model training method. By dynamically screening positive and negative sample pairs in the training data during the model training process, the training data distribution can be kept in a state of real-time matching with the dynamic error distribution of the model, thereby improving the model training efficiency. The model training method can be applied to electronic devices. Electronic devices can include, but are not limited to, tablet computers, laptop computers, desktop computers, servers, etc. Referring to Figure 1 , a flowchart of the model training method provided by some embodiments of the present application. Figure 1 In Step S101, in the current training iteration, input the question in the first sample dataset into the question-and-answer model. The question-and-answer model is used to generate a question response group for each question. Each question response group includes multiple responses to the same question and the current probability of the question-and-answer model generating each response.
[0019] In this embodiment, the problems in the first sample dataset may include programming problems and non-programming problems. Among them, programming problems are used to instruct the question-and-answer model to generate program code, and non-programming problems are used to instruct the question-and-answer model to generate content other than program code. Specifically, programming problems can be used to describe program code features, and the program code features may include but are not limited to the logical functions that the program code needs to implement, the writing rules of the program code, etc. The question-and-answer model outputs program code adapted to the programming problem based on the programming problem. For example, the programming problem can be "Please output a program code for user login authentication in Java language". Based on this programming problem, the question-and-answer model can output a program code written in Java language, and running this program code can perform user login authentication. Non-programming problems refer to other problems not related to program code, including but not limited to mathematical problems, geographical problems, historical problems, language problems, economic problems, philosophical problems, daily life problems, etc. For example, the non-programming problem can be "The length of a rectangle is 10 cm and the width is 5 cm. What is its area in square centimeters?". Based on this non-programming problem, the question-and-answer model can output the answer 50.
[0020] The first sample dataset can be a publicly available sample dataset, such as Eurus-2-RL-Data. For any sample data in the first sample dataset, the sample data may include a question, the reasoning process annotated for the question, and the question answer. When training the question-and-answer model, the sample data in the first sample dataset can be divided into multiple batches (i.e., patches), and the question-and-answer model can be iteratively trained batch by batch. For example, in the first sample dataset, every 256 sample data can be divided into a batch. By running the question-and-answer model multiple times or running the question-and-answer model in parallel, the questions and question prompts in the same batch of sample data can be input into the question-and-answer model to obtain the question response group generated by the question-and-answer model for each question.
[0021] In this embodiment, the question prompt can be similar to the following: A conversation between User and Assistant. The user asks a question, and the Assistant solves it. The assistant first thinks about the reasoning process in the mind and then provides the user with the answer. The reasoning process and answer are enclosed within <think>< / think> and <answer>< / answer>tags, respectively, i.e., <think>reasoning process here< / think> <answer>answer here< / answer> . User:prompt. Assistant:。
[0022] In this embodiment, each question response group may include 16 responses to the same question, and the responses in the same question response group are not completely the same. The responses may include the reasoning process of the question-and-answer model answering the question, the answer obtained based on the reasoning process, and the current probability of the question-and-answer model generating the corresponding response. For example, when 16 question-and-answer models are run in parallel and the question Q1 in the sample data A1 is input into these 16 question-and-answer models, the responses generated by each question-and-answer model for the question Q1 can be obtained. Due to the randomness of the question-and-answer model reasoning, the responses generated by different question-and-answer models for the question Q1 may not be completely the same. The 16 responses generated by these 16 question-and-answer models for the question Q1 can be integrated into a question response group, which serves as the question response group generated by the question-and-answer model for the question Q1. For example, the question response group for the question Q1 may be similar to the following: <think>Reasoning Process 1< / think> <answer>Answer 1: The First Current Probability< / answer> ; <think>Reasoning Process 2< / think> <answer>Answer 2: The Second Current Probability< / answer> ; ……; <think>Reasoning Process 16< / think> <answer>Answer 16: The Sixteenth Current Probability< / answer> ; It should be noted that the so-called current probability refers to the probability of the question-and-answer model generating each response based on the current parameters of the question-and-answer model. It can be understood that since the parameters of the question-and-answer model are dynamically changing during the model training process, the probability of the question-and-answer model generating the same response is dynamically changing in different iterative trainings.
[0023] Step S102, determine the advantage value of each response, where the advantage value characterizes the accuracy of each response relative to other responses in the question response group where it is located.
[0024] In this embodiment, the question annotation can be used as the reference response. For the reference response of any question and the responses generated by the question-and-answer model for this question, the advantage value of each response can be determined based on the degree of proximity of each response to the reference response. Specifically, if one response is relatively close to the reference response and the other responses are relatively far from the reference response, the advantage value of this response can be relatively large. If one response is relatively far from the reference response and the other responses are relatively close to the reference response, the advantage value of this response can be relatively small. If the degrees of proximity of multiple responses to the reference response are relatively similar, the advantage values of these multiple responses can be relatively close. For ease of understanding, the following is illustrated by examples.
[0025] For example, assume that the reference response for question Q1 is reference response R, and responses 1 to 16 are included in the question response group of question Q1. If response 1 is the closest to reference response R, the closeness of responses 2 to 10 to reference response R is lower than that of response 1 to reference response R, and the closeness of each of responses 2 to 10 to reference response R is relatively similar, and the closeness of responses 11 to 16 to reference response R is lower than that of responses 2 to 10 to reference response R, then the advantage value of response 1 can be greater than the advantage values of responses 2 to 10, and the advantage values of responses 2 to 10 can be greater than the advantage values of responses 11 to 16, and the advantage values of each of responses 2 to 10 can be the same.
[0026] It can be understood that in the question response group of the same question, the response with a higher advantage value is more accurate, and the response with a lower advantage value is relatively less accurate.
[0027] Step S103, in each question response group, respectively screen a first response with a first advantage value and a second response with a second advantage value, where the first advantage value is higher than the second advantage value.
[0028] Specifically, in this embodiment, the response with the highest advantage value can be screened as the first response, and the response with the lowest advantage value can be screened as the second response in each question response group respectively. In this way, multiple groups of first responses and second responses can be obtained. For example, the first group of first response and second response are screened from question response group 1; the second group of first response and second response are screened from question response group 2. And so on.
[0029] It should be noted that responses that meet the order of advantage value high and low can also be screened as the first response and the second response according to other screening methods. For example, in each question response group respectively, the response with the highest advantage value is screened as the first response, and the response with an advantage value slightly higher than the lowest advantage value is screened as the second response. Another example is that in each question response group respectively, the response with an advantage value slightly lower than the highest advantage value is screened as the first response, and the response with an advantage value slightly higher than the lowest advantage value is screened as the second response.
[0030] Step S104, update the parameters of the question-and-answer model according to the current probability and reference probability of the first response and the second response, where the reference probability refers to the probability that the question-and-answer model generates the first response and the second response before the current training iteration.
[0031] Specifically, for any question response group, the first response screened from the question response group and the question corresponding to the question response group can be regarded as a positive sample pair, and the second response screened from the question response group and the question corresponding to the question response group can be regarded as a negative sample pair. It can be understood that the purpose of model training is to increase the probability of the question-answering model generating the first response and decrease the probability of the question-answering model generating the second response. Also, since the reference probability is the probability of the question-answering model generating the first response and the second response before the current training iteration, the parameters of the question-answering model can be updated along the direction that the current probability of the first response is greater than the reference probability and the current probability of the second response is lower than the reference probability. In this way, it can be ensured that the parameters of the question-answering model are updated in the correct direction, thereby reducing the problem of increased model training time caused by incorrect parameter updates, and further greatly improving the training efficiency of the question-answering model.
[0032] Step S105, in the next training iteration, use the question-answering model with updated parameters obtained in the current training iteration to regenerate the question response group of the question, and based on the regenerated question response group, re-screen the first response and the second response, and update the parameters of the question-answering model according to the current probabilities and reference probabilities of the re-screened first response and second response.
[0033] Briefly, steps S101 to S104 can be steps executed in multiple iterative loops. For example, when steps S101 to S104 are executed for the first time, the parameters of the question-answering model are updated for the first time. Based on the question-answering model after the first parameter update, steps S101 to S104 can be executed for the second time to update the parameters of the question-answering model for the second time. And so on. The number of times of loop execution can be 50 times, 100 times, etc.
[0034] It can be understood that since in each iterative training, the first response and the second response are re-screened based on the current dynamic error distribution of the question-answering model, the first response and the second response can reflect the current dynamic error distribution of the question-answering model. Furthermore, based on the current probabilities and reference probabilities of the first response and the second response in each iterative training, the model training efficiency can be greatly improved.
[0035] In summary, in the technical solutions of some embodiments of the present application, during the training process of the question-and-answer model, after generating a question response group for each question by the question-and-answer model, the first response and the second response with different superiority values can be respectively selected from each question response group according to the superiority values of each response. Since the superiority value of the first response is higher than that of the second response, the first response can be regarded as the response of the positive sample pair, and the second response can be regarded as the response of the negative sample pair. Also, since the positive and negative sample pairs are selected from the output responses of the question-and-answer model, the positive and negative sample pairs can match the dynamic error distribution of the question-and-answer model. Furthermore, by combining the current probabilities of the question-and-answer model generating the first response and the second response in the current training iteration, and the reference probabilities of the question-and-answer model generating the first response and the second response before the current training iteration, the parameter update direction of the question-and-answer model can be determined more accurately, thereby shortening the training duration of the question-and-answer model and greatly improving the model training efficiency, and solving the problem of low model training efficiency in some technologies.
[0036] The following further elaborates on the above steps S102 and S104.
[0037] In some embodiments, determining the superiority values of each response in the above step S102 may include: Generating a reward value for each response respectively, where the reward value represents the accuracy of the response; In the same question response group, according to the distribution of the reward values among the responses, determine the superiority values of each response. Among them, for any response, the superiority value of the response is related to the degree of dispersion of the reward value distribution and the deviation degree of the reward value of the response. The deviation degree of the reward value refers to the distance of the reward value of the response from the central position of the reward value distribution.
[0038] Specifically, the higher the accuracy of the response, the higher its reward value can be. Conversely, the lower the accuracy of the response, the lower its reward value can be.
[0039] In the same question response group, if the degree of dispersion of the reward value distribution is relatively low (i.e., the reward values are relatively concentrated), and the reward value of a response is greater than the central position of the reward value distribution, then the superiority value of this response can be higher. Conversely, if the reward value of a response is less than the central position of the reward value distribution, then the superiority value of this response can be relatively low.
[0040] In the same question response group, when the degree of dispersion of the reward value distribution is relatively high (i.e., the reward value distribution area is relatively wide), if there are multiple responses with reward values greater than the central position of the reward value distribution, the dominance values of these multiple responses do not need to be relatively high (because there are multiple responses with reward values greater than the central position, so these responses are not particularly dominant). Similarly, if there are multiple responses with reward values less than the central position of the reward value distribution, the dominance values of these multiple responses do not need to be relatively low.
[0041] Specifically, in the same question response group, if the degree of dispersion of the reward value distribution is relatively high and there are relatively few responses with reward values greater than the central position of the reward value distribution, and a response has a reward value greater than the central position of the reward value distribution, the dominance value of this response can be relatively high. Similarly, in the same question response group, if the degree of dispersion of the reward value distribution is relatively high and there are relatively few responses with reward values less than the central position of the reward value distribution, and a response has a reward value less than the central position of the reward value distribution, the dominance value of this response can be relatively low.
[0042] In the above embodiments, based on the reward value distribution in the same question response group and the deviation degree of the reward value of the response, the dominance value of the response is determined, which can quantify the dominance value and ensure the calculation accuracy of the dominance value.
[0043] Specifically, in some embodiments, in the same question response group, the mean value and variance of the reward values of the responses can be determined, and based on the mean value and variance of the reward values, the reward value distribution of this question response group can be determined. Among them, the mean value of the reward values can represent the central position of the reward value distribution, and the variance of the reward values can represent the degree of dispersion of the reward value distribution. The larger the variance of the reward values, the higher the degree of dispersion of the reward value distribution. Conversely, the smaller the variance of the reward values, the lower the degree of dispersion of the reward value distribution. In this way, the reward value distribution of each question response group is determined by mathematical statistical methods, and the accuracy and credibility are relatively high.
[0044] In some embodiments, for any question response group, after obtaining the variance and mean value of the reward values corresponding to this question response group, the dominance value of each response in this question response group can be determined based on Expression (1):
[0045] Among them, represents the dominance value of the i-th response in this question response group, represents the reward value of the i-th response in this question response group, represents the mean value of the reward values of this question response group, represents the variance of the reward values of this question response group.
[0046] The above expression (1) can also be regarded as normalizing the reward values of each response. In this way, the dimensional differences of the reward values of different responses can be eliminated, and the comparability between the reward values can be improved.
[0047] In some embodiments, when the first sample data set includes programming problems, generating reward values for each response respectively may include: For any response, if the response is a response generated by the question-and-answer model for the target programming problem in the first sample data set and the response includes the target program code, at least one test case is run, and the test case is used to detect the correctness of the target program code; Based on the running result of the test case, determine the reward value of the response.
[0048] Specifically, in the first sample data set, at least one test case for each programming problem can be marked. After the question-and-answer model generates a response for the target programming problem, the test case marked for the target programming problem can be run to detect whether the target program code generated by the question-and-answer model is correct. For example, if the test case runs successfully, it means that the target program code is correct, and the reward value of the response can be the third reward value. On the contrary, if the test case fails to run, it means that the target program code is incorrect, and the reward value of the response can be the fourth reward value. The third reward value can be greater than the fourth reward value. For example, the third reward value can be 1, and the fourth reward value can be 0. By running the test case to detect the correctness of the target program code, the logical accuracy of the target program code can be verified more accurately.
[0049] Furthermore, when there are multiple marked test cases for the target programming problem, there may be a problem that some test cases run successfully and some test cases fail to run. In this case, it is impossible to determine the reward value of the response based on whether the test case runs successfully. In view of this, in some embodiments, the above determining the reward value of the response based on the running result of the test case may include: Obtain the total number of test cases for the target program code and the number of test cases that run successfully; Determine the reward value of the response according to the proportion of the number of test cases that run successfully in the total number of test cases, where the proportion is directly proportional to the reward value of the response.
[0050] In this way, when only some test cases run successfully, the reward value of the response can still be determined, improving the applicability of the solution.
[0051] Further, in some embodiments, the importance levels of different test cases may vary. For example, assume that the target programming problem has two labeled test cases, Test Case A and Test Case B. Among them, when Test Case A fails to pass, it indicates that the overall logic of the target program code is incorrect. When Test Case B fails to pass, it indicates that the value range of parameter A in the target program code is out of range. Obviously, the importance level of Test Case A will definitely be higher than that of Test Case B. When Test Case A fails to pass, it can indicate that the target program code is incorrect, and the corresponding reward value can be relatively low. However, when Test Case B fails to pass, it can indicate that only part of the target program code is incorrect, and the corresponding reward value can be relatively high. In view of this, each test case can also have its own corresponding weight. Among them, test cases with a high importance level can have a higher weight, and test cases with a low importance level can have a lower weight. When calculating the reward value, the weights of each test case can be referred to, so that the corresponding reward value can be proportional to the importance level of each test case, thus ensuring the accuracy of the reward value.
[0052] In some embodiments, when the first sample data set includes non-programming problems, generating reward values for each response respectively may include: For any response, if the response is generated by the question-and-answer model for the target non-programming problem in the first sample data set, obtain the reference response of the target non-programming problem; If the response matches the reference response, generate a first reward value for the response. If the response does not match the reference response, generate a second reward value for the response, where the first reward value is greater than the second reward value.
[0053] Specifically, the reference response is the reasoning process and question result annotated for the target non-programming problem. Comparing the response generated by the question-and-answer model with the reference response and generating a reward value according to the comparison result conforms to the conventional training process of model training.
[0054] In some embodiments, a reward model can be pre-trained. After the question-and-answer model generates a response to a question, the response can be input into the trained reward model to obtain the reward values of each response. In this way, the calculation process of the reward value can be reduced, and the model training efficiency can be improved.
[0055] In some embodiments, the step of updating the parameters of the question-and-answer model based on the current probabilities and reference probabilities of the first response and the second response in step S104 may include: Determine the first log-likelihood ratio between the current probability and the reference probability of the first response, and determine the second log-likelihood ratio between the current probability and the reference probability of the second response; Determine the changing trend of the probability difference between the first response and the second response output by the question-answering model based on the first log-likelihood ratio and the second log-likelihood ratio; Update the parameters of the question-answering model according to the changing trend of the probability difference.
[0056] Specifically, a first loss function as shown in expression (2) can be constructed based on the Direct Preference Optimization (DPO) framework.
[0057]
[0058] Among them, represents the question-answering model of the current training iteration, represents the question-answering model before the current training iteration, X represents the question, represents the first response, represents the second response, represents the current probability of the question-answering model of the current training iteration generating the first response, represents the reference probability of the question-answering model before the current training iteration generating the first response, represents the first log-likelihood ratio between the current probability and the reference probability of the first response, represents the current probability of the question-answering model of the current training iteration generating the second response, represents the reference probability of the question-answering model before the current training iteration generating the second response, represents the second log-likelihood ratio between the current probability and the reference probability of the second response, represents the changing trend of the probability difference between the question-answering model outputting the first response and the second response, is the temperature coefficient (such as 0.1), represents the base of the logarithmic function, represents the first loss value.
[0059] As can be seen from Equation (2), the first log-likelihood ratio can represent the relative change in the probabilities of the Q&A model outputting the first response and the second response at the current training iteration, and the second log-likelihood ratio can represent the relative change in the probabilities of the Q&A model outputting the first response and the second response before the current training iteration. For example, before the current training iteration, the probability of the Q&A model outputting the first response is 0.6, and the probability of outputting the second response is 0.4, with a probability change of 0.2. At the current training iteration, the probability of the Q&A model outputting the first response is 0.7, and the probability of outputting the second response is 0.3, with a probability change of 0.4. The change trend between the probability difference before the current training iteration and the probability difference at the current training iteration is the probability difference change trend, such as changing from 0.4 to 0.2. It can be understood that the purpose of model training is to maximize the probability difference between the first response and the second response output by the Q&A model. At the same time, the probability difference change trend can also be maximized. Also, since there is a negative sign in front of the first loss function, the meaning of Equation (2) is that when the value of the first loss value is less than the first threshold, it indicates that the training of the Q&A model is completed.
[0060] In the above embodiment, based on the first log-likelihood ratio and the second log-likelihood ratio, the change trend of the probability difference between the first response and the second response output by the Q&A model is determined, and the parameters of the Q&A model are updated according to the change trend of the probability difference, which can ensure that the parameters of the Q&A model are updated in the correct direction and improve the model training efficiency.
[0061] Referring to Figure 2 , in some embodiments, the model training method of the present application may specifically include the following steps: 1) Construct a second sample data set. Specifically, the second sample data set can be a multi-source heterogeneous open-source inference data set including multiple pieces of training data. The multiple pieces of training data can cover three core fields: mathematics, code generation, and logical reasoning. The training data related to mathematics is centered around the sample data set Open-R1-Math-220k, and the training data related to code generation and logical reasoning can come from the sample data set OpenThoughts-114k. After obtaining the sample data from the sample data sets Open-R1-Math-220k and OpenThoughts-114k, the sample data in the second sample data set can be sorted in the format of problem description, reasoning process, solution method, standard answer, data source, data type, and test case, that is, each piece of sample data in the second sample data set includes content such as problem description, reasoning process, solution method, standard answer, data source, data type, and test case.
[0062] In some embodiments, it is also possible to deduplicate and filter the quality of the sample data in the second sample dataset. For example, if the similarity of the problem descriptions of multiple sample data is relatively high, only one of the sample data can be retained, and other similar sample data can be deleted.
[0063] 2) Based on the second sample dataset, perform initial training on the question-and-answer model. The so-called initial training refers to the supervised fine-tuning of the question-and-answer model based on the second sample dataset.
[0064] Specifically, a model training environment can be constructed using 8 graphics processors. The model of the graphics processor can be NVIDIA A100, and each graphics processor can have 80GB of video memory. During model training, the optimization strategy of data parallelism + ZeRO-3 can be adopted to efficiently utilize the video memory through the DeepSpeed acceleration framework. The maximum sequence length that a single graphics processor can carry is 32k tokens. When processing training data with an ultra-long sequence (>16k), the gradient checkpoint technology can be enabled. The video memory occupancy can be controlled within 65GB / card. The mixed-precision training can adopt the BF16 mode. In this way, compared with FP32, it can save about 40% of the video memory overhead while maintaining numerical stability.
[0065] In some embodiments, each piece of training data in the second sample dataset has its corresponding sequence length. The so-called sequence length is the number of tokens included in the question of the training data. The above initial training of the question-and-answer model based on the second sample dataset may include: Divide multiple pieces of training data into multiple groups according to the sequence length, and the sequence lengths of the training data in different groups are in different length interval ranges; In the order from small to large of the length interval ranges, use the training data of each group to perform initial training on the question-and-answer model in turn.
[0066] For example, according to the sequence length, multiple pieces of training data can be divided into three groups: 4k - 8k, 8k - 16k, and 16k - 32k. That is, the sequence length of the training data in the first group is between 4k and 8k, the sequence length of the training data in the second group is between 8k and 16k, and the sequence length of the training data in the third group is between 16k and 32k. In the first training stage, the training data in the first group can be used to train the question - answering model. And the single - card batch size can be set to 4, and the global batch size can be extended to 256 through gradient accumulation (gradient accumulation step = 8). The initial learning rate can be set to 1e - 5, combined with 2000 - step linear warm - up. In the second training stage, the training data in the second group can be used to train the question - answering model. And the single - card batch size can be reduced to 2, the global batch size can be maintained at 128 (gradient accumulation step = 8), the learning rate can be reduced to 5e - 6, and cosine annealing scheduling can be adopted. In the third training stage, the training data in the third group is used to train the question - answering model. And the single - card batch size can be compressed to 1, the global batch size is maintained at 64 by accumulating 8 steps, and the learning rate can be kept at 2e - 6. In this way, gradient oscillation during training can be avoided.
[0067] Furthermore, to balance video memory and training efficiency, a method of selectively updating parameters can be adopted. For example, assuming that the question - answering model has 28 layers, only the last 14 layers of the question - answering model can be opened for training, and the first 14 layers and the position encoding parameters are completely frozen.
[0068] 3) Execute the above - mentioned step S101 and step S102.
[0069] 4) Update the parameters of the question - answering model according to the advantage values, current probabilities, reference probabilities of each response, and the current probability distribution and reference probability distribution among the responses.
[0070] Specifically, the first parameter update direction of the question - answering model can be determined according to the advantage values, current probabilities, and reference probabilities of each response, and the first parameter update amplitude of the question - answering model can be determined according to the current probability distribution and reference probability distribution among the responses. Update the parameters of the question - answering model according to the first parameter update direction and the first parameter update amplitude.
[0071] Specifically, a second loss function as shown in expression (3) can be constructed based on the advantage values, current probabilities, and reference probabilities of each response, and the first parameter update direction of the question - answering model can be determined according to the second loss function.
[0072]
[0073] In expression (3), each response output by the question-and-answer model is divided into multiple time steps according to the number of tokens. For example, if response A includes 3 tokens, then the first token is the first time step, the second token is the second time step, and the third token is the third time step. The advantage value of each time step is the same and is the advantage value of the response. For example, assuming the advantage value of the above response A is 20, then the advantage values of the above first time step, second time step, and third time step are all 20.
[0074] Based on the above description, in expression (3), represents the t-th time step of the i-th response, represents the model parameter state of the question-and-answer model at the t-th time step, represents the advantage value of the t-th time step of the i-th response, T represents the number of time steps of each question-response group, and N represents the number of question-response groups. represents the question-and-answer model of the current training iteration, represents the question-and-answer model before the current training iteration, represents the state of the question-and-answer model in the current training iteration to generate probability, represents the state of the question-and-answer model before the current training iteration to generate probability, represents the log ratio of the probabilities of the question-and-answer model generating in the current training iteration and before the current training iteration, represents the goodness or badness of generating in the state, and this goodness or badness can be used to determine the first parameter update direction of the question-and-answer model. and functions are used to ensure that the parameter update is not too radical, that is, to ensure the stable update of the model parameters. represents the second loss value.
[0075] Furthermore, based on the current probability distribution and the reference probability distribution among the responses, a third loss function as shown in expression (4) can be constructed, and based on the third loss function, the first parameter update amplitude of the question-and-answer model can be determined.
[0076]
[0077] Specifically, represents the third loss value, represents the probability distribution of all responses output by the question-and-answer model in the state of the current training iteration, indicating the probability distribution of all responses output by the Q&A model before the current training iteration The KL divergence, also known as the Kullback-Leibler divergence, is used to measure the difference between two probability distributions. It is used to quantify the difference between two probability distributions. During the model training process, if the value is less than the first threshold, the sum of the differences between the two probability distributions of the question-response group needs to be less than the second threshold. In this way, the update amplitude of the first parameter can be restricted to avoid the problem of unstable model training caused by excessive parameter changes.
[0078] In some embodiments, the first parameter update direction has a first weight, and the first parameter update amplitude has a second weight; Update the parameters of the Q&A model according to the first parameter update direction and the first parameter update amplitude, including: Determine the second parameter update direction and the second parameter update amplitude of the Q&A model according to the first weight, the second weight, the first parameter update direction, and the first parameter update amplitude; Update the parameters of the Q&A model according to the second parameter update direction and the second parameter update amplitude.
[0079] Specifically, the fourth loss function shown in expression (5) can be constructed according to the first weight, the second weight, the first parameter update direction, and the first parameter update amplitude.
[0080]
[0081] Among them, is the second weight of the first parameter update amplitude The second weight of the first parameter update direction is 1. When training the Q&A model based on the loss function shown in expression (5), the training accuracy of the model can be guaranteed.
[0082] 5) Execute the above steps S103 and S104.
[0083] During the model training process, by repeatedly executing the above steps 3) to 5), the Q&A model can be iteratively trained. During the iterative training process, the dynamic error distribution of the Q&A model can be dynamically changed. Since the first response and the second response used for model parameter update in this application are selected from the output results of the Q&A model, the first response and the second response will also change with the change of the dynamic error distribution, that is, the training data distribution of the Q&A model can be kept in real-time matching with the dynamic error distribution of the Q&A model. Therefore, the model training efficiency can be greatly improved.
[0084] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.
[0085] An embodiment of the present application also provides a model training device. Referring to Figure 3 , it is a schematic diagram of the modules of the model training device provided by some embodiments of the present application. Figure 3 In , the model training device includes: A response generation module 301, in the current training iteration, inputs the questions in the first sample dataset into the question and answer model. The question and answer model is used to generate a question response group for each question respectively. Each question response group includes multiple responses to the same question and the current probabilities of the question and answer model for generating each response; An advantage value determination module 302, which determines the advantage value of each response. The advantage value characterizes the accuracy of each response relative to other responses in the question response group where it is located; A response screening module 303, in each question response group, respectively screens a first response with a first advantage value and a second response with a second advantage value, and the first advantage value is higher than the second advantage value; A parameter update module 304, which updates the parameters of the question and answer model according to the current probabilities and reference probabilities of the first response and the second response. The reference probability refers to the probability that the question and answer model generates the first response and the second response before the current training iteration;
[0086] In some embodiments, the advantage value determination module 302 is used for: Generating a reward value for each response respectively. The reward value characterizes the accuracy of the response; In the same question response group, according to the distribution of the reward values among the responses, the advantage value of each response is determined. Among them, for any response, the advantage value of the response is related to the degree of dispersion of the reward value distribution and the degree of deviation of the reward value of the response. The degree of deviation of the reward value refers to the distance of the reward value of the response from the central position of the reward value distribution.
[0087] In some embodiments, before determining the advantage value of each response, the advantage value determination module 302 is used for: In the same question response group, determine the mean and variance of the reward values of the responses, and based on the mean and variance of the reward values, determine the reward value distribution of the question response group.
[0088] In some embodiments, the questions in the first sample dataset include programming questions, and the programming questions are used to instruct the Q&A model to generate program code; the advantage value determination module 302 is configured to: For any response, if the response is a response generated by the Q&A model for the target programming question in the first sample dataset and the response includes the target program code, then run at least one test case, and the test case is used to detect the correctness of the target program code; Based on the running result of the test case, determine the reward value of the response.
[0089] In some embodiments, the advantage value determination module 302 is configured to: Obtain the total number of test cases for the target program code and the number of test cases that passed the run; Determine the reward value of the response according to the proportion of the number of test cases that passed the run in the total number of test cases, where the proportion is proportional to the reward value of the response.
[0090] In some embodiments, the questions in the first sample dataset include non-programming questions, and the non-programming questions are used to instruct the Q&A model to generate content other than program code; the advantage value determination module 302 is configured to: For any response, if the response is a response generated by the Q&A model for the target non-programming question in the first sample dataset, then obtain the reference response of the target non-programming question; If the response matches the reference response, generate a first reward value for the response, and if the response does not match the reference response, generate a second reward value for the response, where the first reward value is greater than the second reward value.
[0091] In some embodiments, the parameter update module 304 is configured to: Determine the first log-likelihood ratio between the current probability and the reference probability of the first response, and determine the second log-likelihood ratio between the current probability and the reference probability of the second response; Based on the first log-likelihood ratio and the second log-likelihood ratio, determine the probability difference change trend of the Q&A model outputting the first response and the second response; Update the parameters of the Q&A model according to the probability difference change trend.
[0092] In some embodiments, the parameter update module 304 is configured to: Before updating the parameters of the Q&A model based on the current probabilities and reference probabilities of the first response and the second response, update the parameters of the Q&A model based on the advantage values, current probabilities, reference probabilities of each response, and the current probability distribution and reference probability distribution among the responses.
[0093] In some embodiments, the parameter update module 304 is configured to: Determine a first parameter update direction of the Q&A model based on the advantage values, current probabilities, and reference probabilities of each response, and determine a first parameter update amplitude of the Q&A model based on the current probability distribution and reference probability distribution among the responses; Update the parameters of the Q&A model according to the first parameter update direction and the first parameter update amplitude.
[0094] In some embodiments, the first parameter update direction has a first weight, and the first parameter update amplitude has a second weight; the parameter update module 304 is configured to: Determine a second parameter update direction and a second parameter update amplitude of the Q&A model according to the first weight, the second weight, the first parameter update direction, and the first parameter update amplitude; Update the parameters of the Q&A model according to the second parameter update direction and the second parameter update amplitude.
[0095] In some embodiments, the parameter update module 304 is configured to: Before updating the parameters of the Q&A model based on the advantage values, current probabilities, reference probabilities of each response, and the current probability distribution and reference probability distribution among the responses, perform initial training on the Q&A model according to a second sample data set.
[0096] In some embodiments, the second sample data set includes multiple pieces of training data, and each piece of training data has its own corresponding sequence length; the parameter update module 304 is configured to: Divide the multiple pieces of training data into multiple groups according to the sequence length, and the sequence lengths of the training data in different groups are in different length interval ranges; Perform initial training on the Q&A model using each group of training data in turn according to the order of the length interval ranges from small to large.
[0097] For the description of the features in the embodiments corresponding to the model training device, reference may be made to the relevant description in the embodiments corresponding to the model training method, which will not be elaborated here one by one.
[0098] With reference to Figure 4 , an embodiment of the present application further provides an electronic device, including a memory 10 and a processor 20. A computer program is stored in the memory 10, and the processor 20 is configured to run the computer program to execute the steps in any one of the above model training method embodiments.
[0099] Embodiments of the present application also provide a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps in any of the above-described model training method embodiments when running.
[0100] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), external hard drives, magnetic disks, or optical discs that can store computer programs.
[0101] Embodiments of the present application also provide a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps in any of the above-described model training method embodiments.
[0102] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above-described model training method embodiments.
[0103] Those skilled in the art can further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0104] The above has introduced in detail a model training method, apparatus, device, and storage medium provided by the present application. Specific examples are used herein to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A model training method, characterized in that, The method includes: In the current training iteration, input the questions in the first sample dataset into the question-answering model, where the question-answering model is used to generate a group of question responses for each of the questions, and each group of question responses includes multiple responses to the same question and the current probabilities of the question-answering model for generating each of the responses; Determine the advantage value of each of the responses, where the advantage value characterizes the accuracy of each response relative to other responses in the group of question responses where it is located; In each group of question responses, respectively screen out a first response with a first advantage value and a second response with a second advantage value, where the first advantage value is higher than the second advantage value; Update the parameters of the question-answering model according to the current probabilities and reference probabilities of the first response and the second response, where the reference probability refers to the probability of the question-answering model for generating the first response and the second response before the current training iteration; In the next training iteration, use the question-answering model with updated parameters obtained in the current training iteration to regenerate the group of question responses for the questions, and respectively rescreen the first response and the second response according to the regenerated group of question responses, and update the parameters of the question-answering model according to the current probabilities and reference probabilities of the first response and the second response obtained after the rescreening.
2. The method according to claim 1, wherein The determining the advantage value of each of the responses includes: Generate a reward value for each of the responses respectively, where the reward value characterizes the accuracy of the response; In the same group of question responses, determine the advantage value of each of the responses according to the distribution of the reward values among the responses. Among them, for any one of the responses, the advantage value of the response is related to the degree of dispersion of the reward value distribution and the degree of deviation of the reward value of the response. The degree of deviation of the reward value refers to the distance of the reward value of the response from the central position of the reward value distribution.
3. The method according to claim 2, wherein Before determining the advantage value of each of the responses, the method further includes: In the same group of question responses, determine the mean value and variance of the reward values of the responses, and determine the reward value distribution of the group of question responses according to the mean value and variance of the reward values.
4. The method according to claim 2 or 3, characterized in that, The questions in the first sample dataset include programming questions, and the programming questions are used to instruct the question-answering model to generate program codes; The generating a reward value for each of the responses respectively includes: For any one of the responses, if the response is a response generated by the question-answering model for the target programming question in the first sample dataset and the response includes the target program code, then run at least one test case, and the test case is used to detect the correctness of the target program code; Determine the reward value of the response based on the running result of the test case.
5. The method according to claim 4, wherein The determining the reward value of the response based on the running result of the test case includes: Obtain the total number of test cases for the target program code and the number of test cases that passed the running; Determine the reward value of the response according to the proportion of the number of test cases that passed the running in the total number of test cases, where the proportion is proportional to the reward value of the response.
6. The method according to claim 2 or 3, characterized in that The problems in the first sample dataset include non-programming problems, which are used to indicate that the question-answering model generates content other than program code; Respectively generating a reward value for each of the responses includes: For any one of the responses, if the response is a response generated by the question-answering model for a target non-programming problem in the first sample dataset, obtain the reference response of the target non-programming problem; If the response matches the reference response, generate a first reward value for the response; if the response does not match the reference response, generate a second reward value for the response, where the first reward value is greater than the second reward value.
7. The method according to claim 1, wherein Updating the parameters of the question-answering model according to the current probabilities and reference probabilities of the first response and the second response includes: Determining a first log-likelihood ratio between the current probability and the reference probability of the first response, and determining a second log-likelihood ratio between the current probability and the reference probability of the second response; Determining the probability difference change trend of the question-answering model outputting the first response and the second response according to the first log-likelihood ratio and the second log-likelihood ratio; Updating the parameters of the question-answering model according to the probability difference change trend.
8. The method according to claim 1 or 7, characterized in that, The method further includes: Before updating the parameters of the question-answering model according to the current probabilities and reference probabilities of the first response and the second response, update the parameters of the question-answering model according to the advantage values, current probabilities, reference probabilities of each response, and the current probability distribution and reference probability distribution between responses.
9. The method according to claim 8, wherein Updating the parameters of the question-answering model according to the advantage values, current probabilities, reference probabilities of each response, and the current probability distribution and reference probability distribution between responses includes: Determining a first parameter update direction of the question-answering model according to the advantage values, current probabilities and reference probabilities of each response, and determining a first parameter update amplitude of the question-answering model according to the current probability distribution and reference probability distribution between responses; Updating the parameters of the question-answering model according to the first parameter update direction and the first parameter update amplitude.
10. The method according to claim 9, wherein The first parameter update direction has a first weight, and the first parameter update amplitude has a second weight; Updating the parameters of the question-answering model according to the first parameter update direction and the first parameter update amplitude includes: Determining a second parameter update direction and a second parameter update amplitude of the question-answering model according to the first weight, the second weight, the first parameter update direction and the first parameter update amplitude; Updating the parameters of the question-answering model according to the second parameter update direction and the second parameter update amplitude.
11. The method according to claim 8, wherein The method further includes: Before updating the parameters of the question-answering model according to the advantage values, current probabilities, reference probabilities of each response, and the current probability distribution and reference probability distribution between responses, initially train the question-answering model according to a second sample dataset.
12. The method according to claim 11, wherein The second sample dataset includes multiple pieces of training data, and each piece of training data has its own corresponding sequence length; Performing initial training on the Q&A model according to the second sample data set includes: Dividing the multiple pieces of training data into multiple groups according to the sequence length, where the sequence lengths of the training data in different groups are in different length interval ranges; Sequentially performing initial training on the Q&A model using the training data of each group in ascending order of the length interval ranges.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 12 is implemented.
14. An electronic device, characterized in that, The electronic device includes a processor and a memory. The memory is used to store a computer program, and when the computer program is executed by the processor, the method described in any one of claims 1 to 12 is implemented.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the method described in any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Method and device for model training, equipment and storage medium
CN117689003A
Data processing method, device and equipment and computer readable storage medium
CN118981520A
Large language model self-evaluation method and device, electronic equipment and storage medium
CN119337944A
Optimization method and device for improving near-end strategy based on language model, and electronic equipment
CN120068993A