A method for updating a teacher model based on accuracy and related devices

Through the method of dynamically updating the teacher model, combining the accuracy of the student model and the distillation loss of the teacher model, the problem of unreasonable update of the teacher model in the existing technology is solved, and the learning effect and overall accuracy of the student model are improved.

CN115796270BActive Publication Date: 2025-07-22SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211439358.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2025-07-22
Estimated Expiration
2042-11-17

AI Technical Summary

Technical Problem

In the prior art, the update method of teacher model fails to effectively integrate the accuracy of the classification model, resulting in the inability to take into account the stability and performance of the teacher model, affecting the learning effect of the student model.

Method used

By obtaining the training set, validation set and test set, initialize the relevant parameters, calculate the cross entropy loss of the student model and the distillation loss of the teacher model, and backpropagate the student model in the direction of global loss, use the verification set to test the accuracy of the student model, and dynamically update the teacher model to ensure that the accuracy difference is greater than the threshold.

Benefits of technology

It improves the learning effect of the student model in the self-distillation scenario, takes into account the stability and performance of the teacher model, and improves the overall accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115796270B_ABST
    Figure CN115796270B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for updating a teacher model based on accuracy and related devices, mines the model performance (accuracy) as a judgment condition, constructs a more flexible method for updating the teacher model, and further improves the performance of the student model in the self-distillation scenario. The present invention incorporates the performance of the teacher model (such as the accuracy of the classification model) into the decision-making process of the teacher model, and uses the calculated minimum performance difference as the condition for replacing the teacher model, taking into account both the stability and performance of the teacher model, and further improving the learning effect of the student model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a teacher model update method, system, terminal, and computer-readable storage medium based on accuracy rate. Background Art

[0002] Since the pre-trained language models such as BERT (Bidirectional Encoder Representation from Transformers) were proposed, the accuracy rate of tasks such as text classification has been greatly improved. Similar to sequence labeling tasks such as Chinese word segmentation, the text classification framework also includes three main parts: the Embedding part (information embedding), which converts the input characters into a string of information that is distributed and can be understood and computed by a computer; the Encoder part (encoder), which encodes the Embedding information and enables the computer to capture the relationships between each Embedding through specific operations; and the Docoder part (decoder), which decodes the encoded information and then restores it into character information that can be understood by humans.

[0003] Although the existing text classification models already have a high accuracy rate, when additional unlabeled data is used to assist in training, the performance of the same model can be further improved. This indicates that the parameters of the existing text classification models can be further optimized, and the phenomenon that the self-distillation (self-distillation means that without adding a new large model, a teacher model can be found, which can also provide effective gain information to the student model. Here, the teacher model is often not more complex than the student model, but the gain information provided is effective incremental information for the student model to improve the efficiency of the student model) method produces a better model also verifies this point; in the self-distillation scenario, the quality of the teacher model's performance directly affects the learning effect of the student model, and the two are not simply linearly related. Therefore, in the dynamic change stage of model training, it is particularly important to quantify the selection conditions of the teacher.

[0004] In existing research, there are two sources of the teacher model. One is to use the student model in the previous iteration process, as Figure 1 shown, and the other is to use the student model with the best performance in all previous iteration processes. The latter is more stable, and the experimental effect is slightly better than the former. However, both of the above methods for selecting the teacher model have the same problem, that is, the update of the teacher model and the student model is synchronous. To avoid that in the iteration process, a student model with a slightly improved accuracy rate is updated into a teacher model, some researchers manually specify a minimum update threshold to ignore the student models with slightly improved performance.

[0005] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0006] The main objective of the present invention is to provide a method, system, terminal, and computer-readable storage medium for updating a teacher model based on accuracy, aiming to solve the problem that the existing technology does not incorporate the accuracy of the classification model into the decision-making process of the teacher model, and thus cannot balance the stability and performance of the teacher model.

[0007] To achieve the above objective, the present invention provides a method for updating a teacher model based on accuracy, and the method for updating a teacher model based on accuracy includes the following steps:

[0008] Obtain a training set, a validation set, and a test set, where the validation set is randomly selected from the training set;

[0009] Initialize training-related parameters, where the training-related parameters include an initialized balance factor and a threshold coefficient;

[0010] Obtain the prediction information of the student model and calculate the cross-entropy loss of the student model;

[0011] If there is a teacher model, obtain the prediction information of the teacher model, calculate the distillation loss of the teacher model, and incorporate the distillation loss into the global loss;

[0012] Update the student model in the direction of reducing the global loss through backpropagation;

[0013] When the student model has learned all of the training set, use the validation set to test the current student model and obtain the accuracy of the current student model;

[0014] If the accuracy of the current student model is greater than or equal to the accuracy corresponding to the student model with the best historical performance during the training process, save the current student model, and at the same time update the accuracy corresponding to the student model with the best historical performance during the training process and the degree of improvement in the performance of the current student model during the training process;

[0015] If the degree of improvement in the performance of the current student model during the updated training process is greater than or equal to the minimum accuracy update threshold, update the teacher model, and at the same time update the accuracy of the current teacher model and the minimum accuracy update threshold.

[0016] Optionally, in the method for updating a teacher model based on accuracy, the distillation loss uses KL divergence or mean squared error to measure the difference in the probability distributions output by the student model and the teacher model.

[0017] Optionally, in the teacher model update method based on accuracy, the step of using the validation set to test the current student model and obtaining the accuracy of the current student model is specifically as follows:

[0018] Use the current student model to predict the validation set, compare the obtained pseudo-labels with the true labels, and if they are equal, the prediction is correct. Among them, the ratio of the number of correctly predicted samples to the total number of samples is the accuracy.

[0019] Optionally, in the teacher model update method based on accuracy, after using the validation set to test the current student model and obtaining the accuracy of the current student model when the student model has learned all the training sets, the following steps are further included:

[0020] If the accuracy of the current student model is less than the accuracy corresponding to the student model with the best historical performance during the training process, end the current iteration process.

[0021] Optionally, in the teacher model update method based on accuracy, the calculation formula for the minimum accuracy update threshold is:

[0022] Acc min_gap = γ · (1 - Acc tea ) ;

[0023] Where Acc min_gap represents the minimum accuracy update threshold, γ represents the threshold coefficient, and Acc tea represents the accuracy of the current teacher model.

[0024] Optionally, in the teacher model update method based on accuracy, the student model with the best historical performance during the training process is the student model with the highest accuracy on the validation set.

[0025] Optionally, in the teacher model update method based on accuracy, the training set and the test set belong to the same data set.

[0026] In addition, to achieve the above object, the present invention also provides a teacher model update system based on accuracy, where the teacher model update system based on accuracy includes:

[0027] A data set acquisition module, configured to acquire a training set, a validation set, and a test set, and the validation set is randomly selected from the training set;

[0028] A parameter initialization module, configured to initialize training-related parameters, and the training-related parameters include initializing a balance factor and a threshold coefficient;

[0029] A student model calculation module, configured to obtain prediction information of a student model and calculate the cross-entropy loss of the student model;

[0030] A teacher model calculation module, configured to, if there is a teacher model, obtain prediction information of the teacher model, calculate the distillation loss of the teacher model, and incorporate the distillation loss into the global loss;

[0031] A backpropagation update module, configured to update the student model by backpropagation in a direction of reducing the global loss;

[0032] A student model verification module, configured to, after the student model has learned all of the training set, use the validation set to verify the current student model and obtain the accuracy of the current student model;

[0033] A student model saving module, configured to, if the accuracy of the current student model is greater than or equal to the accuracy corresponding to the student model with the best historical performance during the training process, save the current student model, and at the same time update the accuracy corresponding to the student model with the best historical performance during the training process and the degree of improvement in the performance of the current student model during the training process;

[0034] A teacher model update module, configured to, if the degree of improvement in the performance of the current student model during the updated training process is greater than or equal to a minimum accuracy update threshold, update the teacher model, and at the same time update the accuracy of the current teacher model and the minimum accuracy update threshold.

[0035] In addition, to achieve the above object, the present invention further provides a terminal, where the terminal includes: a memory, a processor, and an accuracy-based teacher model update program stored on the memory and executable on the processor, and when the accuracy-based teacher model update program is executed by the processor, the steps of the accuracy-based teacher model update method as described above are implemented.

[0036] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores an accuracy-based teacher model update program, and when the accuracy-based teacher model update program is executed by a processor, the steps of the accuracy-based teacher model update method as described above are implemented.

[0037] In the present invention, a training set, a validation set, and a test set are obtained, and the validation set is randomly selected from the training set; training-related parameters are initialized, and the training-related parameters include an initialized balance factor and a threshold coefficient; prediction information of a student model is obtained, and the cross-entropy loss of the student model is calculated; if there is a teacher model, prediction information of the teacher model is obtained, and the distillation loss of the teacher model is calculated, and the distillation loss is incorporated into the global loss; the student model is updated by backpropagation in the direction of reducing the global loss; when the student model has learned all of the training set, the validation set is used to test the current student model, and the accuracy of the current student model is obtained; if the accuracy of the current student model is greater than or equal to the accuracy corresponding to the student model with the best historical performance during the training process, the current student model is saved, and at the same time, the accuracy corresponding to the student model with the best historical performance during the training process and the degree of improvement in the performance of the current student model during the training process are updated; if the updated degree of improvement in the performance of the current student model during the training process is greater than or equal to the minimum accuracy update threshold, the teacher model is updated, and at the same time, the accuracy of the current teacher model and the minimum accuracy update threshold are updated. The present invention incorporates the performance of the teacher model into the decision-making process of the teacher model, and the calculated minimum performance difference is used as the condition for replacing the teacher model, taking into account both the stability and performance of the teacher model, and further improving the learning effect of the student model. Description of the Drawings

[0038] Figure 1 is a schematic diagram of the teacher model selection method in traditional self-distillation in the prior art;

[0039] Figure 2 is a schematic diagram of the teacher model selection method of the present invention;

[0040] Figure 3 is a flowchart of a preferred embodiment of the teacher model update method based on accuracy of the present invention;

[0041] Figure 4 is a schematic diagram of the influence of different values on the model in a preferred embodiment of the teacher model update method based on accuracy of the present invention;

[0042] Figure 5 is a schematic diagram of the principle of a preferred embodiment of the teacher model update system based on accuracy of the present invention;

[0043] Figure 6 is a schematic diagram of the operating environment of a preferred embodiment of the terminal of the present invention. Detailed Embodiments

[0044] To make the objectives, technical solutions and advantages of the present invention more clear and definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.

[0045] The method for updating a teacher model based on accuracy rate according to a preferred embodiment of the present invention is as Figure 2 and Figure 3 shown. The method for updating a teacher model based on accuracy rate includes the following steps:

[0046] Step S10: Obtain a training set, a validation set and a test set, and the validation set is randomly selected from the training set.

[0047] Specifically, prepare some pre-operations for training. For example, set the warmup proportion to 0.1, select the BertAdam optimizer, load the BERT configuration file, use publicly recognized open-source data for training, randomly select 10% of the data from the training set as the validation set, and the actual test set remains unchanged. The training set and the test set belong to the same data set (both are open-source data).

[0048] Step S20: Initialize training-related parameters, where the training-related parameters include an initialized balance factor and a threshold coefficient.

[0049] Specifically, the training-related parameters include an initialized balance factor α and a threshold coefficient γ, and initialize the balance factor α and the threshold coefficient γ.

[0050] Step S30: Obtain the prediction information of the student model and calculate the cross-entropy loss of the student model.

[0051] Step S40: If there is a teacher model, obtain the prediction information of the teacher model, calculate the distillation loss of the teacher model, and incorporate the distillation loss into the global loss.

[0052] Specifically, if there is a teacher model, obtain the prediction information of the teacher model, calculate the distillation loss, and incorporate the distillation loss into the global loss loss. The distillation loss uses the KL (Kullback Lleibler) divergence or the mean square error MSE (Mean Square Error) to measure the difference in the probability distributions output by the student model and the teacher model. Since in the experimental process, using MSE to align the logits of the teacher and student models (i.e., the output information of the last layer of the model) has a better effect, the MSE method is adopted here.

[0053] Step S50: Update the student model in the direction of reducing the global loss loss by backpropagation, and repeat Steps S20 - S50 until the student model has learned all the training data (i.e., the training set).

[0054] Step S60: When the student model has learned all the training sets, use the validation set to test the current student model and obtain the accuracy Acc of the current student model now , and increment the variable num no_improvement by 1.

[0055] Step S70: If the accuracy Acc of the current student model now is greater than or equal to the accuracy Acc corresponding to the student model with the best historical performance during the training process best (i.e., Acc now ≥Acc best ), then save the current student model and simultaneously update the accuracy Acc corresponding to the student model with the best historical performance during the training process best and the degree of improvement in the performance of the current student model during the training process Acc gap .

[0056] Step S80: After meeting the conditions in Step S70, if the degree of improvement in the performance of the current student model during the updated training process Acc gap is greater than or equal to the minimum accuracy update threshold Acc min_gap (i.e., Acc gap ≥Acc min_gap ), then update the teacher model and simultaneously update the accuracy Acc of the current teacher model tea and the minimum accuracy update threshold Acc min_gap . Since there is a certain lag in knowledge propagation during the process of distilling the knowledge of the teacher model to the student model, the num no_improvenent variable is set to zero to extend the number of iterations.

[0057] If the conditions in Step S70 are not met, end the current iteration process.

[0058] If num no_improvenent ≥ the iteration parameter patient, end the model training in advance and test the model with the test set; otherwise, train until the preset maximum number of iterations is reached and then end.

[0059] For example Figure 2As shown, first assume that there is no fallback situation in the teacher model, that is, the teacher model at time t is the student model at time t - i. Then, the teacher model at time t + 1 can only be the student model at time t - i or t, and cannot be the student model at time t - i - 1 or an earlier time. Here, t >= i + 1, and i >= 1.

[0060] Based on the above conditions, there is an obvious sequence between these two actions. At the same moment, the action of saving the student model must precede the action of updating the teacher model. In addition, the quality of the teacher model can be reflected to a certain extent on the validation set and quantified by the accuracy. Subtract the accuracies of two teacher models. If the difference is greater than the minimum threshold, update; otherwise, do not update. The specific formula is as follows:

[0061] Acc min_gap = γ·(1 - Acc tea ); (1)

[0062] where Acc min_gap represents the minimum accuracy update threshold, γ represents the threshold coefficient, and Acc tea represents the accuracy of the current teacher model.

[0063] Different from the update idea of the student model, the above formula (1) only specifies the threshold coefficient γ, and the more specific minimum threshold is dynamically changing. The idea expressed by the above formula is that in the first few iterations of training, due to the large learning rate, the model parameters are updated quickly, and there is a large room for improvement in the model. Therefore, a larger minimum threshold is required. On the contrary, in the last few iterations of training, the learning rate decays to a smaller value, the model parameters are updated more slowly, and the model is more stable with a smaller room for improvement. Therefore, a smaller minimum threshold is required.

[0064] The algorithm flow for dynamically selecting the teacher model is as follows:

[0065] Input: training set text collection, training set label collection, balance factor α, threshold coefficient γ, and iteration parameter patient;

[0066] Initialization: Let num no_improvement = 0, Acc best = 0, Acc tea = 0, Acc gap = 0, Acc min_gap = 0;

[0067] 1: for epoch = 1,..., epoch max do;

[0068] 2: Randomly initialize the training set;

[0069] 3: For b = 1, ..., B do;

[0070] 4: slogits = model s ·forword(x) / / x is the training batch text set;

[0071] 5: loss ce = CELoss(slogits, y) / / y is the corresponding label of the training batch;

[0072] 6: If model t is Not None then;

[0073] 7: tlogits = model t ·forword(x);

[0074] 8: loss mse = MSELoss(slogits, tlogits);

[0075] 9: loss = loss ce + α * loss mse ;

[0076] 10: Else;

[0077] 11: loss = loss ce ;

[0078] 12: End if;

[0079] 13: loss.backward();

[0080] 14: End for;

[0081] 15: Use the validation set to test the student model and obtain Acc now ;

[0082] 16: num no_improvement += 1;

[0083] 17: If Acc now ≥ Acc best then;

[0084] 18: Acc best = Acc now ;

[0085] 19: Acc gap = Acc best - Acc tea ;

[0086] 20: Save the student model;

[0087] 20: if Acc gap ≥ Acc min_gap then;

[0088] 21: num no_improvement = 0;

[0089] 22: Acc tea = Acc best ;

[0090] 23: Acc min_gap = γ·(1 - Acc tea )

[0091] 24: Update the teacher model;

[0092] 25: end if;

[0093] 26: end if;

[0094] 27: end for (if num no_improvement ≥ patient, end training in advance).

[0095] As can be seen from the above algorithm and the described training steps, num no_improvement and patient are set to terminate training in advance.

[0096] The present invention incorporates the performance of the teacher model, such as the accuracy of the classification model, into the decision-making process of the teacher model, and uses the calculated minimum performance difference as the condition for replacing the teacher model, taking into account both the stability and performance of the teacher model, and can further improve the learning effect of the student model.

[0097] The present invention mines the model performance (accuracy) as a judgment condition, constructs a more flexible method for updating the teacher model, further improves the performance of the student model in the self-distillation scenario, and in addition, this method can also be extended to a general method for early termination of model training.

[0098] Furthermore, the experimental data is as follows:

[0099]

[0100] Table 1: Summary of experimental parameters

[0101]

[0102] Table 2: Details of the text classification dataset

[0103]

[0104] Table 3: Fixing the α value at 1.0 and exploring the impact of different γ values on the model

[0105] Table 1 summarizes the specific parameters during the experiment. Since the impact of different γ values on the model cannot be directly observed in Table 3 when the α value is fixed, the iteration number is taken as the abscissa and the accuracy rate is taken as the ordinate to obtain Figure 4 , where MR, R8, R52, and 20NG are four different English text classification datasets. The specific details can be seen in Table 2, where AVG represents taking the average value of the length. From a macroscopic perspective, as the γ value gradually increases, the accuracy rate of the classification model first rises and then falls, while the iteration number shows a gradually decreasing trend. In particular, the performance of the classification model on each dataset when the γ value is 0.05 is marked with a five-pointed star for easy observation of the impact of different γ values on the experiment later. From a microscopic perspective, there are significant differences in the performance of the classification model on different datasets. For example, the changes in the 20NG dataset are scattered, but the changes are more concentrated in the R8 dataset. It is speculated that the reason is that the length of a single sample in the 20NG dataset is longer, and when the parameter max_seq_length is 256, it cannot well contain all the information of the sample, while most of the sample content in the R8 dataset is within 256 words.

[0106]

[0107] Table 4: Fixing the γ value at 0.05 and exploring the impact of different α values on the model

[0108] Table 4 records three situations during the self-distillation process when the proportion of the distillation loss is lower than, equal to, or higher than the cross-entropy loss. The bold part in Table 4 records the best accuracy rate values among all records on each dataset. It can be found that when the α value is too small, the performance of the model is poor, and the effect of distillation is not obvious at this time. In addition, the performance around the α value of 1.0 is generally better.

[0109]

[0110] Table 5: Comparison with related advanced research

[0111] Table 5 collates and summarizes the advanced research results on the same datasets in recent years. Although the accuracy rate of the present invention (i.e., this work) on most datasets still has a certain gap from the optimal performance, after further considering the variable of the model parameters, it can be found that the results obtained in this work are the most cost-effective among all the studies.

[0112] The present invention can be applied to most model training scenarios. As long as there is a way to quantify performance during the model training stage, such as the accuracy metric in a text classification task, the formula (1) can be transformed for application. For example, text classification is a specific task scenario. The network framework constructed to solve this task is the model, and the teacher model is the selected student model used to guide the subsequent training process of the student model. However, this method is more recommended for application in the self-distillation scenario, especially for tasks where the model has low performance at the beginning of training and high performance at the end. In this case, the effect of filtering the teacher model is better. If it is applied in a general training scenario, a suitable γ value needs to be selected because an overly large γ value (such as 1.0) will cause the min_gap value to be too large, resulting in the model ending training relatively quickly; an overly small γ value (such as 0.0) will make the min_gap value always 0, obtaining the same effect as when this method is not used.

[0113] Furthermore, as Figure 5 shown, based on the above-mentioned teacher model update method based on accuracy, the present invention also correspondingly provides a teacher model update system based on accuracy. Among them, the teacher model update system based on accuracy includes:

[0114] A dataset acquisition module 51, configured to acquire a training set, a validation set, and a test set, and the validation set is randomly selected from the training set;

[0115] A parameter initialization module 52, configured to initialize training-related parameters, and the training-related parameters include an initialization balance factor and a threshold coefficient;

[0116] A student model calculation module 53, configured to obtain the prediction information of the student model and calculate the cross-entropy loss of the student model;

[0117] A teacher model calculation module 54, configured to, if there is a teacher model, obtain the prediction information of the teacher model, calculate the distillation loss of the teacher model, and incorporate the distillation loss into the global loss;

[0118] A backpropagation update module 55, configured to update the student model in the direction of reducing the global loss by backpropagation;

[0119] A student model verification module 56, configured to, after the student model has learned all the training sets, use the validation set to verify the current student model and obtain the accuracy of the current student model;

[0120] The student model saving module 57 is configured to save the current student model if the accuracy rate of the current student model is greater than or equal to the accuracy rate of the student model with the best historical performance during the training process, and simultaneously update the accuracy rate of the student model with the best historical performance during the training process and the degree of improvement in the performance of the current student model during the training process;

[0121] The teacher model updating module 58 is configured to update the teacher model if the degree of improvement in the performance of the current student model during the updated training process is greater than or equal to the minimum accuracy rate update threshold, and simultaneously update the accuracy rate of the current teacher model and the minimum accuracy rate update threshold.

[0122] Further, as Figure 6 shown, based on the above accuracy rate-based teacher model updating method and system, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20, and a display 30. Figure 6 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0123] The memory 20 may be an internal storage unit of the terminal in some embodiments, such as a hard disk or a memory of the terminal. The memory 20 may also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk equipped on the terminal, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as program codes for installing the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a teacher model updating program 40 based on accuracy rate is stored on the memory 20, and the teacher model updating program 40 based on accuracy rate can be executed by the processor 10, so as to implement the accuracy rate-based teacher model updating method in this application.

[0124] The processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chips in some embodiments, and is used to run program codes stored in the memory 20 or process data, such as executing the accuracy rate-based teacher model updating method, etc.

[0125] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) toucher, etc. The display 30 is used to display information on the terminal and to display a visual user interface. The components 10-30 of the terminal communicate with each other via a system bus.

[0126] In one embodiment, when the processor 10 executes the accuracy-based teacher model update program 40 in the memory 20, the steps of the above-mentioned accuracy-based teacher model update method are implemented.

[0127] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an accuracy-based teacher model update program, and when the accuracy-based teacher model update program is executed by a processor, the steps of the above-mentioned accuracy-based teacher model update method are implemented.

[0128] In summary, the present invention provides an accuracy-based teacher model update method and related devices. The method includes: obtaining a training set, a validation set, and a test set, where the validation set is randomly selected from the training set; initializing training-related parameters, where the training-related parameters include initializing a balance factor and a threshold coefficient; obtaining prediction information of a student model and calculating the cross-entropy loss of the student model; if there is a teacher model, obtaining prediction information of the teacher model and calculating the distillation loss of the teacher model, and incorporating the distillation loss into the global loss; backpropagating and updating the student model in the direction of reducing the global loss; when the student model has learned all of the training set, using the validation set to test the current student model and obtaining the accuracy of the current student model; if the accuracy of the current student model is greater than or equal to the accuracy corresponding to the student model with the best historical performance during the training process, saving the current student model, and simultaneously updating the accuracy corresponding to the student model with the best historical performance during the training process and the degree of improvement in the performance of the current student model during the training process; if the updated degree of improvement in the performance of the current student model during the training process is greater than or equal to the minimum accuracy update threshold, updating the teacher model, and simultaneously updating the accuracy of the current teacher model and the minimum accuracy update threshold. The present invention incorporates the performance of the teacher model into the decision-making process of the teacher model, and uses the calculated minimum performance difference as the condition for replacing the teacher model, taking into account both the stability and performance of the teacher model, and further improving the learning effect of the student model.

[0129] It should be noted that in this text, the term "including", "comprising" or any other variants thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or terminal including that element.

[0130] Of course, those of ordinary skill in the art can understand that all or part of the processes in the above-described method embodiments can be implemented by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium readable by a computer. When the program is executed, it can include the processes of the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0131] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for updating a teacher model based on accuracy, characterized in that The accuracy-based teacher model updating method is applied to MR, R8, R52, and 20NG text classification datasets. The accuracy-based teacher model updating method includes: Obtain a training set, a validation set, and a test set, where the validation set is randomly sampled from the training set; Initialize training-related parameters, where the training-related parameters include an initialized balance factor and a threshold coefficient; Obtain the prediction information of the student model and calculate the cross-entropy loss of the student model; If there is a teacher model, obtain the prediction information of the teacher model and calculate the distillation loss of the teacher model, and incorporate the distillation loss into the global loss; Update the student model by backpropagation in the direction of reducing the global loss; After the student model has learned all of the training set, use the validation set to test the current student model and obtain the accuracy of the current student model; If the accuracy of the current student model is greater than or equal to the accuracy corresponding to the student model with the best historical performance during the training process, save the current student model, and at the same time update the accuracy corresponding to the student model with the best historical performance during the training process and the degree of improvement in the performance of the current student model during the training process; If the degree of improvement in the performance of the current student model during the updated training process is greater than or equal to the minimum accuracy update threshold, update the teacher model, and at the same time update the accuracy of the current teacher model and the minimum accuracy update threshold; The formula for the minimum accuracy update threshold is: ; Among them, represents the minimum accuracy update threshold, represents the threshold coefficient, represents the accuracy of the current teacher model.

2. The method for updating a teacher model based on accuracy according to claim 1, wherein The distillation loss uses KL divergence or mean squared error to measure the difference in the probability distributions output by the student model and the teacher model.

3. The method for updating a teacher model based on accuracy according to claim 1, wherein The step of using the validation set to test the current student model and obtaining the accuracy of the current student model is specifically: Use the current student model to predict the validation set, compare the obtained pseudo-labels with the true labels, and if they are equal, the prediction is correct. Among them, the number of correctly predicted samples divided by the total number of samples is the accuracy.

4. The method for updating a teacher model based on accuracy according to claim 1, wherein After the step of, after the student model has learned all of the training set, using the validation set to test the current student model and obtaining the accuracy of the current student model, it further includes: If the accuracy of the current student model is less than the accuracy corresponding to the student model with the best historical performance during the training process, end the current iteration process.

5. The method for updating a teacher model based on accuracy rate according to claim 1, wherein The student model with the best historical performance during the training process is the student model with the highest accuracy on the validation set.

6. The method for updating a teacher model based on accuracy according to claim 1, wherein The training set and the test set belong to the same dataset.

7. A teacher model update system based on accuracy, characterized in that The accuracy-based teacher model updating system is applied to the accuracy-based teacher model updating method according to any one of claims 1-6. The accuracy-based teacher model updating system includes: A dataset acquisition module for obtaining a training set, a validation set, and a test set, where the validation set is randomly sampled from the training set; A parameter initialization module for initializing training-related parameters, where the training-related parameters include an initialized balance factor and a threshold coefficient; A student model calculation module for obtaining the prediction information of the student model and calculating the cross-entropy loss of the student model; A teacher model calculation module, configured to, if there is a teacher model, obtain prediction information of the teacher model, calculate a distillation loss of the teacher model, and incorporate the distillation loss into a global loss; A backpropagation update module, configured to update the student model in the direction of reducing the global loss through backpropagation; A student model verification module, configured to, when the student model has learned all of the training sets, use the validation set to verify the current student model and obtain an accuracy rate of the current student model; A student model saving module, configured to, if the accuracy rate of the current student model is greater than or equal to the accuracy rate corresponding to the student model with the best historical performance during the training process, save the current student model, and simultaneously update the accuracy rate corresponding to the student model with the best historical performance during the training process and the degree of improvement in the performance of the current student model during the training process; A teacher model update module, configured to, if the degree of improvement in the performance of the current student model during the updated training process is greater than or equal to a minimum accuracy rate update threshold, update the teacher model, and simultaneously update the accuracy rate of the current teacher model and the minimum accuracy rate update threshold.

8. A terminal, characterized in that, The terminal includes: a memory, a processor, and an accuracy rate-based teacher model update program stored on the memory and executable on the processor. When the accuracy rate-based teacher model update program is executed by the processor, the steps of the accuracy rate-based teacher model update method according to any one of claims 1-6 are implemented.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an accuracy rate-based teacher model update program. When the accuracy rate-based teacher model update program is executed by a processor, the steps of the accuracy rate-based teacher model update method according to any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Student model training method and device and electronic equipment

    CN113435208A

  • Knowledge distillation method and system based on multi-student discussion

    CN114049513A