Residual knowledge distillation method based on heterogeneous soft labels
Through the residual knowledge distillation method based on heterogeneous soft labels, the problem that traditional knowledge distillation method cannot fully improve the performance of students' models is solved, and the efficient learning and performance improvement of students' models is achieved, while reducing the impact of teacher model errors.
Patent Information
- Application Number
- CN202510133987.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional knowledge distillation method cannot fully consider the personalized learning needs of the student model, which leads to limited improvement in the performance of the student model. At the same time, the errors of the teacher model will affect the learning effect of the student model, and the student model lacks knowledge review, resulting in knowledge forgetting.
A residual knowledge distillation method based on heterogeneous soft labels is designed. By finely dividing the soft labels and performing residual distillation according to their characteristics, the student model can flexibly select the required knowledge and effectively learn the correct knowledge of the teacher model.
It significantly improves the performance of the student model, reduces the negative impact of teacher model error prediction on student model, and effectively avoids knowledge forgetting.
Smart Images

Figure CN119990256A_ABST
Abstract
Description
Technical Field
[0003] The present invention relates to the field of machine learning, and more specifically to a model compression method that transfers knowledge in a teacher model to a student model through knowledge distillation to enhance student performance. Background Art
[0005] Traditional knowledge distillation methods usually rely on preset fixed distillation points for learning, which cannot fully consider the personalized learning needs of the student model, thus limiting the performance improvement of the student model. At the same time, there are certain errors in the prediction results of the teacher model. Directly using the soft labels provided by the teacher model for learning will affect the learning effect of the student model. In addition, the student model lacks review of the learned knowledge during the learning process, resulting in knowledge forgetting, which further restricts its performance improvement.
[0006] The present invention designs a residual knowledge distillation method based on heterogeneous soft labels, which allows the student model to flexibly select the required knowledge according to the current learning state, and effectively learn the correct knowledge of the teacher model, thereby significantly improving the performance of the student model. This method effectively reduces the negative impact of the teacher model's incorrect prediction on the student model by finely dividing the soft labels and performing residual distillation based on their characteristics, rather than passively accepting the soft label knowledge provided by the teacher model. Summary of the invention
[0008] The present invention designs a residual knowledge distillation method based on heterogeneous soft labels. The method not only uses the samples predicted correctly by the teacher model, but also uses the samples predicted incorrectly by the teacher model, because these incorrect samples also contain some important knowledge. The method first subdivides the soft labels into two types of heterogeneous soft labels according to the accuracy of the teacher model's prediction of the target class: soft labels predicted correctly by the target class and soft labels predicted incorrectly by the target class. Subsequently, the residual distillation method is used to guide the student model to focus on learning the correct knowledge in the heterogeneous soft labels. For the soft labels predicted correctly by the target class, they are directly used for the training of the student model, because such labels represent the data features extracted by the teacher model very accurately, which is conducive to the learning of the student model. For the soft labels predicted incorrectly by the target class, it is necessary to eliminate the target class predicted incorrectly, and only retain the classification probability of the non-target class for the student model to learn. This is because such labels represent that there is a deviation in the data features extracted by the teacher model. If they are directly applied to the training of the student model, it will mislead the student model. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 The test accuracy of the student model under the condition of heterogeneous teacher and student models (%); DETAILED DESCRIPTION
[0012] The residual knowledge distillation method based on heterogeneous soft labels is characterized by:
[0013] (1) The training data set containing C categories and N data is represented as D = ; The teacher model predicts M samples correctly for the target class, expressed as ; There are NM samples with incorrect predictions for the target class, expressed as ;
[0014] (2) For For a training sample in the data set, the prediction vector is expressed as = [ ],in is the probability of the ith class:
[0015] =
[0016] in, Indicates that the sample is predicted to be the logit component of the i-th category;
[0017] (3) Use ] and ] represent the teacher model and the student model in The output result for a sample in the data set;
[0018] = , =
[0019] in, represents the logit of the i-th category of the sample in the teacher model, Indicates the logit component of the sample predicted as the i-th category in the student model;
[0020] (4) The objective function of the student model learning the target class to predict the correct soft label is:
[0021]
[0022] in, Represents the teacher model pair The prediction result of the i-th sample in the data set, Represents the student model pair The prediction result of the i-th sample in the data set, KL represents the KL divergence loss function;
[0023] (5) For For a training sample in the data set, if R represents the target category of the sample, then the probability value of the target category is eliminated. The predicted vector is expressed as = [ ];
[0024] =
[0025] (6) Use ] and ] represent the teacher model and the student model respectively. The output result for a sample in the data set; then:
[0026] = , =
[0027] (7) The objective function of the student model learning the soft label of the target class prediction error is:
[0028]
[0029] in, Represents the teacher model pair The prediction result of the i-th sample in the data set, Represents the student model pair The prediction result of the i-th sample in the data set;
[0030] (8) The overall objective function for residual distillation of soft labels is:
[0031] .
[0032] Validity Verification
[0033] The validity verification experiment of the present invention is as follows:
[0034] (1) Experimental setup
[0035] The present invention verifies the effectiveness of this scheme under the condition of teacher-student heterogeneity on the CIFAR-10 dataset and the SVHN dataset. Teacher-student heterogeneity means that the teacher model and the student model use different types of network architectures. In the teacher-student heterogeneity, the teacher model we use is ResNet18, and the student models are AlexNet, Tiny-AlexNet and Tiny-VggNet.
[0036] (2) Experimental results and analysis
[0037] Combination Figure 1 It can be seen that under the heterogeneous teacher-student experimental setting, the student model trained by the scheme in this paper achieved an accuracy of 79.75% on the CIFAR-10 dataset and an accuracy of 89.99% on the SVHN dataset. The experimental results show that the present invention effectively reduces the negative impact of the teacher model's incorrect prediction on the student model while ensuring the performance of the student model by finely dividing the soft labels and performing residual distillation according to their characteristics, rather than passively accepting the soft label knowledge provided by the teacher model.
Claims
1. The residual knowledge distillation method based on heterogeneous soft labels is characterized by: (1) The training data set containing C categories and N data is represented as D = ; The teacher model predicts M samples correctly for the target class, expressed as ; There are NM samples with incorrect predictions for the target class, expressed as ; (2) For For a single training sample in the dataset, the prediction vector is represented as = [ ],in is the probability of the ith class: = in, Indicates that the sample is predicted to be the logit component of the i-th category; (3) Use ] and ] represent the teacher model and the student model in The output result for a sample in the data set is = , = in, Indicates the logit component of the sample predicted as the i-th category in the teacher model, Indicates the logit component of the sample predicted as the i-th category in the student model; (4) The objective function of the student model learning the target class to predict the correct soft label is: in, Represents the teacher model pair The prediction result of the i-th sample in the data set, Represents the student model pair The prediction result of the i-th sample in the data set, KL represents the KL divergence loss function; (5) For For a training sample in the data set, if R represents the target category of the sample, then the probability value of the target category is eliminated. The predicted vector is expressed as = [ ]; = (6) Use ] and ] represent the teacher model and the student model respectively. The output result for a sample in the data set; then: = , = (7) The objective function of the student model learning the soft label of the target class prediction error is: in, Represents the teacher model pair The prediction result of the i-th sample in the data set, Represents the student model pair The prediction result of the i-th sample in the data set; (8) The overall objective function for residual distillation of soft labels is: 。