Residual knowledge distillation method based on heterogeneous soft labels

Through the residual knowledge distillation method based on heterogeneous soft labels, the problem that traditional knowledge distillation method cannot fully improve the performance of students' models is solved, and the efficient learning and performance improvement of students' models is achieved, while reducing the impact of teacher model errors.

CN119990256AInactive Publication Date: 2025-05-13鞠礼惠
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510133987.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional knowledge distillation method cannot fully consider the personalized learning needs of the student model, which leads to limited improvement in the performance of the student model. At the same time, the errors of the teacher model will affect the learning effect of the student model, and the student model lacks knowledge review, resulting in knowledge forgetting.

Method used

A residual knowledge distillation method based on heterogeneous soft labels is designed. By finely dividing the soft labels and performing residual distillation according to their characteristics, the student model can flexibly select the required knowledge and effectively learn the correct knowledge of the teacher model.

Benefits of technology

It significantly improves the performance of the student model, reduces the negative impact of teacher model error prediction on student model, and effectively avoids knowledge forgetting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990256A_ABST
    Figure CN119990256A_ABST
Patent Text Reader

Abstract

The invention designs a residual knowledge distillation method based on heterogeneous soft labels, which not only utilizes a teacher model to predict correct samples, but also utilizes the teacher model to predict wrong samples, and utilizes important knowledge contained in the wrong samples. According to the method, a residual distillation method is applied to guide a student model to focus on learning correct knowledge in a heterogeneous soft label. And the soft label with correct prediction of the target class is directly used for training a student model. And for the soft label with the prediction error of the target class, the target class with the prediction error needs to be eliminated, and only the classification probability of the non-target class is reserved for the student model learning. The performance of the student model is ensured, and meanwhile, the negative influence on the student model caused by misprediction of the teacher model is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0003] The present invention relates to the field of machine learning, and more specifically to a model compression method that transfers knowledge in a teacher model to a student model through knowledge distillation to enhance student performance. Background Art

[0005] Traditional knowledge distillation methods usually rely on preset fixed distillation points for learning, which cannot fully consider the personalized learning needs of the student model, thus limiting the performance improvement of the student model. At the same time, there are certain errors in the prediction results of the teacher model. Directly using the soft labels provided by the teacher model for learning will affect the learning effect of the student model. In addition, the student model lacks review of the learned knowledge during the learning process, resulting in knowledge forgetting, which further restricts its performance improvement.

[0006] The present invention designs a residual knowledge distillation method based on heterogeneous soft labels, which allows the student model to flexibly select the required knowledge according to the current learning state, and effectively learn the correct knowledge of the teacher model, thereby significantly improving the performance of the student model. This method effectively reduces the negative impact of the teacher model's incorrect prediction on the student model by finely dividing the soft labels and performing residual distillation based on their characteristics, rather than passively accepting the soft label knowledge provided by the teacher model. Summary of the invention

[0008] The present invention designs a residual knowledge distillation method based on heterogeneous soft labels. The method not only uses the samples predicted correctly by the teacher model, but also uses the samples predicted incorrectly by the teacher model, because these incorrect samples also contain some important knowledge. The method first subdivides the soft labels into two types of heterogeneous soft labels according to the accuracy of the teacher model's prediction of the target class: soft labels predicted correctly by the target class and soft labels predicted incorrectly by the target class. Subsequently, the residual distillation method is used to guide the student model to focus on learning the correct knowledge in the heterogeneous soft labels. For the soft labels predicted correctly by the target class, they are directly used for the training of the student model, because such labels represent the data features extracted by the teacher model very accurately, which is conducive to the learning of the student model. For the soft labels predicted incorrectly by the target class, it is necessary to eliminate the target class predicted incorrectly, and only retain the classification probability of the non-target class for the student model to learn. This is because such labels represent that there is a deviation in the data features extracted by the teacher model. If they are directly applied to the training of the student model, it will mislead the student model. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 The test accuracy of the student model under the condition of heterogeneous teacher and student models (%); DETAILED DESCRIPTION

[0012] The residual knowledge distillation method based on heterogeneous soft labels is characterized by:

[0013] (1) The training data set containing C categories and N data is represented as D = ; The teacher model predicts M samples correctly for the target class, expressed as ; There are NM samples with incorrect predictions for the target class, expressed as ;

[0014] (2) For For a training sample in the data set, the prediction vector is expressed as = [ ],in is the probability of the ith class:

[0015] =

[0016] in, Indicates that the sample is predicted to be the logit component of the i-th category;

[0017] (3) Use ] and ] represent the teacher model and the student model in The output result for a sample in the data set;

[0018] = , =

[0019] in, represents the logit of the i-th category of the sample in the teacher model, Indicates the logit component of the sample predicted as the i-th category in the student model;

[0020] (4) The objective function of the student model learning the target class to predict the correct soft label is:

[0021]

[0022] in, Represents the teacher model pair The prediction result of the i-th sample in the data set, Represents the student model pair The prediction result of the i-th sample in the data set, KL represents the KL divergence loss function;

[0023] (5) For For a training sample in the data set, if R represents the target category of the sample, then the probability value of the target category is eliminated. The predicted vector is expressed as = [ ];

[0024] =

[0025] (6) Use ] and ] represent the teacher model and the student model respectively. The output result for a sample in the data set; then:

[0026] = , =

[0027] (7) The objective function of the student model learning the soft label of the target class prediction error is:

[0028]

[0029] in, Represents the teacher model pair The prediction result of the i-th sample in the data set, Represents the student model pair The prediction result of the i-th sample in the data set;

[0030] (8) The overall objective function for residual distillation of soft labels is:

[0031] .

[0032] Validity Verification

[0033] The validity verification experiment of the present invention is as follows:

[0034] (1) Experimental setup

[0035] The present invention verifies the effectiveness of this scheme under the condition of teacher-student heterogeneity on the CIFAR-10 dataset and the SVHN dataset. Teacher-student heterogeneity means that the teacher model and the student model use different types of network architectures. In the teacher-student heterogeneity, the teacher model we use is ResNet18, and the student models are AlexNet, Tiny-AlexNet and Tiny-VggNet.

[0036] (2) Experimental results and analysis

[0037] Combination Figure 1 It can be seen that under the heterogeneous teacher-student experimental setting, the student model trained by the scheme in this paper achieved an accuracy of 79.75% on the CIFAR-10 dataset and an accuracy of 89.99% on the SVHN dataset. The experimental results show that the present invention effectively reduces the negative impact of the teacher model's incorrect prediction on the student model while ensuring the performance of the student model by finely dividing the soft labels and performing residual distillation according to their characteristics, rather than passively accepting the soft label knowledge provided by the teacher model.

Claims

1. The residual knowledge distillation method based on heterogeneous soft labels is characterized by: (1) The training data set containing C categories and N data is represented as D = ; The teacher model predicts M samples correctly for the target class, expressed as ; There are NM samples with incorrect predictions for the target class, expressed as ; (2) For For a single training sample in the dataset, the prediction vector is represented as = [ ],in is the probability of the ith class: = in, Indicates that the sample is predicted to be the logit component of the i-th category; (3) Use ] and ] represent the teacher model and the student model in The output result for a sample in the data set is = , = in, Indicates the logit component of the sample predicted as the i-th category in the teacher model, Indicates the logit component of the sample predicted as the i-th category in the student model; (4) The objective function of the student model learning the target class to predict the correct soft label is: in, Represents the teacher model pair The prediction result of the i-th sample in the data set, Represents the student model pair The prediction result of the i-th sample in the data set, KL represents the KL divergence loss function; (5) For For a training sample in the data set, if R represents the target category of the sample, then the probability value of the target category is eliminated. The predicted vector is expressed as = [ ]; = (6) Use ] and ] represent the teacher model and the student model respectively. The output result for a sample in the data set; then: = , = (7) The objective function of the student model learning the soft label of the target class prediction error is: in, Represents the teacher model pair The prediction result of the i-th sample in the data set, Represents the student model pair The prediction result of the i-th sample in the data set; (8) The overall objective function for residual distillation of soft labels is: 。