A Knowledge Distillation Method Without Data Based on Student Feedback
By introducing auxiliary classifiers and self-supervised enhancement tasks into the student model, and adjusting the difficulty of synthesized pictures in combination with feedback from students and teachers, the problem of unconsidered learning ability of students in distillation of data-free knowledge is solved, and efficient training and performance improvement of student model is achieved.
Patent Information
- Application Number
- CN202211028120.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-25
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-08-25
AI Technical Summary
The existing data-free knowledge distillation method fails to effectively consider the learning ability of the student model during the image synthesis process, resulting in the synthesis of pictures being too simple and weakening the final performance of the student model.
By adding auxiliary classifiers to the student model, the noise vector and generator are trained jointly with the loss function of student feedback and teacher feedback, and adaptively adjust the difficulty of the synthetic picture to match the current learning ability of the student model, and optimize the generation of the synthetic picture with self-supervised enhancement tasks and the data prior information of the teacher model.
Effectively train students' models to improve their final performance, ensure that the synthetic pictures meet the learning needs of students' models, and improve the learning efficiency and accuracy of the models.
Smart Images

Figure CN115409157B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge distillation, and particularly relates to a data-free knowledge distillation method based on student feedback. Background Art
[0002] In recent years, convolutional neural networks have achieved remarkable success in various practical applications. However, their high storage and computational costs make it difficult to deploy models on mobile devices. Therefore, Hinton et al. proposed the knowledge distillation technique to achieve model compression, and its main idea is to transfer the dark knowledge from a pre-trained heavyweight teacher model to a lightweight student model.
[0003] Typical knowledge distillation methods are all based on a strong premise that the original data used to train the teacher model can be directly used to train the student model. However, in some practical scenarios, due to reasons such as privacy, intellectual property rights, or large datasets, the data is not publicly shared. Thus, data-free knowledge distillation was proposed to solve this problem. The existing related work mainly uses the feedback of the teacher model to achieve image synthesis, and then uses the synthesized images to replace the original images in the knowledge distillation process.
[0004] However, the existing work does not explicitly consider the learning ability of the student during the image synthesis process, and the synthesized images may be too simple for the current ability of the student, resulting in the student model not learning new knowledge, thus weakening the final performance of the model. Summary of the Invention
[0005] The main purpose of the present invention is to overcome the above-mentioned defects in the prior art, and propose a data-free knowledge distillation method based on student feedback, which uses a self-supervised enhanced auxiliary task to estimate the current learning ability of the student, thereby adaptively adjusting the content of the synthesized images to generate samples that are difficult for the student model, enabling the student model to continuously acquire new knowledge to improve the final performance of the student model.
[0006] The present invention adopts the following technical solutions:
[0007] A data-free knowledge distillation method based on student feedback, comprising the following steps:
[0008] S1: Initialize the student model, and add an auxiliary classifier after the feature extractor of the student model;
[0009] S2: Use the auxiliary classifier to feedback the current learning ability of the student model, and at the same time jointly train the noise vector and the generator according to the loss functions of the student feedback and the teacher feedback to obtain the best synthesized images;
[0010] S3: Use the synthetic images obtained in S2 to train the student model through knowledge distillation, and at the same time independently train an auxiliary classifier to learn the auxiliary task;
[0011] S4: Repeat S2 and S3 until the student model is trained to convergence.
[0012] Specifically, the current learning ability of the student model is fed back by the auxiliary classifier in S2, and the specific process includes:
[0013] Randomly generate a noise vector z and input it into the generator network The synthetic images can be obtained Then for the synthetic images Rotate by a certain angle, and input the rotated images Into the feature extractor Φ of the student model So as to obtain the feature representation And input it into the auxiliary classifier Use the output result of the auxiliary classifier Calculate the loss function to quantify the current learning ability of the student model, that is, the loss function fed back by the student, specifically:
[0014]
[0015] Among them, k represents the class label of the self-supervised augmentation task, and the self-supervised augmentation task regards a self-supervised rotation task and the original image classification task as a joint task.
[0016] Specifically, the specific definition of the categories of the self-supervised augmentation task is as follows:
[0017] Given that the total number of categories of the original image classification task is N, and the total number of categories of the self-supervised rotation task is M; assuming that the synthetic image Is of class n in the image classification task, and its rotated version is of class m in the self-supervised rotation task, then its category in the self-supervised augmentation task is n*M + m.
[0018] Specifically, the loss function fed back by the teacher in S2 is specifically:
[0019]
[0020] Among them, Is the output of the teacher model And the label of the predefined image classification task The cross entropy between them, and its formula is expressed as:
[0021]
[0022] The l2-norm distance between the synthetic image and the real image feature statistics, which is expressed by the formula:
[0023]
[0024] Wherein, and are the mean and variance of the feature map of the synthetic image at the l-th layer of the teacher model; μ l and are the mean and variance of the feature map stored in the l-th layer of the teacher model, that is, representing the feature statistics information of the real image.
[0025] Specifically, in S2, the noise vector and the generator are jointly trained according to the loss functions of both the student feedback and the teacher feedback, and the overall loss function is:
[0026]
[0027] Wherein, α is a hyperparameter used to balance the weights of the two loss terms.
[0028] Specifically, in S3, the overall loss function for training the student model through knowledge distillation is:
[0029]
[0030] Wherein, β is a hyperparameter used to balance the weights of the three loss terms; is the conventional loss term in the original image classification task, which is used to calculate the cross-entropy between the output of the student model and the predefined label;
[0031] is the KL divergence between the teacher and student outputs, and its formula is expressed as:
[0032]
[0033] Wherein, σ(·) is the softmax function, and τ is the hyperparameter for smoothing the output distribution;
[0034] is the feature map of the last layer of the teacher model and the feature map of the last layer of the student model The mean square error between them is expressed by the formula:
[0035]
[0036] Wherein, r(·) is a mapping operation for aligning the dimensions between the feature maps.
[0037] Specifically, in S3, the independent training of the auxiliary classifier specifically includes: after the student completes each training iteration, fixing the parameters of the student model, and then training and updating the parameters of the auxiliary classifier according to the loss function
[0038] As can be seen from the above description, compared with the prior art, the advantages of the present invention are as follows:
[0039] During the process of image synthesis, the student model is also made to be one of the contributors. According to the current ability feedback by the student, the content of the synthesized image is adaptively adjusted to generate samples that are more difficult for the current ability of the student, avoiding the situation that overly simple samples prevent the student model from learning new knowledge all the time, and training the student more effectively to improve the final performance.
[0040] In the case of no original training data, the present invention adaptively adjusts the content of the synthesized image according to the current state of the student model, tailors the synthesized image for the student model, thereby training the student model more effectively to improve the final performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The method of the present invention will be further described below with reference to the drawings:
[0043] As Figure 1 shown, a data-free knowledge distillation method based on student feedback includes the following steps:
[0044] S1: Initialize the student model and add an auxiliary classifier after the feature extractor of the student model;
[0045] In a specific implementation, the auxiliary classifier is composed of two fully connected layers.
[0046] S2: Use the auxiliary classifier to feedback the current learning ability of the student model, and at the same time jointly train the noise vector and the generator according to the loss functions of the student feedback and the teacher feedback, so as to obtain the best synthesized image;
[0047] The process of using the auxiliary classifier to feedback the current learning ability of the student model specifically includes:
[0048] Randomly generate a noise vector z and input it into the generator network to obtain a synthesized image Then rotate the synthesized image by a certain angle, and input the rotated image into the student model of the feature extractor Φ, so as to obtain the obtained feature representation Input to the auxiliary classifier Utilize the output result of the auxiliary classifier Calculate the loss function to quantify the current learning ability of the student model, that is, the loss function of the student feedback, specifically as follows:
[0049]
[0050] Among them, k represents the class label of the self-supervised augmentation task, and the self-supervised augmentation task regards a self-supervised rotation task and the original image classification task as a joint task.
[0051] It is necessary to adaptively synthesize pictures according to the current ability of the student model. The ability of the model to capture semantic information can be well used as an indicator of the student model's ability, and an auxiliary task can reflect the degree to which the student model understands semantic information from the side. If only the self-supervised rotation task is used as the auxiliary task, the ability evaluation may be inaccurate. For example, the number "6" rotates 180° for the data "9", while it rotates 0° for the number "6" itself. Therefore, the self-supervised augmentation task is adopted as the auxiliary task of this method, so that the model can identify the category while recognizing the rotation angle.
[0052] During the picture synthesis process, the optimization objective generates difficult samples by expanding the loss value, that is, those samples that are difficult for the student model to understand in terms of semantic information.
[0053] The specific definition of the category of the self-supervised augmentation task is as follows:
[0054] Given that the total number of categories of the original image classification task is N, and the total number of categories of the self-supervised rotation task is M; assume that the synthesized picture is of class n in the image classification task, and its rotated version is of class m in the self-supervised rotation task, then its category in the self-supervised augmentation task is n*M + m.
[0055] However, if the picture synthesis only depends on the student feedback, it will lead to the distribution of the synthesized pictures being far from the distribution of the real pictures due to the lack of prior knowledge of the original data. Therefore, to consider the quality of the synthesized pictures, the picture synthesis also needs to depend on the teacher feedback, and its loss function is specifically as follows:
[0056]
[0057] Among them, expresses a one-hot hypothesis, that is, if the synthesized picture has the same distribution as the original training picture, then the output of the teacher model for the synthesized picture will be similar to the form of a one-hot vector. Therefore, is defined as the teacher model Output and the labels of predefined image classification tasks The cross entropy between them is expressed by the formula:
[0058]
[0059] Effectively utilize the statistics stored in the batch normalization layer of the teacher model as data prior information. So is defined as the l2 norm distance between the statistical features of the synthetic image and the real image, and its formula is expressed as:
[0060]
[0061] Where and are the mean and variance of the feature map of the synthetic image at the l-th layer of the teacher model; μ l and are the mean and variance of the feature map stored in the l-th layer of the teacher model, which represent the feature statistics of the real image.
[0062] Therefore, during the image synthesis process, the noise vector and the generator are jointly trained according to the loss functions of both student feedback and teacher feedback. The overall loss function is:
[0063]
[0064] Where α is a hyperparameter used to balance the weights of the two loss terms. In a specific implementation, α is set to 10.
[0065] S3: Use the synthetic images obtained in S3 to train the student model through knowledge distillation, and at the same time independently train the auxiliary classifier to learn the auxiliary task;
[0066] The overall loss function for training the student model through knowledge distillation is:
[0067]
[0068] Where β is a hyperparameter used to balance the weights of the three loss terms. In a specific implementation, β is set to 30; is the conventional loss term in the original image classification task, which is used to calculate the cross entropy between the output of the student model and the predefined labels;
[0069] is the KL divergence between the teacher and student outputs, and its formula is expressed as:
[0070]
[0071] Among them, σ(·) is the softmax function, and τ is the hyperparameter for smoothing the output distribution. In a specific implementation, τ is set to 20;
[0072] is the feature map of the last layer of the teacher model and the feature map of the last layer of the student The mean squared error between them is expressed by the formula:
[0073]
[0074] Among them, r(·) is a mapping operation. To align the dimensions between feature maps, in a specific implementation, the mapping operation consists of a three-layer convolutional block of 1×1, 3×3, and 1×1.
[0075] Independently training the auxiliary classifier specifically includes: after the student completes each training iteration, fixing the parameters of the student model, and then training and updating the parameters of the auxiliary classifier according to the loss function to improve the evaluation ability of the auxiliary classifier itself, so as to more accurately estimate the learning ability of the student during the image synthesis process.
[0076] S4: Repeat S2 and S3 until the student model is trained to convergence.
[0077] Experiments were conducted using the data-free knowledge distillation method based on student feedback on two publicly available image classification datasets, namely CIFAR10 and CIFAR100. Their image sizes are both 32×32. CIFAR10 consists of 10 classes, with each class including 5000 training images and 1000 test images. CIFAR100 consists of 100 classes, with each class including 500 training images and 100 test images. The training images are only used to train the teacher model to obtain a pre-trained teacher model, which is invisible to the student model. The test images are used to evaluate the prediction accuracy of the model. The teacher model selects the network structure of WRN-40-2, and the student model selects the network structure of WRN-16-1.
[0078] First, to prove the superiority of the present invention, comparisons were made with other existing technical methods. The experimental results are shown in Table 1. The present invention is significantly superior to other technical methods and obtains a student model with better performance.
[0079] Table 1 Prediction accuracies of student models of various methods
[0080]
[0081] Among them, DFAL comes from Chen, Hanting, Yunhe Wang, Chang Xu, Zhaohui Yang, Chuanjian Liu, Boxin Shi, Chunjing Xu, Chao Xu, Qi Tian, Data-Free Learning of Student Networks;
[0082] ZSKT comes from Micaelli, Paul, Amos J. Storkey, Zero-shot Knowledge Transfer via Adversarial Belief Matching;
[0083] ADI comes from Yin, Hongxu, Pavlo Molchanov, Zhizhong Li, José Manuel Arun Mallya, Derek Hoiem, Niraj Kumar Jha, Jan Kautz, Dreaming to Distill: Data-Free Knowledge Transfer via Deep Inversion;
[0084] CMI comes from Fang, Gongfan, Jie Song, Xinchao Wang, Chen Shen, Xingen Wang, Mingli Song, Contrastive Model Inversion for Data-Free Knowledge Distillation.
[0085] To further prove the importance of student feedback in the image synthesis process, an ablation experiment was conducted. That is, during the data synthesis process, when training the noise vector and the generator, the student feedback loss function was removed and only the teacher feedback was used. The experimental results are shown in Table 2, which shows that a better-performing student model was obtained with student feedback.
[0086] Table 2 Results of the ablation experiment
[0087]
[0088] Another embodiment of the present invention provides an implementation of image classification using a data-free knowledge distillation method based on student feedback, which includes:
[0089] Obtain a pre-trained teacher model: Obtain it from a model that has been trained on the training set of an image classification dataset;
[0090] Image synthesis: According to the feedback of the student model and the teacher model simultaneously, the noise vector and the generator are jointly trained. After the training is completed, the noise vector is input into the generator, and the obtained output is the required synthesized image;
[0091] Knowledge distillation: The student model is trained by knowledge distillation using the synthesized image, and at the same time, the auxiliary classifier used to feedback the student state in the image synthesis stage is independently trained. The two stages of the image synthesis process and the knowledge distillation process are alternately carried out until the student model converges.
[0092] Image classification: The image to be predicted is input into the trained student model, and the obtained output is the category to which the image belongs.
Claims
1. A data-free knowledge distillation method based on student feedback, comprising the following steps: S1: Initialize the student model and add an auxiliary classifier after the feature extractor of the student model; S2: Use the auxiliary classifier to feedback the current learning ability of the student model, and at the same time jointly train the noise vector and the generator according to the loss functions of the student feedback and the teacher feedback, so as to obtain the best synthetic image; S3: Use the synthetic image obtained in S2 to train the student model through knowledge distillation, and at the same time independently train the auxiliary classifier to learn the auxiliary task; S4: Repeat S2 and S3 until the student model is trained to convergence.
2. The method for knowledge distillation without data based on student feedback according to claim 1, wherein, The process of using the auxiliary classifier in S2 to feedback the current learning ability of the student model specifically includes: Randomly generate a noise vector z and input it into the generator network A synthesized image can be obtained Then for the synthesized image Rotate it by a certain angle, and input the rotated image into the feature extractor Φ of the student model so that the obtained feature representation can be input into the auxiliary classifier Using the output result of the auxiliary classifier calculate the loss function to quantify the current learning ability of the student model, that is, the loss function fed back by the student, specifically: where k represents the class label of the self-supervised augmentation task, and the self-supervised augmentation task regards a self-supervised rotation task and the original image classification task as a joint task.
3. A method for knowledge distillation without data based on student feedback according to claim 2, characterized in that, The specific definition of the class of the self-supervised augmentation task is as follows: Given that the total number of categories for the original image classification task is N, and the total number of categories for the self-supervised rotation task is M; assume that the synthetic image is of n categories in the image classification task, and its rotated version is of m categories in the self-supervised rotation task, then its category in the self-supervised augmentation task is n*M + m.
4. A method for knowledge distillation without data based on student feedback according to claim 3, characterized in that The loss function of the teacher feedback in S2 is specifically: Among them, is the output of the teacher model and the label of the predefined image classification task The cross-entropy between them is expressed by the formula as follows: The l2-norm distance between the synthetic image and the real image feature statistics, which is expressed by the formula: Among them, and are the mean and variance of the feature map of the synthesized image at the l-th layer of the teacher model; μ l and are the mean and variance of the feature map stored in the l-th layer of the teacher model, which represent the feature statistics of the real image.
5. A method for knowledge distillation without data based on student feedback according to claim 4, characterized in that , in S2, jointly train the noise vector and the generator according to the loss functions of the student feedback and the teacher feedback, and the overall loss function is: where α is a hyperparameter used to balance the weights of the two loss terms.
6. A data-free knowledge distillation method based on student feedback according to claim 5, characterized in that , in S3, the overall loss function for training the student model through knowledge distillation is: Among them, is a conventional loss term in the original image classification task, which is used to calculate the cross-entropy between the output of the student model and the predefined labels; Output the KL divergence between the teacher and the student, and its formula is expressed as: where σ(·) is the softmax function and τ is the hyperparameter for smoothing the output distribution; The feature map of the last layer of the teacher model and the feature map of the last layer of the student model The mean square error between them is expressed by the formula as follows: where r(·) is a mapping operation to align the dimensions between the feature maps.
7. A method for knowledge distillation without data based on student feedback according to claim 6, characterized in that , in S3, the independent training of the auxiliary classifier specifically includes: after the student completes each training iteration, fixing the parameters of the student model, and then training and updating the parameters of the auxiliary classifier according to the loss function .
Citation Information
Patent Citations
Visual interpretation method and system for deep neural network model.
CN112861933A
Adversarial Teacher-Student Learning for Unsupervised Domain Adaptation
US20190287515A1