A robust image classification method based on multi-model adversarial distillation
Through multi-model distillation and adversarial training techniques, the classification and defense capabilities of complex models are passed to lightweight networks, solving the problems of excessive memory and computing requirements of deep neural network models when deploying on resource-limited devices and threatened security, achieving efficient and robust image classification performance.
Patent Information
- Application Number
- CN202210488306.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-06
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-05-06
AI Technical Summary
Existing deep neural network models face the problem of excessive memory and computing requirements when deployed on devices with limited resources. At the same time, their security is also threatened by adversarial attacks, making it difficult to take into account both model accuracy and defense.
The deep neural network training method based on multi-model distillation combined with adversarial training is adopted to transfer the classification and defense capabilities of complex models to the lightweight network through knowledge distillation, and the adversarial training samples are generated through multi-step gradient descent method to improve the adversarial robustness of the model.
It realizes the model's adversarial robustness while maintaining excellent classification accuracy, avoiding the problem that single knowledge distillation cannot take into account accuracy and defense as well as the reduction accuracy and generalization of the adversarial training, making the student model performance obtained by training more comprehensive.
Smart Images

Figure CN114842257B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a robust image classification method, in particular to an image classification method based on a deep neural network, and belongs to the field of deep learning and artificial intelligence. Background Art
[0002] In recent years, with the rapid development of deep learning technology, multi-layer neural network structures have emerged one after another, and deep neural networks have achieved unprecedented results in image classification tasks. The design of neural network models should achieve a good trade-off between model performance and model complexity. However, in practice, it is difficult for researchers to determine the right balance, so they prefer to choose neural network models that are over-parameterized, have sufficient expressive power, and are easy to optimize. Although neural networks can have better expressive power as their depth increases and their structures become more complex, they also face more resource consumption. If it is transplanted to the edge or mobile devices considering actual landing applications, it will be constrained by many aspects such as large memory usage, high computing power, and high energy consumption. In order to solve the problem that the huge memory and computing requirements of deep neural networks seriously hinder their deployment in resource-limited devices and reduce the pressure of model memory and inference training, knowledge distillation (Reference [1]: Geoffrey Hinton, Oriol Vinyals and Jeff Dean, Distilling the Knowledge in a Neural Network. NIPS Deep Learning Workshop, 2014, Geoffrey Hinton, Oriol Vinyals and Jeff Dean, Knowledge Distillation in Neural Network, NIPS Deep Learning Workshop, 2014) has emerged as an effective method for compressing large models into small models. Knowledge distillation, network pruning and parameter quantization are both mainstream methods for model lightweighting.
[0003] At the same time, with the continuous development of deep neural network models, their security is also seriously threatened. The emergence of adversarial attack algorithms has posed a serious threat to the application of deep neural networks. Adversarial attacks for image classification add perturbations that are difficult for human vision to detect to benign samples, causing the classifier to make incorrect judgments. This has become a major obstacle to the large-scale deployment of deep learning models in production. Against attacks from adversarial samples, the most widely used defense strategy is adversarial training. Training is performed on perturbed data to improve the robustness of the model against malicious samples. The idea of adversarial training was first proposed in 2005 (reference [2]: Daniel Lowd and Christopher Meek. Adversarial learning, In Proceedings of the eleventh ACM SIGKDD international conference on Knowledge discovery in data mining, 2005, namely Daniel Lowd and Christopher Meek. Adversarial learning, Adversarial learning, In Proceedings of the eleventh ACM SIGKDD international conference on Knowledge discovery in data mining, 2005.). Today, there are many forms of optimization, for example, PGD adversarial generation of adversarial samples to participate in training to improve the model's adversarial robustness (reference [3]: Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu, Towards deep learning models resistant to adversarial attacks. In Proc. of ICLR, 2018, namely Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu, Deep Learning Models Resilience to Adversarial Attacks, In Proc. of ICLR, 2018.)
[0004] As a requirement for deployment applications, it is of great significance to design lightweight and robust deep neural network models in image classification tasks. Therefore, distillation learning and adversarial training, as methods to reduce model size and improve model robustness, respectively, provide technical support for the robust application of lightweight deep neural networks. In the current distillation learning framework, a large model with good performance is usually used as a teacher to distill a lightweight small model. The student model after knowledge distillation has relatively simple functions and it is difficult to have excellent performance in both accuracy and defense. The adversarial training framework can effectively bring strong defense to the model, but it will cause the model to lose model accuracy and reduce the generalization of the model. Therefore, a training method that can take into account both model accuracy and security is an urgent need.
[0005] The present invention combines model distillation learning and adversarial training techniques to achieve the following two goals: (1) obtain a lightweight model with excellent classification performance; and (2) ensure the security and robustness of the model while ensuring excellent classification accuracy. Summary of the invention
[0006] The present invention aims to overcome the defects of the prior art in that the model structure is complex and the generalization ability of the classification model is not high, and provides a robust image classification method based on multi-model adversarial distillation.
[0007] In order to effectively balance the accuracy and defense of the model in the training of lightweight deep neural networks, the present invention proposes a deep neural network training method based on multi-model distillation combined with adversarial training. This method transfers the classification and defense capabilities in the complex model to the lightweight network through the information transmission channel of the knowledge distillation model, and obtains an image classifier with excellent generalization and adversarial properties.
[0008] The knowledge distillation method of the present invention updates the student model parameters based on the gap in the Logit output of the teacher and student models as the loss function term; the adversarial training samples are generated by a multi-step gradient descent method. Finally, through a multi-model distillation framework, different weight allocations are used to distinguish the importance of the teacher model knowledge and perform distillation training.
[0009] The technical solutions adopted by the present invention to achieve the above-mentioned invention objects are as follows:
[0010] S1: pre-trained complex model;
[0011] Normally train the complex model under the given training data set X and obtain model T 1 Complex models with excellent classification capabilities should be selected based on prior knowledge.
[0012] S2: Generate adversarial training samples;
[0013] Given a data set X, perturbations are added to each sample through multiple iterations according to the gradient of the loss function to obtain adversarial training samples. Adversarial samples are generated according to formula (1).
[0014]
[0015] where x t is the original sample; α is the perturbation coefficient, which determines the step size in each iteration; sign(·) is the sign function, which specifies the direction in which the image pixels change; J(x,y) is the loss function of the model; is the gradient of the loss function with respect to the image pixel values.
[0016] S3: Adversarial training of complex models;
[0017] Use the perturbation samples and correct labels generated in step S2 as adversarial training data. Select a complex model framework and directly adopt adversarial training to generate the teacher model T 2 In step 3, each new iteration of the adversarial training process regenerates a new batch of perturbation samples according to formula (1) as training samples for this round of adversarial training.
[0018] S4: Knowledge distillation;
[0019] The step S4 specifically includes:
[0020] S4.1: Select the lightweight model structure as the student model S;
[0021] S4.2: Input the training sample x(x∈X) into the student model S and the teacher model T 1 In the above equation, we get the Logit outputs of the two models. The real labels are used as hard labels for the student model training, and the outputs of the teacher model are used as soft labels for the student model. The loss is calculated according to formula (2): T1 , where lamba is the weight to measure the importance of the loss function, and loss is calculated based on the cross entropy loss nat ;
[0022] loss T =(out s -out T ) 2 *lamba (2)
[0023] S4.3: Generate adversarial samples based on the student model S using formula (1) during knowledge distillation Same as step S4.2, the adversarial dataset Input teacher model T 2 And the student model S gets the loss T2and loss adv .
[0024] S4.4: Calculate the total loss value according to formula (3).
[0025] LOSS=loss nat +loss adv +loss T1 +loss T2 (3)
[0026] During the experiment, users can change lamba to make adjustments based on the different teacher models and the different requirements for generating lightweight model functions.
[0027] Specifically, the method described in the present invention has the following beneficial effects:
[0028] Since it is difficult to directly find a model with excellent classification accuracy and defensiveness as a teacher model, it is difficult to simultaneously improve the classification accuracy and defensiveness of the student model in knowledge distillation. The training strategy described in the present invention adopts adversarial training and multi-model distillation methods to significantly improve its adversarial robustness while taking into account the accuracy of the model. It avoids the problem that single knowledge distillation cannot take into account both accuracy and defensiveness, and that adversarial training significantly reduces accuracy and generalization, making the performance of the trained student model more comprehensive. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 A basic knowledge distillation framework.
[0030] Figure 2 A framework for robust deep neural network training based on multi-model knowledge distillation. DETAILED DESCRIPTION
[0031] The specific implementation of the present invention is further described below with reference to the accompanying drawings and taking CIFAR100 image data as an example.
[0032] Reference Figure 2 , starting with complex model pre-training, the specific steps are as follows:
[0033] S1: pre-trained complex model;
[0034] The complex model framework uses Densenet-121, and the training data set uses CIFAR100. The data set consists of 100 categories of 32×32 3-channel RGB color images, including 50,000 training sets and 10,000 test sets. The model T is trained 1 .
[0035] S2: Generate adversarial training samples;
[0036] According to formula (1), the PGD method is used to set the perturbation coefficient, the step size in each iteration and the number of iterations, and the training set in the same data set in step S1 is used to generate adversarial training samples.
[0037]
[0038] S3: Adversarial training of complex models;
[0039] Use the perturbed images and correct labels generated in step 2 as adversarial training data. The model used is Densenet-121 without pre-training, and adversarial training is directly adopted to generate model T 2 In step S3, each new iteration of the adversarial training process regenerates a new batch of perturbation samples according to formula (1) as the training samples for this round.
[0040] S4: Knowledge distillation;
[0041] S4.1: ResNet-34 is selected as the lightweight student model structure, whose network depth and parameter quantity are much smaller than the teacher model Densenet-121.
[0042] S4.2: Input the training sample x(x∈X) into the student model S and the teacher model T 1 In the above equation, we get the Logit outputs of the two models. The real labels are used as hard labels for the student model training, and the outputs of the teacher model are used as soft labels for the student model. The loss is calculated according to formula (2): T1 , where lamba is the weight to measure the importance of the loss function, and loss is calculated based on the cross entropy loss nat ;
[0043] S4.3: Generate adversarial samples based on the student model S using formula (1) during knowledge distillation Same as step S4.2, the adversarial dataset Input teacher model T 2 And the student model S gets the loss T2 and loss adv .
[0044] S4.4: Calculate the total loss value according to formula (3).
[0045] loss T =(out s -out T ) 2 *lamba (2)
[0046] LOSS=loss nat +loss adv+loss T1 +loss T2 (3)
[0047] As described above, the present invention is a robust image classification method based on multi-model adversarial distillation. Compared with the existing traditional adversarial training methods that sacrifice standard accuracy to improve model adversarial defense, the present invention combines knowledge distillation and adversarial training technology to improve model performance in four common indicators: model lightweight, classification accuracy, model robustness and generalization. While greatly improving the model's adversarial defense, it can maintain or even improve the model's standard classification accuracy. Our method brings new ideas to the training of lightweight deep learning image classification networks, and we hope that it can help researchers better improve model performance so as to deploy safe and reliable application deep learning models in situations where hardware resources are limited, such as at the edge.
Claims
1. A robust image classification method based on multi-model adversarial distillation, characterized by: The extraction method comprises the following steps: S1: pre-trained complex model; Under the given training data set X, the complex model is trained normally to obtain model T1; the complex model should select a complex model with excellent classification ability based on prior knowledge; S2: Generate adversarial training samples; Given a data set X, perturbations are added to each sample through multiple iterations according to the gradient of the loss function to obtain adversarial training samples. Adversarial samples are generated according to formula (1); where x t is the original sample; α is the perturbation coefficient, which determines the step size in each iteration; sign(·) is the sign function, which specifies the direction in which the image pixels change; J(x,y) is the loss function of the model; is the gradient of the loss function with respect to the image pixel value; S3: Adversarial training of complex models; Use the perturbed samples and correct labels generated in step S2 as adversarial training data; select a complex model framework, directly adopt adversarial training, and generate the teacher model T2; in step S3, each new iteration of the adversarial training process regenerates a new batch of perturbed samples according to formula (1) as training samples for this round of adversarial training; S4: Use multiple teacher models to perform knowledge distillation to obtain a lightweight student model; Step S4 specifically includes: S4.1: Select the lightweight model structure as the student model S; S4.2: Input the training sample x, x∈X, into the student model S and the teacher model T1, and obtain the Logit output of the two models respectively; the true label is used as the hard label for the student model training, and the output of the teacher model is used as the soft label for the student model; the loss is calculated according to formula (2) T1 , where lamba is the weight to measure the importance of the loss function, and loss is calculated based on the cross entropy loss nat ; loss T =(out s -out T ) 2 *lamba (2) S4.3: Generate adversarial samples based on the student model S using formula (1) during knowledge distillation Same as step S4.2, the adversarial dataset Input teacher model T2 and student model S to get loss T2 and loss adv ; S4.4: Calculate the total loss value according to formula (3); LOSS=loss nat +loss adv +loss T1 +loss T2 (3) In step S4.2, the user determines the specific size of the lamba weight according to the function of the model to be generated. The user can generate a strong adversarial robustness model or a high classification accuracy model by different weight allocations.
Citation Information
Patent Citations
Neural network black box aggressive defense method based on knowledge distillation
CN111027060A
Knowledge distillation method and system based on multi-student discussion
CN114049513A