Back door defense method based on adversarial pruning and knowledge distillation
Through the combination of anti-pruning and knowledge distillation, the problem of effectively defending backdoor attacks without sacrificing model accuracy is solved, and the recovery of model accuracy and significant improvement of defense capabilities is achieved.
Patent Information
- Application Number
- CN202510018839.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to effectively defend against backdoor attacks without sacrificing model accuracy, especially the possible loss of model accuracy during pruning.
The method of anti-pruning combined with knowledge distillation is adopted to enhance the model's defense ability by anti-pruning marking and erasing suspicious neurons, and the knowledge distillation technology is used to restore the model's accuracy sacrificed due to pruning.
While ensuring the accuracy of the model, it significantly improves the model's defense capabilities, can effectively defend against backdoor attacks, and improves the stability and robustness of the model when facing complex attacks.
Smart Images

Figure CN119940471A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning security, and in particular to a backdoor defense method for neural networks. Background Art
[0002] Deep Neural Networks (DNNs) have been widely used in target detection, natural language processing, and smart healthcare. Its success is mainly due to the huge model parameters and massive training data, but this is extremely difficult for some resource-constrained users because it requires a lot of manpower and material resources. So people began to obtain computing power and training data from third-party platforms, such as Github and Hugging face.
[0003] However, deep learning models are highly sensitive to training data. If an attacker impersonates a third-party platform and adds some poisoned data to the training data, it may cause the model to be implanted with a backdoor. A backdoor attack will make the model perform well on clean data, but for poisoned data containing triggers, it will cause the model to output a malicious result pre-specified by the attacker. The essence of a backdoor attack is to use the redundant classification ability of the neural network model to train the model to learn the potential relationship between the trigger and the target category. In order to defend against backdoor attacks, the current mainstream methods are mainly divided into two categories: detection-based methods and erasure-based methods. Detection-based methods aim to identify whether there is a backdoor attack and what the backdoor trigger and target label are. These methods are usually based on anomaly detection, cluster analysis, adversarial perturbation and other technologies to discover abnormal behavior or abnormal parameters in the model. Although detection-based methods can provide useful information, they cannot repair infected models, so they need to be used in conjunction with erasure-based methods. Erasure-based methods aim to remove the impact of backdoor attacks from the model to restore the normal function of the model.
[0004] Knowledge distillation is a common machine learning model compression technique, whose main goal is to transfer the knowledge learned from a large, complex model (usually called the teacher model) to a smaller, more concise model (i.e., the student model), thereby improving the performance and generalization ability of the student model. Through this process, the student model can learn the efficient representation and decision-making mode of the teacher model without significantly increasing the computational complexity and resource consumption. The core idea of knowledge distillation is to help the student model improve its understanding and classification ability of data by transferring the deep knowledge of the teacher model, especially its internal feature representation and decision logic.
[0005] Specifically, knowledge distillation is not just about passing the output labels of the teacher model directly to the student model, but helps the student model capture more complex feature relationships and potential patterns through the teacher model's intermediate layer output, probability distribution or attention mechanism information. This knowledge transfer enables the student model to make accurate judgments like the teacher model when facing complex data, while maintaining lower computational overhead.
[0006] In the present invention, the application of knowledge distillation is used to solve the problem of model accuracy loss that may be caused by the pruning process. In the process of defense against pruning, the pruning operation will inevitably lead to the loss of accuracy of some normal neurons due to the need to remove neurons that may be affected by backdoor attacks. This loss is usually manifested as a decrease in classification accuracy or an unstable performance of the model on the input data. In order to overcome this problem, the present invention adopts a knowledge distillation method to restore the model accuracy sacrificed by pruning.
[0007] Through knowledge distillation, the student model can learn more robust and effective feature representations from the teacher model, thereby compensating for the performance degradation caused by pruning. Between the pruned model (teacher model) and the unpruned backdoor model (student model), knowledge distillation helps the student model better understand and process data by aligning the output distribution of the teacher model and the student model. In this way, the student model can not only recover in terms of accuracy, but also enhance its defense capabilities when facing different types of attacks.
[0008] In addition, knowledge distillation can also help reduce the computational resources required by pruning operations. During the pruning process, although the network parameters and computational complexity are reduced, the model's performance on certain tasks may also deteriorate. By using knowledge distillation, the student model can efficiently inherit the knowledge of the teacher model, allowing it to maintain high performance and accuracy while reducing computational complexity.
[0009] In general, the present invention not only recovers the accuracy loss caused by the removal of neurons during the pruning process, but also enhances the defense capability of the model by introducing knowledge distillation technology. Knowledge distillation enables the pruned model to ensure performance while improving the security and reliability of the model in practical applications. Summary of the invention
[0010] The technical problem to be solved by the present invention is to provide a backdoor defense framework that can effectively defend against backdoor attack methods without sacrificing model accuracy.
[0011] The technical solution adopted by the present invention to solve the technical problem is: adversarial pruning combined with knowledge distillation, including the following steps:
[0012] S1, add a mask to all neurons, the mask value is all 1, and initialize the adversarial neuron perturbation;
[0013] S2. Input the initialization data obtained in step S1 into the model, and train the mask with clean data under the condition of adversarial neuron perturbation. Since the backdoor neurons are more likely to collapse than normal neurons under the condition of adversarial perturbation, the backdoor neurons can be marked in subsequent training.
[0014] S3. After a certain number of training rounds, the threshold value is determined by observing the performance of the model at this time, and the trained mask is compared with it. Neurons with values less than the threshold value are marked (i.e., the mask value is set to 0 as a mark), and suspicious neurons are erased without affecting the accuracy of the model too much;
[0015] S4, the model with the best performance after adversarial pruning is used as the teacher network, and the original model without pruning is used as the student model;
[0016] S5. Perform weighted alignment on the intermediate attention layer of the student model and the intermediate attention layer of the teacher model, and judge whether the defense effect meets expectations through the accuracy and attack success rate, so as to achieve the defense effect.
[0017] Specifically, in step S1, the input data is an image of size 3×224×224 pixels, which is usually an RGB image with a height and width of 224 pixels and 3 color channels (red, green, and blue). In order to achieve adversarial pruning and backdoor defense, masks are first set for all neurons, and the initial value of the mask is 1. The role of the mask is to mark the neuron area that needs attention in the network, and to adjust and update it through the subsequent training process. The mask value of 1 means that these neurons are not disabled in the initial stage.
[0018] In actual operation, in order to achieve adversarial pruning and ensure network stability, step S1 also requires adding a batch normalization (BN) layer in each layer. The role of the batch normalization layer is to standardize the input neurons to reduce internal covariate shift, thereby accelerating training and improving the stability of the model. The number of neurons in this batch normalization layer is consistent with the original number of neurons in the layer, ensuring that the structure of the network will not change unnecessarily due to the addition of the batch normalization layer.
[0019] After adding the batch normalization layer, in order to make the neural network adapt to disturbances during training, the initial value of the batch normalization layer is set to all 1. This means that in the initial state, the scaling factor (gamma) and offset (beta) of all neurons in the batch normalization layer are set to 1, so as to ensure that the output of the network will not be affected too much in the early stage. The purpose of setting it to all 1 is to prevent the disturbance in the early stage of training from causing too much disturbance to the network performance.
[0020] In addition, in order to achieve perturbation training, all neurons in the batch normalization layer will be perturbed, and the perturbation will be continuously optimized during the training process. The initialization value of the perturbation is usually random, and a small value is usually selected to ensure that the perturbation does not cause drastic changes in network performance in the early stages of training. The role of the perturbation is to disrupt the weights and activations of the neural network so that the network can learn more robust features in adversarial environments. Through training, the perturbation will be gradually adjusted according to the feedback of the loss function, thereby enhancing the network's ability to resist backdoor attacks. The loss function for training is:
[0021]
[0022] Specifically, in step S2, the value of the mask is updated using the gradient descent method to optimize the performance of the model and enhance its defense against backdoor attacks. The gradient descent method is a common optimization algorithm that gradually reduces the loss function by continuously adjusting the parameters of the mask and perturbation. The core purpose of this method is to minimize the loss of the activation value of the neural network under the perturbation during the training process, thereby improving the robustness of the model in the face of potential attacks. During the training process, the update rule of the gradient descent method optimizes the value of the mask as an adjustable parameter. Specifically, the value of the mask will be continuously adjusted as the training progresses so that the network can more effectively "perceive" potential backdoor triggers. Through the gradient descent algorithm, the adjustment of the mask value not only helps to reduce the error in the loss function, but also ensures that the model maintains its original function while enhancing its sensitivity to backdoor attacks, making the trigger easier to be recognized by the neural network. Its formal definition is shown in formula (2). By minimizing the total loss function, especially by reducing the loss of the activation value under the perturbation condition, the gradient descent method can effectively optimize the mask, so that the network can better identify and defend against backdoor attacks under the perturbation. At the same time, the updated mask not only improves the model's performance in defending against attacks, but also ensures that the model maintains efficient and accurate functionality when processing normal data. In this process, the mask update process actually implicitly enhances the "perception ability" of the neural network. By adjusting the value of the mask, the network gradually learns how to detect backdoor triggers by strengthening the response of sensitive neurons. When the network can more easily identify triggers through learned masks and perturbation strategies, the robustness and defense capabilities of the model are significantly improved, while still being able to maintain its original functionality and ensure the completion of normal tasks.
[0023]
[0024] Specifically, the perturbation update method used in step S4 is attention distillation, the main purpose of which is to restore the accuracy of the model and further remove the backdoor problem. Attention distillation guides the student network to imitate the attention distribution of the teacher network, so that the student network can learn a more robust feature representation, thereby improving its ability to resist attacks. In this process, the student network not only pays attention to the final output of the teacher network, but also pays special attention to the attention mechanism of the intermediate layers of the teacher model. By weighted alignment of multiple intermediate layers of the student network, the student network can gradually adjust its attention distribution and learn the effective feature selection and attention mode of the teacher network, thereby recovering the accuracy loss caused by pruning or perturbation. At the same time, the student network gradually enhances its ability to recognize backdoor triggers during the learning process, and can more effectively remove potential backdoor attacks. This attention alignment weighted mechanism enables the student network to converge faster and improve robustness while ensuring accuracy, and finally obtains a powerful model that can both accurately classify and effectively defend against backdoor attacks. Its formal definition is shown in formula (3).
[0025]
[0026] The beneficial effects of the present invention are as follows:
[0027] The backdoor defense method described in the present invention shows significant advantages in backdoor erasure. Through an effective erasure strategy, the present invention can eliminate the backdoor in the backdoor model, thereby destroying the strong correlation between the backdoor neuron and the trigger, thereby significantly improving the defense effect of the model. This process ensures the removal of backdoor features, so that the model can maintain efficient robustness in the face of potential attacks. At the same time, the present invention combines knowledge distillation technology to effectively restore the precision loss that may be caused during the pruning process. By utilizing the knowledge of the teacher model during the model training process, the student model can restore the precision loss caused by the pruning operation as much as possible while maintaining high defense performance. Especially in the face of heat map analysis, the model processed by the present invention does not show obvious differences when feeding poisoned data and clean data, which shows that the sensitivity of the model to the backdoor is greatly reduced and the defense capability is significantly improved. In the backdoor attack scenario, the model processed by the defense of the present invention can successfully defend against various poisoned data and backdoor attack methods, further verifying its powerful attack defense capability. Therefore, the backdoor defense method of the present invention can not only improve the performance of the model with a small loss of accuracy, but also significantly improve the defense effect, ensuring the stability and robustness of the model when facing complex attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a technical flow chart of the present invention;
[0029] Figure 2A schematic diagram of a backdoor attack application scenario of the present invention;
[0030] Figure 3 It is a flowchart of the backdoor defense method for neural networks of the present invention;
[0031] Figure 4 Schematic diagram of the attention alignment weighting process of the present invention;
[0032] Figure 5 This is a schematic diagram showing the effect of the attention weighted adjustment mechanism of the present invention. DETAILED DESCRIPTION
[0033] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0034] Figure 1 is a technical flow chart of the present invention. The flow chart describes a method for removing backdoor attacks in deep neural networks through perturbation generation, attention extraction and distillation techniques. First, input training data and pre-trained models with backdoors, and then generate perturbations to change the input data to avoid backdoor activation. By extracting and aligning the attention features of the model, it focuses on the features of normal data rather than backdoor features. Finally, the model is optimized through attention distillation and regularization, the backdoor effect is removed and a clean model is output to ensure that it can still accurately classify data with backdoors.
[0035] Figure 2 It is a schematic diagram of the backdoor attack application scenario of the present invention. With the continuous increase in the scale of model parameters and the scale of training data, the use of third-party pre-trained models and open source data sets has become more and more common for fine-tuning pre-trained models. In this process, although the user can control the fine-tuning process and model structure by himself, it is impossible to conduct a comprehensive inspection of each open source data, which provides the possibility for the backdoor attack of the present invention. The attacker uses the method of the present invention to upload poisoned data to the open source data, so that the model suitable for downstream tasks trained based on these data is implanted with a backdoor. This is an extremely common scenario, so the attack of the present invention is of great significance to the security research of deep neural networks.
[0036] Figure 3This is a flowchart of the backdoor defense method for neural networks of the present invention. It mainly includes the following steps: S1, initialization of mask and perturbation, S2, training model, updating mask and perturbation, S3, pruning suspicious neurons, S4, introducing attention distillation to erase backdoors. In the model initialization stage, the defender needs to add a mask to all neurons, the mask value is all 1, and add perturbations to all neurons.
[0037] After initialization, the input clean data is sent to the model for training, with the goal of updating the perturbed mask. The core idea is that the backdoor neurons will be more likely to collapse under perturbations, causing the model to misclassify clean samples. Here we can use masks as markers to mark suspicious neurons, making it easier for the model to find triggers, ensuring the efficiency of backdoor defense. However, such pruned models may sacrifice the accuracy of classifying clean data, as suspicious neurons may also be clean neurons, resulting in a decrease in accuracy. Therefore, the present invention introduces the knowledge distillation method. This will not only improve the accuracy, but also further reduce the success rate of attacks, which ensures the robustness of backdoor defense.
[0038] Figure 4 is a schematic diagram of the attention weighted alignment mechanism of the present invention, which is Figure 2 A refinement of the flowchart of the knowledge distillation method in . The computable trigger needs to find a "potential" neuron - it will undergo a significant change in activation value due to the presence of the trigger, making it easier for the model to perceive the presence of the trigger. The present invention proposes an improved neuron search method, which iteratively calculates the mean square error between the activation value output by each neuron and the target activation value as the loss function, and uses the gradient descent algorithm to update the value of the trigger so that the value of the loss function gradually decreases. At this time, the activation value of the neuron will gradually increase and approach the target activation value. Its formal definition is shown in formula (3), and the specific process of the present invention is shown in Algorithm 1.
[0039]
[0040] Figure 5 This is a schematic diagram of the ablation study of the attention alignment weighted mechanism of the present invention. Taking BadNet as an example, the ACC and ASR indicators tested under clean data and poisoned data found that the mechanism can not only improve a certain amount of accuracy but also make the model converge quickly.
[0041] Table 1 shows the test results of different backdoor defense methods on different backdoor attack methods. The test indicators include accuracy (ACC) and attack success rate (ASR), which are defined as shown in formula (4) and formula (5). The attack effect of the present invention has achieved the best results on each data set.
[0042] Table 1
[0043]
[0044] Among them, TP (True Positives) represents the number of samples correctly predicted as positive; TN (True Negatives) represents the number of samples correctly predicted as negative; FP (False Positives) represents the number of samples incorrectly predicted as positive; FN (False Negatives) represents the number of samples incorrectly predicted as negative; card (D poison ) represents the poisoned sample set D poison The number of elements in .
[0045]
[0046] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A backdoor defense method based on adversarial pruning and knowledge distillation, characterized in that: The method comprises the following steps: S1, add a mask to all neurons, the mask value is all 1, and initialize the adversarial neuron perturbation; S2. Input the initialization data obtained in step S1 into the model, and train the mask with clean data under the condition of adversarial neuron perturbation. Since the backdoor neurons are more likely to collapse than normal neurons under the condition of adversarial perturbation, the backdoor neurons can be marked in subsequent training. S3. After a certain number of training rounds, the threshold value is determined by observing the performance of the model at this time, and the trained mask is compared with it. Neurons with values less than the threshold value are marked (i.e., the mask value is set to 0 as a mark), and suspicious neurons are erased without affecting the accuracy of the model too much; S4, the model with the best performance after adversarial pruning is used as the teacher network, and the original model without pruning is used as the student model; S5. Perform weighted alignment on the intermediate attention layer of the student model and the intermediate attention layer of the teacher model, and judge whether the defense effect meets expectations through the accuracy and attack success rate, so as to achieve the defense effect.
2. The backdoor defense method based on adversarial pruning and knowledge distillation as claimed in claim 1, characterized in that: The defense mechanism is implemented by setting a mask for all neurons. In practice, the defense strategy only requires adding a batch normalization layer to each layer. The number of neurons in the batch normalization layer should be consistent with the number of neurons added in this layer, so as to ensure that the structure and dimension of the network remain consistent. In order to ensure the effectiveness of the perturbation, the values of all batch normalization layers are initialized to all 1s so that there will not be too much interference on the output of the network during the training process. On this basis, perturbations are added to the neurons in all batch normalization layers and trained. The initialization value of the perturbation usually consists of small random values to ensure that the network does not deviate excessively in the early stage of training, and at the same time can effectively capture the attack mode, thereby enhancing the robustness of the model. In the subsequent training process, these perturbations will be continuously adjusted according to the feedback of the loss function, so as to achieve the purpose of adversarial pruning, gradually reduce the impact of backdoor attacks and improve the model's defense capabilities against unknown attacks. In addition, combined with knowledge distillation technology, the pre-trained model knowledge can be transferred to the target model. In this way, the model can not only effectively remove the backdoor, but also maintain its original performance.
3. The backdoor defense method based on adversarial pruning and knowledge distillation as claimed in claim 1, characterized in that: Step S2 uses the gradient descent method to update the values of the perturbation and mask to ensure the effectiveness of the defense strategy. Specifically, in this step, the parameters of the perturbation and mask are first updated by the gradient descent optimization algorithm, so that the perturbation can effectively affect the output of the neural network and gradually approach the characteristics of the backdoor neurons through the training process. In addition, the output performance of each neuron on clean data after the perturbation is added will be recorded during the training process. In particular, by testing the performance of the neural network on clean data, the response changes of neurons under the influence of perturbations can be observed. Since backdoor neurons often exhibit abnormal behaviors under certain conditions, especially on clean data, they are more susceptible to perturbations, causing neuron output to collapse or misclassify. This abnormal behavior provides clues for identifying potential backdoor neurons. By observing the response of neurons under perturbations, the model can quickly identify neurons that behave unstable or misclassify when faced with clean data, which are likely to be carriers of backdoor attacks. Therefore, the process of perturbation updates not only helps the implementation of adversarial pruning, but also effectively locks and removes suspicious neurons, thereby effectively clearing backdoor attacks. By continuously updating the gradient descent and monitoring the performance of neurons, we can quickly focus on neurons that behave abnormally under perturbations. These abnormal neurons are marked as potential backdoor neurons and removed in the subsequent pruning step. This process makes backdoor defense more efficient and accurate while keeping the model's performance stable on clean data. The loss function for training is:
4. The method based on adversarial pruning and knowledge distillation as claimed in claim 1, characterized in that: An adaptive threshold is set to accurately determine which neurons should be marked and eventually pruned. The threshold determination method uses an adaptive strategy and dynamically adjusts according to the performance of the model during training to ensure that normal neurons and neurons that may be affected by backdoor attacks can be effectively distinguished. Specifically, the threshold calculation is based on two key indicators: classification accuracy (ACC) and attack success rate (ASR). ACC reflects the performance of the model on clean data, while ASR measures the backdoor effect of the model under attack conditions. Through the comprehensive analysis of these two indicators, the threshold value can be adaptively adjusted to fully consider the impact of different attacks during the defense process. When ACC is high and ASR is low, the model's defense effect is better. At this time, the threshold value can be set to a higher value to ensure that only those neurons that significantly deviate from normal performance are marked and pruned. On the contrary, when ACC is low or ASR is high, the threshold value may be lowered to more sensitively detect neurons affected by the backdoor. During the training process, the mask value of each neuron is compared with the current threshold value. If the mask value is less than the threshold value, the neuron is marked as a suspicious neuron. Labeled neurons often exhibit abnormal behavior, such as misclassification when faced with clean data, or obvious backdoor characteristics when attacked. These neurons are likely to be carriers of backdoor attacks, so they need to be removed from the network through pruning operations. Once neurons are marked as suspicious, they will be removed in the subsequent pruning step. Pruning not only helps remove possible backdoor neurons, but also reduces the complexity of the network and improves the efficiency of the model. At the same time, by using adaptive thresholds, the defense process can be flexibly adjusted and the accuracy and effect of pruning can be optimized according to actual conditions. Finally, combined with knowledge distillation technology, the model can retain its original classification performance while effectively eliminating the risk of backdoor attacks. Define the optimization goal:
5. The method based on adversarial pruning and knowledge distillation as claimed in claim 1, characterized in that: During the defense process, adversarial pruning is first used to remove suspicious backdoor neurons from the backdoor model, resulting in a pruned model. The pruned model is used as a teacher model, as a "teacher" to guide the student model's learning. The teacher model represents the ideal performance after pruning and removing backdoor neurons, with high robustness and low attack success rate (ASR). The original backdoor model without any defense processing is regarded as the student model. The student model learns from the teacher model through knowledge distillation during the training process. The core idea of knowledge distillation is to transfer the knowledge of the teacher model to the student model, helping the student model to learn the robustness and robustness of the teacher model while maintaining its original performance. In this process, the output of the teacher model is used as the target of the student model, especially in the early stages of model training, the output of the teacher model contains not only the category label, but also the deep feature representation and probability distribution information. The student model gradually learns the behavioral characteristics of the teacher model by minimizing the difference between its output and the output of the teacher model. In this way, the student model can effectively learn how to maintain high classification accuracy while reducing the risk of being affected by backdoor attacks. As an ideal version after defense, the teacher model provides a clear learning goal, allowing the student model to self-adjust during the learning process and gradually improve its performance on clean data sets. At the same time, the student model can also adapt to possible disturbances during the learning process, further enhancing its defense capabilities in the face of unknown attacks. Therefore, by using the pruned model as the teacher model and the unprocessed backdoor model as the student model, combined with knowledge distillation technology, the defense strategy can effectively remove backdoor attacks without losing too much original performance, thereby improving the security and robustness of the overall model. This method not only effectively removes backdoor neurons, but also enhances the student model's defense capabilities against backdoor attacks, achieving an ideal defense effect.
6. The method based on adversarial pruning and knowledge distillation as claimed in claim 1, characterized in that: Combined with knowledge distillation technology, a robust network is distilled to enhance the defense effect and recover the accuracy loss caused by pruning suspicious neurons. The main purpose of introducing knowledge distillation is to compensate for the performance degradation caused by the accuracy loss of normal neurons during the pruning process. Specifically, the knowledge distillation method trains the pruned model (i.e., teacher model) with the original backdoor model (i.e., student model) so that the student model can learn robust representations and decision-making strategies from the teacher model, thereby recovering the accuracy loss caused by the pruning process. To achieve this goal, the attention layer of the student model is first aligned with the attention layer of the teacher model. Attention alignment refers to adjusting the attention mechanism of the student model to align it with the attention distribution of the teacher model when processing the input data. This step is achieved by introducing an attention alignment weighting mechanism. This mechanism uses a weighting strategy to enable the student model to quickly move closer to the attention distribution of the teacher model during training, thereby better capturing the key features and patterns learned by the teacher model. Specifically, the attention layers of the student model and the teacher model will adjust the attention weight of the student model by calculating the similarity or distance between the two (such as KL divergence or mean square error). In this process, the student network continuously learns the attention distribution of the teacher network, so that the student model can gradually understand the importance of the teacher model in feature selection, and on this basis, process the input data more reasonably. Since the attention alignment weighting mechanism can make the student network closer to the attention distribution of the teacher network faster during training, it not only improves the training efficiency of the model, but also accelerates the convergence process of the model. In addition, the attention alignment weighting mechanism can effectively avoid the accuracy loss caused by normal neurons during pruning. In the traditional pruning process, some neurons that should be retained may be mistakenly deleted, resulting in a decrease in network accuracy. Through knowledge distillation, especially the introduction of the attention alignment mechanism, the student model can restore the accuracy loss that may be caused by pruned neurons by learning from the attention distribution of the teacher model, thereby ensuring the overall performance and robustness of the network. In general, the combination of knowledge distillation and attention alignment weighting mechanism enables the pruned model to maintain efficient defense while maximally recovering the accuracy loss caused by pruning. This not only improves the performance of the model on normal data sets, but also enhances the model's defense capabilities when facing backdoor attacks, ensuring the efficiency and accuracy of the defense process. The optimization objectives are defined as:
Citation Information
Cited By
Convolutional neural network back door defense method and device based on protection channel constraint
CN121582689A