A Backdoor Defense Method Based on Model-Level Contrastive Learning

Through the model-level comparison learning two-stage defense method MCLDef, the problem of original data accuracy loss caused by backdoor defense in the existing technology is solved, and efficient backdoor erasing and attack success rate reduction is achieved.

CN115329329BActive Publication Date: 2025-07-11EAST CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210753091.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-29
Publication Date
2025-07-11
Estimated Expiration
2042-06-29

AI Technical Summary

Technical Problem

When erasing backdoor attacks, existing backdoor defense methods often lead to loss of original data accuracy and fail to effectively reduce the attack success rate.

Method used

The two-stage defense method MCLDef based on model-level comparison learning is adopted to generate poisoned data through trigger inversion, and positive and negative pairs are defined in the feature space. The model-level comparison learning loss is used to optimize and purify the model to achieve backdoor erasing.

Benefits of technology

It effectively reduces the success rate of backdoor attacks while minimizing the loss of accuracy of original data as much as possible, and has better performance than existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115329329B_ABST
    Figure CN115329329B_ABST
Patent Text Reader

Abstract

The present invention discloses a backdoor defense method based on model-level contrastive learning. The backdoor defense method includes two stages, called MCLDef. In the first stage, MCLDef inverses the trigger to obtain the inverted trigger. In this way, the obtained trigger can be used to generate poisoned data. Based on the positive and negative pairs in the backdoor defense proposed in the present invention, in the second stage, MCLDef performs model-level contrastive learning on the positive and negative pairs, where the feature representation of the poisoned data is pulled closer to its clean data counterpart, while shrinking or even destroying the cluster corresponding to the feature representation of the original poisoned data. In this way, the backdoor of the attacked model can be erased, while greatly reducing the loss of accuracy for the original data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and relates to a backdoor defense method based on model-level contrast learning, specifically a defense method against backdoor injection attacks. The method performs backdoor defense through two-stage steps to achieve the erasure of the backdoor in the attacked model, that is, to reduce the attack success rate while minimizing the loss of the accuracy of the original data as much as possible. Background Art

[0002] In recent years, the prosperity of artificial intelligence (AI) has promoted the increasing deployment of deep neural networks (DNNs) in various security-critical fields (e.g., autonomous driving, medical diagnosis, industrial control). Due to the wide application of deep neural networks, any small error can lead to irreparable consequences. However, DNNs have been proven to be vulnerable to potential threats at different stages of their life cycle. A novel attack method called backdoor attack mainly puts a small part of data with triggers into the training set for training. Utilizing the powerful learning ability of deep neural networks, a strong connection is established between the specific trigger and the target label, thereby embedding the backdoor into the DNN. This enables the DNN to predict it as the target label if the trigger appears in the test data during the model inference stage, thus completing the attack. What is terrifying about the backdoor attack is its concealment. It can not only achieve the specified attack prediction for poisoned samples, but also basically maintain the prediction accuracy for the original samples. Therefore, it is difficult for users to discover whether the model has been attacked by a backdoor, and it is also somewhat difficult to defend against the attacked model.

[0003] Currently, the danger of backdoor attacks has attracted extensive attention from many researchers, and they have proposed a series of backdoor defense methods. There are mainly two mainstream backdoor defense methods. The first is the detection-based method. The goal of the defense is often to detect whether the model has been implanted with a backdoor and find the target label among them, or to detect whether the data sample is poisoned data, so as to filter out the poisoned data when training the model. However, the focus of this type of method is on detection rather than erasing the backdoor. Although it can help identify potential threats, the backdoor model still needs to be repaired. The second defense method, that is, the method based on backdoor erasure, aims to erase the backdoor by purifying the negative impact of the attacked model. However, some existing methods in this category focus more on how to erase the backdoor and do not focus on how to protect the original accuracy of the original data in the purified model. Therefore, their purification process may cause a large change in the feature space of the model, thus causing a certain damage to the accuracy of the original samples. How to make full use of the distribution characteristics of the feature vectors of the original data and poisoned data in the feature space (that is, poisoned data often forms a new cluster in the special space) to correct the distribution of the feature vectors of poisoned samples is a challenge in backdoor defense methods. Summary of the Invention

[0004] To solve the deficiencies of the prior art, the object of the present invention is to provide a backdoor defense method based on model-level contrast learning.

[0005] The object of the present invention is to use the method of model-level contrast learning to achieve backdoor erasure while minimizing the loss of the original accuracy of the model. So far, most existing erasure-based backdoor defense methods have not started from the perspective of the feature space. The present invention is inspired by an observation that poisoned data often forms a new cluster in the feature space of the attacked network. The present invention will correct the feature vectors of poisoned samples from the perspective of the feature space.

[0006] To achieve this goal, considering the idea of the contrast learning method, which is to use unlabeled data to better learn good feature representations. Traditional contrast learning is based on instances, taking different transformations of an instance as positive pairs and different transformations of different instances as negative pairs, making the feature representations of positive pairs close and the feature representations of negative pairs far away, so as to learn very good feature representations. This idea can just be used in backdoor defense to correct the feature vectors of poisoned samples. Therefore, the present invention defines positive and negative pairs in backdoor defense according to the distribution characteristics of the feature vectors of the original samples and poisoned samples in the feature space. And on this basis, the present invention proposes a two-stage backdoor defense method MCLDef. As Figure 2As shown, the architecture has two main parts: 1. Trigger inversion (i.e., the first stage) to obtain the inverted trigger; 2. Model-level contrastive backdoor erasure (i.e., the second stage) to obtain the purified model and achieve the defense goal. In the second stage, the concept of positive and negative pairs, which is core in contrastive learning, needs to be defined. Therefore, the present invention defines positive and negative pairs in backdoor defense.

[0007] The specific technical solution for achieving the object of the present invention is:

[0008] Positive and negative pairs in backdoor defense defined based on the feature space

[0009] Figure 1 The feature space of the backdoored neural network obtained by the classic backdoor attack BadNets is shown. There are 10 classes of original data, represented by 0 - 9 respectively; the poisoned data is an additional class. It can be seen in the present invention that the original data forms respective clusters according to their classes, that is, the backdoored model can well predict the original data. For the poisoned data, their feature representations in the feature space do not lie in the clusters corresponding to their true labels but form a new cluster. The goal of backdoor defense is to shrink or even destroy the newly formed cluster of the poisoned data while making the feature representations of the poisoned data return to the corresponding original clusters. Considering this goal, since the prediction behavior of the poisoned model for the original data is normal, therefore, the present invention defines the feature representation of a clean sample in the poisoned model and the feature representation of its corresponding poisoned sample in the purified model as a positive pair. Through the definition of the positive pair, the present invention hopes that the distance between the positive pairs gets closer and closer, that is, the feature representation of the poisoned data will approach the cluster where its corresponding clean data is located. In addition, the present invention hopes to shrink or even destroy the cluster where the poisoned data is in the poisoned model. Therefore, the present invention defines the feature representations of the poisoned sample in the poisoned model and the purified model as negative pairs. Through the definition of the negative pair, the present invention hopes that the distance between the negative pairs gets farther and farther, that is, the feature representation of the poisoned data moves away from the cluster it forms in the poisoned model.

[0010] Formally, the present invention assumes that \(x\) is an original data, \(\hat{x}\) is its corresponding poisoned data, \(h(\cdot)\) is the output of the neural network feature extractor (the neural network removes the last layer classifier), \(\omega\) and \(\hat{\omega}\) correspond to the parameters of the purified model and the poisoned model respectively. The present invention uses the feature generator of the network to generate the following feature representations, which are respectively: the feature representation of the poisoned data (poisoned sample) in the purified model the feature representation of the poisoned data (poisoned sample) in the poisoned model the feature representation of the original data (clean sample) in the poisoned model The present invention constructs \((z, \hat{z})\) as a positive pair, \((z, z)\) cle \((\hat{z}, \hat{z})\) as negative pairs.poi )Constructed as a negative pair.

[0011] Trigger inversion based on the general formula for poisoned data (first stage)

[0012] Figure 2 The upper part shows the workflow of the trigger inversion stage, the goal of which is to obtain an optimized trigger to construct poisoned data.

[0013] The method of the present invention needs to obtain poisoned data. However, it is not practical to assume that the trigger of the attacked backdoor model is known. Therefore, the present invention needs to perform trigger inversion to obtain the trigger. The present invention uses an optimized method for trigger inversion. For the original data x, the present invention can obtain poisoned data using the following general formula:

[0014]

[0015] where Δ is the trigger pattern, m is a binary mask used to determine the position of the trigger, and ⊙ denotes the dot product of tensors. The combination of Δ and m constitutes the trigger for the attack.

[0016] For a given part of the clean training set data where |·| represents the number of the set, N represents the number of set elements, that is, N pieces of clean data. The present invention intends to use the clean training set data and the trigger obtained by trigger inversion to obtain the poisoned data set through formula (1) If the present invention uses the existing backdoor attack detection method, the present invention can obtain the target label y of the backdoor attack t , that is where 1 ≤ i ≤ N. The present invention can optimize the trigger pattern Δ and the mask m through the following formula:

[0017]

[0018] where represents the cross-entropy loss function; f(·) is the output of the neural network; represents the parameters of the poisoned model, λ represents the hyperparameter that controls the weight of the regularization term optimization objective; |·| represents the L1 norm of a vector.

[0019] Based on formula (2), the present invention can obtain the optimized trigger pattern Δ * and the mask m * , so as to generate corresponding poisoned data using clean data.

[0020] Backdoor Erasure Based on Model-Level Contrastive Learning Method (Second Stage)

[0021] Figure 2 The lower part shows the workflow of this stage. The goal of this stage is to use the positive and negative pairs defined in the present invention and model-level contrastive learning to achieve the erasure of the backdoor of the poisoned model, while minimizing the loss of the accuracy of the original data as much as possible.

[0022] For a given part of clean training set data where The present invention uses the optimized trigger pattern Δ obtained in the first stage * and the mask m * , and obtains the poisoned dataset according to Equation (1) For one of the clean data x i and its corresponding poisoned data The present invention obtains the positive pair according to the positive and negative pairs defined above negative pair Next, the present invention uses the model-level contrastive learning loss to erase the backdoor in the attacked model. The model-level contrastive learning loss is as follows:

[0023]

[0024] where τ is a temperature coefficient used to control the smoothness of the soft label (i.e., the feature vector obtained after the sample is input into the model); is the cosine similarity function used to measure the similarity of two feature vectors.

[0025] At the beginning of this stage, the present invention sets the parameters ω of the purification model equal to the parameters of the poisoned model i.e., The present invention optimizes ω using the following formula:

[0026]

[0027] Based on Equation (4), the present invention can obtain the optimized purification model ω * , which can achieve backdoor erasure and minimize the loss of the model's accuracy for the original data, thus achieving the defense goal against backdoor attacks.

[0028] The present invention also proposes an application of the above backdoor defense method in defending against backdoor attacks.

[0029] The present invention also proposes a backdoor defense system for implementing the above backdoor defense method. The backdoor defense system includes a trigger inversion module and a model-level contrastive learning module;

[0030] The trigger inversion module uses the poisoned model attacked by the backdoor and a part of clean data to optimize and obtain the inverted trigger;

[0031] The model-level contrastive learning module is used to, through the obtained positive and negative pairs, make the feature representation of the poisoned data approach the cluster where the corresponding clean data is located, while moving away from the cluster formed in the poisoned model, thereby shrinking or even destroying the newly generated cluster of the poisoned data, obtaining a purified model, and achieving the defense goal.

[0032] The beneficial effects of the present invention include: defending against backdoor injection attacks, being able to erase the backdoor in the attacked model, that is, reducing the attack success rate, and at the same time minimizing the loss of the accuracy of the original data as much as possible.

[0033] The defense performance of the present invention against the classic backdoor attack method BadNets is that the attack success rate is reduced from 99.64% to 2.22%, and the loss of the accuracy of the original data is 0.82%. Compared with the existing outstanding defense methods (Fine-Tuning, Neural Attention Distillation, Implicit Backdoor Adversarial Unlearning, NeuralCleanse), the defense performance of the present invention for reducing the attack success rate is increased by an average of 51.59%, and the performance of the loss of the accuracy of the original data is reduced by an average of 1.71%. Brief Description of the Drawings

[0034] Figure 1 It is a visualization diagram of the hidden space of the deep neural network obtained by the backdoor attack, that is, the motivation diagram of the present invention.

[0035] Figure 2 It is the workflow diagram of the backdoor defense method based on model-level contrastive learning of the present invention. Detailed Embodiments

[0036] Combined with the following specific embodiments and drawings, the present invention will be further described in detail. The processes, conditions, experimental methods, etc. for implementing the present invention, except for the specifically mentioned content below, are all common knowledge and well-known common sense in the art, and the present invention has no special limiting content.

[0037] The present invention discloses a backdoor attack defense method based on model-level contrastive learning. With the prosperous development of artificial intelligence technology, the present invention has witnessed a series of designed backdoor injection attacks, which maliciously threaten deep neural networks by attackers. Although there are various existing defense methods that can effectively erase the backdoors of the attacked networks, they still suffer from a high attack success rate or a large and unavoidable loss of the original data accuracy. The present invention is inspired by an observation that poisoned data often forms a new cluster in the feature space of the attacked network. The present invention proposes an effective two-stage defense method based on model-level contrastive learning, named MCLDef. In the first stage, MCLDef obtains the inverted trigger based on trigger inversion. In this way, the obtained trigger can be used to generate poisoned data. Based on the positive and negative pairs in the backdoor defense proposed by the present invention, in the second stage, MCLDef performs model-level contrastive learning on the positive and negative pairs, where the present invention pulls the feature representation of the poisoned data closer to its clean data counterpart, while shrinking or even destroying the cluster corresponding to the original poisoned data feature representation. In this way, the backdoor of the attacked model can be erased, while greatly reducing the loss of the original data accuracy. Experimental results show that the method of the present invention can not only greatly reduce the backdoor attack success rate by using 5% of clean data, but also minimize the loss of the original data accuracy as much as possible.

[0038] The focus of the present invention is to use some clean data to erase the backdoors in the backdoor-attacked network. So far, many attackers have implemented backdoor injection attacks, that is, by using the powerful learning ability of neural networks, a certain proportion of poisoned data with triggers is injected into the training set data. The backdoor model obtained through backdoor attack training can predict the pictures with triggers as the specified labels, but maintain good prediction behavior for the original data. There are many existing backdoor defense methods at present, but these methods often only consider how to erase the backdoors of the attacked models, without starting from the perspective of the feature space to protect the original data accuracy of the purified models.

[0039] The present invention defines the positive and negative pairs in backdoor defense from the phenomenon in the feature space, and under the guidance of the positive and negative pairs, proposes a two-stage backdoor defense method MCLDef.

[0040] The present invention provides an effective two-stage defense method MCLDef based on model-level contrastive learning for backdoor injection attacks. In the first stage, an inverted trigger is obtained based on trigger inversion. In this way, the obtained trigger can be used to generate poisoned data. Based on the positive and negative pairs in the backdoor defense proposed in the present invention, in the second stage, MCLDef performs model-level contrastive learning on the positive and negative pairs, in which the feature representation of the poisoned data is brought closer to its clean data counterpart, while shrinking or even destroying the cluster corresponding to the original feature representation of the poisoned data. In this way, the backdoor of the attacked model can be erased, while greatly reducing the loss of accuracy of the original data.

[0041] The backdoor attack defense method based on model-level contrastive learning proposed in this invention:

[0042] Step 1: Trigger inversion based on the general formula for poisoned data (first stage)

[0043] The process of backdoor attack is to add triggers to a part of clean data to construct poisoned data. The trigger is often a change in certain pixels in the photo, such as a specific block, logo image, subtle pixel changes, etc. The implementation of the backdoor attack enables the deep neural network to learn the connection between the trigger and the target label, so that when the test image contains a trigger, the deep neural network can predict the image as the target label, thereby completing the attack. The present invention requires the use of poisoned data, but the trigger used by the attacker is unknown to the defender. Therefore, this step implements trigger inversion, that is, using the poisoned model attacked by the backdoor and a part of the clean data to optimize and generate the trigger used by the attacker.

[0044] This stage sees Figure 2 In the first half, the purpose is to obtain the trigger for the attack through an optimized method, so that the trigger can be used to generate poisoned data.

[0045] The trigger for trigger inversion of the present invention is the change of certain pixels in the photo, which can be represented by a trigger pattern Δ and a binary mask m, where Δ is a pixel image of the same size as the image, such as a square in the lower right corner of the image, a picture of the "Hello Kitty" logo; m is a 0, 1 tensor of the same size as the image, which is used to determine the location of the trigger. Therefore, the present invention uses m⊙Δ to construct the final trigger, which is a tensor of the same size as the image, so that the trigger can be added to the clean image to construct the poisoned data.

[0046] For the original data x, the present invention can use the following general formula to obtain the poisoned data:

[0047]

[0048] Among them, Δ is the trigger pattern, m is a binary mask used to determine the position of the trigger, and ⊙ represents the dot product of tensors. The combination of Δ and m constitutes the trigger for the attack.

[0049] For a given part of clean training set data where (|·| represents the number of elements in the set, N represents the number of elements in the set, that is, N pieces of clean data). The present invention intends to use the clean training set data and the trigger obtained by trigger inversion to obtain the poisoned data set through Equation (1) If the present invention uses existing backdoor attack detection methods, the present invention can obtain the target label y of the backdoor attack t , that is where 1 ≤ i ≤ N. For example, the present invention traverses all possible labels, uses Equation (2) for optimization to obtain the corresponding trigger pattern Δ and mask m for each label, and then obtains the L1 norm of m corresponding to each possible label. The present invention can use the median absolute deviation to find the smaller abnormal m, and the corresponding label can be regarded as the target label y t . The present invention takes minimizing Equation (2) as the optimization goal. In this stage, the trigger pattern Δ and mask m will be iteratively optimized, and the optimization algorithm is used to minimize Equation (2), that is, to make the deep neural network predict the poisoned data constructed by the trigger pattern Δ and mask m as the target label, and at the same time make the trigger size as small as possible.

[0050]

[0051] where represents the cross-entropy loss function; f(·) is the output of the neural network; represents the parameters of the poisoned model, λ represents the hyperparameter that controls the weight of the regularization term optimization goal; |·| represents the L1 norm of a vector.

[0052] The poisoned model is a deep neural network, and its function is to predict the category of data. However, when facing data containing triggers, its prediction behavior for poisoned data is specific, so as to complete the attacker's attack goal. The deep neural network used in the present invention is a convolutional neural network, which mainly includes a convolutional layer and a fully connected layer. Therefore, its parameters mainly include the convolutional kernel parameters of the convolutional layer, the weight parameters and offset parameters of the fully connected layer.

[0053] Through the optimization of Equation (2), the present invention can obtain the optimized trigger pattern Δ * and mask m * , so that the corresponding poisoned data can be generated using clean data.

[0054] Step 2: Backdoor Defense Positive and Negative Pairs Defined Based on the Feature Space

[0055] Based on Figure 1 the distribution characteristics of the feature vectors obtained by the original samples and poisoned samples in the feature space after dimensionality reduction mapping through the network output, that is, the poisoned data often forms a new cluster in the feature space. Since the poisoned model's prediction behavior for the original data is normal, and the goal of the present invention is to obtain a purified model, that is, a model that is consistent with the structure of the poisoned model, has normal prediction behavior for the original data, and also has normal prediction behavior for the poisoned data containing the trigger. The present invention defines the feature representation of a clean sample in the poisoned model and the corresponding feature representation of the poisoned sample obtained by using Equation (1) in the purified model as positive pairs. Through the definition of positive pairs, the present invention hopes that the distance between positive pairs will become closer and closer, that is, the feature representation (feature vector) of the poisoned data will approach the cluster where the feature representation of its corresponding clean data is located. In addition, the present invention hopes to shrink or even destroy the cluster where the poisoned data is located in the poisoned model. Therefore, the present invention defines the feature representations of the poisoned sample in the poisoned model and the purified model as negative pairs. Through the definition of negative pairs, the present invention hopes that the distance between negative pairs will become farther and farther, that is, the feature representation of the poisoned data will move away from the cluster it forms in the poisoned model, realizing the shrinkage or even destruction of the cluster where the poisoned data is located in the poisoned model.

[0056] Specifically, the present invention assumes that x is an original data, is its corresponding poisoned data, h(·) is the output of the neural network feature extractor (the neural network removes the last layer classifier), ω and correspond to the parameters of the purified model and the poisoned model respectively. The parameters are the convolution kernel parameters of the convolutional layer, the weight parameters of the fully connected layer, and the offset parameters in the model. The present invention uses the feature generator of the network to generate the following feature representations, where z, z poi , z cle respectively represent the feature vector obtained after the poisoned sample is input into the purified model, the feature vector obtained after the poisoned sample is input into the poisoned model, and the feature vector obtained after the clean sample is input into the poisoned model. The present invention constructs (z, z cle ) into positive pairs and (z, z poi ) into negative pairs.

[0057] Step 3: Backdoor Erasure Based on the Model-Level Contrastive Learning Method (Second Stage)

[0058] This stage is shown in Figure 2 the lower part. The purpose is to use the positive and negative pairs obtained in the previous step to move the feature representation of the poisoned data closer to the cluster where its corresponding clean data is located and away from the cluster it forms in the poisoned model, thereby shrinking or even destroying the newly generated cluster of the poisoned data and achieving the defense goal of the present invention.

[0059] For a given part of clean training set data wherein The present invention utilizes an optimized trigger pattern Δ * and a mask m * , and obtains a poisoned data set according to equation (1) For one of the clean data x i and its corresponding poisoned data The present invention obtains positive pairs according to the positive and negative pairs defined in step 2 negative pairs Then, the invention uses a model-level contrastive learning loss to erase the backdoor in the attacked model. The model-level contrastive loss is as follows:

[0060]

[0061] where τ is a temperature coefficient used to control the smoothness of the soft labels (i.e., the feature vectors obtained after the samples are input into the model); is a cosine similarity function used to measure the similarity between two feature vectors, where ||·|| represents the L2 norm of a vector. The cosine similarity function is essentially the dot product operation of two vectors after the L2 norm.

[0062] At the initial stage of this phase, the present invention sets the parameters ω of the purification model equal to the parameters of the poisoned model i.e., The present invention takes minimizing equation (3) as the optimization objective, so as to make the feature representation (feature vector) of the poisoned data approach the cluster where the feature representation of its corresponding clean data is located, and at the same time move the feature representation of the poisoned data away from the cluster formed in the poisoned model. The present invention will perform iterative optimization through an optimization algorithm to achieve the following optimization objective:

[0063]

[0064] This equation will obtain the model parameters that minimize equation (3) through an optimization method, that is, the final purification model parameters. Through the optimization of equation (4), the present invention can obtain an optimized purification model ω * , thereby achieving the defense objective against backdoor attacks.

[0065] The present invention assumes that there is currently an attacked backdoor model, which is trained by an unreliable third-party data set. In addition, the present invention assumes that there is a small part (5%) of clean training set data available for fine-tuning the model to erase the backdoor of the attacked model.

[0066] In the first stage, the inventive method of the present invention needs to obtain poisoning data. However, it is not practical to assume that the trigger of the attacked backdoor model is known. Therefore, the present invention needs to perform trigger inversion to obtain the trigger. The present invention uses an optimized method for trigger inversion. Based on Step 1, the present invention implements this stage, and the first-stage algorithm is as follows:

[0067]

[0068] The first line initializes the trigger pattern Δ and the mask m. Lines 2-7 iteratively perform optimization until a trigger pattern and a mask with relatively good quality are obtained. Line 2 indicates that the loop iteratively optimizes E times. The present invention uses stochastic gradient descent for optimization, that is, Lines 3-7. Among them, Line 3 indicates randomly taking a mini-batch of data b{x, y} from the data set where x represents the data and y represents the label; Line 4 uses the trigger pattern Δ and the mask m to obtain the poisoned data Line 5 calculates the loss value, where l(·) represents the cross-entropy loss function; Lines 6 and 7 use the stochastic gradient descent algorithm to optimize the trigger pattern Δ and the mask m.

[0069] In the second stage, the inventive method of the present invention needs to utilize the positive and negative pairs defined in Step 2 and guide the execution of model-level contrast backdoor erasure based on the model-level loss defined in Step 3. Based on Steps 2 and 3, the present invention implements this stage, and the second-stage algorithm is as follows:

[0070]

[0071] The first line indicates that at the beginning of Stage 2, the purified model ω is equal to the poisoned model The present invention needs to implement the purification process by optimizing and updating the purified model. Lines 2-9 iteratively perform purification until a purified model that achieves the defense goal is obtained. Line 2 indicates that the loop iteratively optimizes E times. The present invention uses stochastic gradient descent for optimization, that is, Lines 3-9. Among them, Line 3 indicates randomly taking a mini-batch of data b{x, y} from the data set where x represents the data and y represents the label; Line 4 uses the trigger pattern Δ and the mask m to obtain the poisoned data Lines 5-7 obtain the corresponding feature vectors of the positive and negative pairs, that is, z, z poi 、z cle , where h(·) represents the feature vector output by the model; Line 8 calculates the contrast loss value, τ is a temperature coefficient, and sim(·, ·) is a cosine similarity function; Line 9 uses the stochastic gradient descent algorithm to update the purified model ω.

[0072] Embodiment

[0073] The method of the present invention focuses on the defense scenarios against various backdoor injection attacks. To carry out a specific embodiment, the present invention defends against the most classic backdoor attack method, BadNets, and the process is as follows:

[0074] 1. Select CIFAR-10 as the experimental dataset of the present invention. The training set contains 50,000 images, and the test set contains 10,000 images. Both the training set and the test set contain 10 categories, such as categories like dogs, cats, birds, etc.

[0075] 2. To simulate the scenario of defending against backdoor injection attacks, the present invention will attack 10% of the training dataset (that is, randomly select 500 images for attack in each category, a total of 5,000 images), that is, add a trigger with a size of 3*3 at the lower right corner of the image. The present invention uses this dataset to train for 50 rounds to obtain the poisoned model under attack.

[0076] 3. The present invention assumes that 5% of the clean training set data is obtained for defense, a total of 2,500 images.

[0077] 4. The present invention uses the clean data and the first stage (i.e., the trigger inversion stage) to optimize the trigger pattern and mask. The present invention optimizes for 100 rounds to obtain the finally optimized trigger pattern and mask for generating poisoned data in the second stage.

[0078] 5. The present invention uses the clean data in 4 and the optimized trigger pattern and mask to purify the poisoned model based on the second stage (model-level contrast backdoor erasure) to obtain the purified model. The present invention trains the fine-tuning network for 20 rounds to obtain the finally purified model.

[0079] 6. The present invention uses all the test set data (10,000 images) to test the accuracy of the original data of the purified model; uses all the non-attack target class test set data (9,000 images), adds the attack trigger to obtain poisoned data, and tests the attack success rate of the purified model. Table 1 shows the defense performance of the present invention, where the "Before" row represents the performance of the model without any defense.

[0080] 7. As a comparison of the experimental effects, use relatively mainstream backdoor defense algorithms as benchmarks, namely Fine-Tuning (FT), Neural Attention Distillation (NAD), Implicit Backdoor Adversarial Unlearning (I-BAU), Neural Cleanse (NC).

[0081] Table 1: Defense Performance

[0082] Attack success rate (%) Accuracy of original data (%) Before 99.64 83.01 FT 12.78 80.64 NAD 5.76 81.71 I-BAU 2.82 79.54 NC 3.58 81.37 MCLDef 2.22 82.19

[0083] According to the experimental results in Table 1, it can be seen that compared with the current mainstream backdoor defense methods (FT, NAD, I-BAU, NC), the method proposed in the present invention has a lower attack success rate, a higher original data accuracy, and much better performance.

[0084] Specifically: (1) In terms of attack success rate, the present invention reduces the attack success rate to a minimum and eliminates the backdoor in the attacked model as much as possible. Compared with the mainstream defense method, the backdoor defense method in the present invention reduces the attack success rate by 82.63%, 64.46%, 21.28%, and 37.99%, respectively; (2) In terms of the original data accuracy method, the present invention reduces the loss of original data accuracy to a minimum. Compared with the mainstream defense method, the present invention improves the original data accuracy by 1.92%, 0.59%, 3.33%, and 1.01%, respectively.

[0085] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the present invention, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the attached claims.

Claims

1. A backdoor defense method based on model-level contrastive learning, characterized in that, Defense is carried out by using the distribution characteristics of the feature representations of the original data and the poisoned data, including the following steps: Step 1, inversion of the trigger based on the poisoning model: The method of optimizing the trigger is adopted for trigger inversion: for the original data x, the poisoned data is obtained by using the following formula: where Δ is the trigger pattern, m is a binary mask used to determine the position of the trigger, and ⊙ refers to the dot product of tensors; the combination of Δ and m constitutes the trigger for the attack; For a given part of the clean training set data where |·| represents the number of the set, N represents the number of set elements, that is, N pieces of clean data; using the clean training set data and the trigger obtained by inverting the trigger, the poisoned data set is obtained through Equation (1) Use the backdoor attack detection method to obtain the target label y of the backdoor attack t , that is where 1 ≤ i ≤ N; The trigger pattern Δ and the mask m are optimized by the following formula: Among them, represents the cross-entropy loss function; f(·) is the output of the neural network; represents the parameters of the poisoned model; λ represents the hyperparameter that controls the weight of the regularization objective; |·| represents the L1 norm of a vector; Obtain the optimized trigger pattern Δ based on Equation (2) * and the mask m * , so as to generate corresponding poisoned data using clean data; Step 2, backdoor erasure based on the model-level contrast learning method: For a given part of the clean training set data wherein using the optimized trigger pattern Δ * and the mask m * , and the poisoned dataset For one of the clean data x i and its corresponding poisoned data Obtain positive pairs according to positive and negative pairs negative pairs The positive pair is the feature representation of a clean sample in the poisoned model and the feature representation of its corresponding poisoned sample in the purification model; the negative pair is the feature representation of the poisoned sample in the poisoned model and the purification model; The model-level contrast learning loss is used to erase the backdoor in the attacked model.

2. The backdoor defense method based on model-level contrastive learning according to claim 1, wherein The model-level contrast learning loss in Step 2 is as follows: where τ is a temperature coefficient used to control the smoothness of the soft label; sim(·,·) is the cosine similarity function used to measure the similarity of two feature vectors; the soft label refers to the feature vector obtained after the sample is input into the model.

3. The backdoor defense method based on model-level contrastive learning according to claim 2, characterized in that, At the beginning, let the parameters ω of the purification model be equal to the parameters of the poisoned model That is Optimize ω using the following formula: Based on Equation (4), the optimized purification model ω is obtained * , thus achieving the defense goal against backdoor attacks.

4. The backdoor defense method based on model-level contrastive learning according to claim 1, wherein The poisoning model is a deep neural network used to predict the class of data; when facing data containing a trigger, its prediction behavior for the poisoned data is specified, so as to achieve the attacker's attack goal; the deep neural network is a convolutional neural network, including a convolutional layer and a fully connected layer, and its parameters include the convolutional kernel parameters of the convolutional layer, the weight parameters of the fully connected layer, and the offset parameters.

5. The backdoor defense method based on model-level contrastive learning according to claim 1, wherein The feature generator of the network is used to generate the following feature representations: Among them, z, z poi , z cle respectively represent the feature vector obtained after the poisoned sample is input into the purification model, the feature vector obtained after the poisoned sample is input into the poisoning model, and the feature vector obtained after the clean sample is input into the poisoning model.

6. The backdoor defense method based on model-level contrastive learning according to claim 5, wherein, The feature vector of the clean sample in the poisoning model and the feature vector of its corresponding poisoned sample in the purification model are defined as the positive pair, and the feature vectors of the poisoned sample in the poisoning model and the purification model respectively are defined as the negative pair.

7. Application of the backdoor defense method according to any one of claims 1-6 in the defense against backdoor attacks.

8. A backdoor defense system for implementing the backdoor defense method according to any one of claims 1-6, characterized in that, The backdoor defense system includes a trigger inversion module and a model-level contrast learning module; The trigger inversion module uses the poisoned model attacked by the backdoor and a part of clean data to optimize and obtain the inverted trigger; The model-level contrast learning module is used to, through the obtained positive and negative pairs, realize moving the feature representation of the poisoned data closer to the cluster where its corresponding clean data is located, and at the same time moving away from the cluster formed in the poisoning model, so as to shrink or even destroy the newly generated cluster of the poisoned data, obtain the purification model, and achieve the defense goal.

Citation Information

Patent Citations

  • Method for a generative adversarial network with different hierarchical function combinations

    CN113988293A

  • Backdoor attack defense method and defense system based on security training

    CN114238975A