Back door defense method for filtering poisoning data based on combined back door trigger
By calculating the absolute difference value of the data set, filtering the poisoned data and adding benign triggers, the problem of incomplete separation of poisoned data and clean data loss in the prior art is solved, and efficient backdoor defense effect is achieved.
Patent Information
- Application Number
- CN202411379328.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-07-08
AI Technical Summary
The existing backdoor defense methods cannot effectively separate poisoned data and clean data, and it is easy to sacrifice too much clean data during the separation process, resulting in a decrease in model accuracy and the inability to cope with existing common backdoor attacks.
By calculating the absolute difference in the dataset, some poisoning and clean data are selected, benign triggers are added and labels are modified, and the poisoning data is isolated in the model by combining triggers to ensure the integrity of the clean data.
It realizes effective separation of poisoned data without sacrificing too much clean data, maintaining the high accuracy of the model, and resisting multiple backdoor attacks.
Smart Images

Figure CN120277658A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence security, and in particular to a backdoor defense method based on filtering poisoned data by combining backdoor triggers. Background Art
[0002] Nowadays, deep learning has made remarkable progress in fields such as image classification, autonomous driving, and object detection. For example, in the field of image classification, its performance has even exceeded the ability of human discrimination. Therefore, deep learning has gradually been applied in the real world in these fields. As a result, the security issue of the model has also emerged. Among them, the backdoor attack is one of the most common attack methods on the model. The attacker completes the production of poisoned data by selecting some pictures, adding triggers, and modifying the labels, and then injects them into the training data set. When the victim normally trains the model with the poisoned data set, the attacker can complete the backdoor attack on the model. The backdoor attack has the characteristics of high attack success rate and imperceptibility. The attacked model can show a similar effect to the clean model on clean pictures, but when encountering data with the trigger made by the attacker, it will mislead the model prediction to the target label selected by the attacker.
[0003] With the continuous progress of the research on backdoor attacks, attackers can produce more concealed poisoned data, making it more difficult to detect whether the data is poisoned. The original attacks were easier to detect because they modified the labels, resulting in inconsistent semantic information and label information in the pictures. To achieve a more concealed attack, the attacker proposed the clean label attack, that is, only adding triggers without modifying the labels, making the picture information and label information consistent. At the same time, due to more refined attacks, the attacker can achieve a higher attack success rate with a smaller injection amount of poisoned data.
[0004] Due to the high attack success rate and high concealment of the backdoor attack, it may lead to major security accidents in reality. The current defenses are mainly divided into the suppression of the poisoning function in the model and the filtering of poisoned data. For poisoning suppression, the current defenses often reduce the accuracy of the model or cannot defend against the backdoor well when resisting the backdoor attack. For the filtering of poisoned data, the existing backdoor defense methods cannot completely separate the poisoned data and the clean data, and often require a part of the previously obtained clean data set, and it is not easy to obtain some clean data in the real environment. In addition, most of the existing poisoned data filtering methods often cannot cope with the existing common backdoor attacks, resulting in only part of the poisoned data being selected, so that a large amount of poisoned data still remains in the remaining data for training, leading to the failure of the defense, and sacrificing too much clean data when separating the poisoned data, resulting in a decrease in the model accuracy due to the lack of some clean data when training the model.
[0005] Due to the deficiencies of existing backdoor defense methods for poisoned data filtering, there is an urgent need for a backdoor defense method for poisoned data filtering that can defend against existing common backdoor attack methods. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a backdoor defense method for filtering poisoned data based on a combined backdoor trigger, to resist backdoor attacks against poisoned training data, and utilize the characteristics of the backdoor trigger during training. By using a benign trigger and the trigger designed by the attacker in the poisoned dataset through a combined trigger method to separate the poisoned data, so as to solve the problems existing in the existing poisoned data filtering methods.
[0007] The technical solution adopted by the present invention is as follows: A backdoor defense method for filtering poisoned data based on a combined backdoor trigger, comprising the following steps:
[0008] Step 1: Back up the entire original dataset; on the backup dataset, perform preliminary screening of partial poisoned data on the dataset, and screen out a small number of poisoned samples from the dataset; after training the proxy model on the entire dataset for a small number of iterations, calculate the absolute value of the difference between the two largest predicted values when each image passes through the filtering model output, to obtain the absolute difference; on the backup dataset, screen out a part of the data with the largest absolute difference as the partial poisoned dataset, and screen out a part of the data with the smallest absolute difference as the partial clean dataset;
[0009] Step 2: On the backup dataset, add benign triggers to the partial poisoned dataset and partial clean dataset obtained in the above steps and modify the labels, then merge them with the remaining data to train the model, and obtain the filtering model after a certain training cycle;
[0010] Step 3: Add benign triggers to all the data in the backup dataset, and then the filtering model makes predictions on it; separate the backup dataset into a poisoned dataset and a clean dataset through the predictions;
[0011] Step 4: Find the corresponding data in the original dataset through the clean dataset separated on the backup dataset, and train the model with this part of the original clean dataset to obtain the clean model.
[0012] In step 1 of the present invention, the calculation of the absolute difference is as follows:
[0013] ΔTop2 diff =|max(f θ (x)) - 2 nd max(f θ (x))|,
[0014] where f θf(x) is the model θ For all the output values corresponding to the labels of the input x, max(·), 2 nd max(·) represents taking the maximum value and the second - largest value respectively, ΔTop2 diff represents the absolute difference.
[0015] In step 2 of the present invention, the benign trigger is a 3×3 black - and - white pixel square in the upper - left corner in the CIFAR10 dataset, and a 14×14 black - and - white pixel square on the ImageNet - 12 dataset.
[0016] In step 2 of the present invention, the modified label of some poisoned datasets is 10.
[0017] In step 2 of the present invention, the modified label of some clean datasets is 11.
[0018] In step 2 of the present invention, after adding the benign trigger to some poisoned datasets and modifying the labels to the new target class labels, it will induce a filtering model after training. When the benign trigger and the attacker's trigger appear simultaneously, the induced filtering model predicts the new target class.
[0019] In step 2 of the present invention, after adding the benign trigger to some clean datasets and modifying the labels to the new benign target class labels, it will induce a filtering model after training. When the benign trigger appears alone, the induced filtering model predicts the benign target class.
[0020] In step 2 of the present invention, if the prediction result is the new target class, it is considered poisoned data; if the prediction is the benign target class, it is considered clean data.
[0021] In step 2 of the present invention, the training period is 30 epochs.
[0022] In step 1 of the present invention, 3% of the data with the largest absolute difference is selected as some poisoned datasets, and 1% of the data with the smallest absolute difference is selected as some clean datasets.
[0023] The present invention has the following advantages compared with the prior art: The defender of the present invention sorts and separates data by calculating the gap between the top values output by each data model, obtaining a small amount of clean data and poisoned data; through the idea of combining triggers, that is, by adding a benign trigger to the clean data and modifying the label to the benign target class, making the benign trigger the most prominent feature of the benign target class. In order to achieve that when the most prominent feature (benign trigger) of the benign target class and the most prominent feature (attacker's trigger) in the attacker's target class appear simultaneously, it is predicted as a new target class. A benign trigger is also added to the poisoned data and the label is modified to the new target class. In this way, after all data is added with benign triggers, the clean data, due to the added benign trigger, is predicted as the benign target class, while the poisoned data, because it contains the attacker's trigger, when the two triggers appear simultaneously, the model will predict it as the new target class. Thus, the data in the dataset predicted as the new target class by the model is filtered into the poisoned dataset, and the data predicted as the benign target class is filtered into the clean dataset.
[0024] It can be seen that the present invention uses the backdoor set by itself to filter poisoned data, so as to achieve that when facing the existing mainstream backdoor attacks, almost all poisoned data can be accurately isolated, while reducing the sacrifice number of clean samples, and maintaining a high model accuracy without the model being poisoned. In addition, the present invention can reduce the sacrifice of clean data while ensuring that the clean data contains as little poisoned data as possible. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is the overall flowchart of the experiment in the embodiment of the present invention;
[0026] Figure 2 It is the effect diagram of the image being attacked in the experiment in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0028] I. Embodiment
[0029] In the present invention, there are an attacker, a defender, and a training dataset CIFAR-10. The model uses Resnet18. The attacker can inject poisoned data into the training set without the defender's knowledge, and the defender needs to train a clean model in the dataset with poisoned samples.
[0030] The present invention proposes a backdoor defense method based on filtering poisoned data with a combined backdoor trigger. Taking a standard dataset as an example, no additional dataset is required, nor is it necessary to have the prerequisite of obtaining a small portion of clean data in the dataset. The poisoner has completed an unknown backdoor poisoning attack on it, while the defender needs to train a clean model on this possibly contaminated training set without prior knowledge, so that the model can achieve backdoor defense without significantly reducing the classification accuracy.
[0031] Specifically, it includes the following steps:
[0032] Step 1: Back up the entire original dataset. On the backed-up dataset, it is necessary to perform preliminary screening of some poisoned data in the dataset, conduct a preliminary separation of poisoned data and clean data, and screen out a small number of poisoned samples from the dataset.
[0033] When the filtering model predicts poisoned data and clean data, the gap between the two largest output values for each poisoned sample by the filtering model is often larger than that of clean samples. Based on this observation, the absolute difference between the two largest predicted values when each image passes through the filtering model can be calculated to distinguish clean samples and poisoned samples. After training the proxy model on the entire dataset for a small number of iterations, calculate the absolute difference by taking the absolute value of the difference between the two largest predicted values when each image passes through the filtering model. The specific calculation of the absolute difference is as follows:
[0034] ΔTop2 diff =|max(f θ (x)) - 2 nd max(f θ (x))|,
[0035] where f θ (x) is the output value corresponding to all labels of the input x by the filtering model f θ , and max(·), 2 nd max(·) respectively represent taking the maximum value and the second largest value. Among them, the larger the absolute difference ΔTop2 diff value, the more likely it is a poisoned sample.
[0036] On the backed-up dataset, sort the data in the backed-up dataset from largest to smallest according to the obtained absolute difference, and select a small portion of data with the largest absolute difference (such as 3%) as the partially poisoned dataset, and select a portion of data with the smallest absolute difference (such as 1%) as the partially clean dataset.
[0037] In this embodiment, on the backup dataset, first, a poisoned model for screening is trained using the poisoned CIFAR-10 dataset for 5 epochs. Then, for each input data, the absolute value of the difference between the two values with the largest predicted values in the model output is calculated.
[0038] Step 2: On the backup dataset, benign triggers are added to the partially poisoned dataset and the partially clean dataset separated in the above step. The benign trigger added to the partially poisoned dataset is a 3*3 black-and-white pixel square in the upper left corner in the CIFAR10 dataset and a 14*14 black-and-white pixel square in the ImageNet-12 dataset. After adding the benign trigger to the partially poisoned dataset, the label is modified to the new target class label. Here, to avoid overlapping with the attacker's target class, an additional new class is selected. For example, in CIFAR10, the labels are 0-9, and we set 10 as the new target class label. The addition in the partially clean dataset is the same benign trigger as in the partially poisoned dataset. After adding the benign trigger to the partially clean dataset, the label is modified to the benign target class label, where this benign target class label is also an extra-created class and is not the same as the line target class. For example, in CIFAR10, the labels are 0-9, and we set 11 as the benign target class label.
[0039] The partially poisoned dataset and the partially clean dataset after adding the benign trigger and modifying the label are merged with the remaining data to train the model, and then a filtering model for filtering poisoned data is obtained after a small number of training epochs. In this embodiment, the number of training epochs is 30.
[0040] After adding the benign trigger to the partially poisoned dataset and modifying the label to the new target class label, the filtering model will be induced after training. When the benign trigger and the attacker's trigger appear simultaneously, the filtering model is induced to predict the new target class; after adding the benign trigger to the partially clean dataset and modifying the label to the new benign target class label, the filtering model will be induced after training. When the benign trigger appears alone, the filtering model is induced to predict the benign target class. At the same time, keep the filtered poisoned data containing as little clean data as possible.
[0041] Step 3: According to the filtering model obtained in Step 2, benign triggers are added to all the data in the backup dataset, and then the filtering model makes predictions on it; if the prediction result is the new target class 10, it is considered poisoned data; if the prediction is the benign target class 11, it is considered clean data. The backup dataset is separated into a poisoned dataset and a clean dataset through prediction.
[0042] Step 4: Since the first three steps are carried out on the backup data, where a small part of the poisoned data and clean data have their labels modified and since benign triggers are added to all the data, the obtained clean set also has benign triggers. Directly using it to train a clean model may affect the model accuracy and implant a benign trigger backdoor. Therefore, the corresponding data in the original dataset can be found through the clean dataset separated from the backup dataset. This part of the data does not carry our benign trigger and its label has not been modified. Thus, the original clean dataset is obtained, and the clean model is trained using this part of the original clean dataset.
[0043] II. Dataset and Model Structure.
[0044] This scheme uses two publicly available standard datasets, CIFAR-10 and a subset of 12 classes of ImageNet, ImageNet-12, as the datasets for testing the effectiveness of this defense method. The two datasets differ in size and quantity. The detailed information is shown in Table 1.
[0045] Table 1
[0046] Dataset Number of classes Pixel value Training quantity Testing quantity Model CIFAR-10 10 32*32*3 40000 50000 ResNet-18 ImageNet-12 12 224*224*3 4320 1440 ResNet-18 。
[0047] III. The following is the experimental process of this embodiment.
[0048] 1. Experimental Setup.
[0049] Experimental Setup: In this embodiment, five advanced backdoor attacks are used to test the effectiveness of the defense of this experiment. In order to first successfully implement the backdoor attack in this embodiment, the SGD optimizer is used, the learning rate is set to 0.01, the momentum is set to 0.9, the weight decay is 1e-4, the epoch is set to 100, and the learning rate is reduced by 10 times when the epoch reaches 50 and 75. In this paper, the poisoning rate for the dirty label attack is set to 0.1, and the poisoning rate α for the clean label attack is set to 0.07. For the target label, label 0 is selected as the target label in both CIFAR-10 and ImageNet-12. In order to ensure that the attack effect can be achieved for all types of attacks, the poisoning rate for the dirty label backdoor attacks such as BadNet, Blend, and Bpp is set to 10%, and the poisoning rate for the clean label attacks SIG and Refool is set to 7%. The specific parameter settings of the attack methods are shown in Table 2.
[0050] Table 2
[0051] Attack method Type Pattern Target Poisoning ratio BadNets Fixed Square 0 10% Blend Fixed Blended image 0 10% Bpp Change Pixel jitter 0 10% WaNet Change Image distortion 0 10% Refool Fixed Reflection 0 7% SIG Change Sin perturbation 0 7% 。
[0052] 2. Experimental Results.
[0053] Experimental Results of Defense Effectiveness:
[0054] See Figure 1 and Figure 2 In this experiment, six common backdoor attacks were defended on two datasets, CIFAR-10 and ImageNet-12. The six backdoor attacks included dirty label attacks and clean label attacks. At the same time, four defense methods were compared horizontally. The two metrics for evaluating backdoor attacks and defenses were the attack success rate (ASR) and the clean data prediction accuracy (ACC). The specific details are shown in Table 3.
[0055] To verify the effectiveness of the method for separating poisoned data and clean data, the true positive rate (TPR) was used to represent the proportion of poisoned samples in the separated poisoned data among the total poisoned samples, and the false positive rate (FPR) was used to represent the proportion of clean samples in the poisoned data among the total clean samples for statistics. In addition, the results of two poisoning filtering methods were compared horizontally. The specific results are shown in Table 4:
[0056] Table 3
[0057]
[0058]
[0059] Table 4
[0060]
[0061]
[0062] It can be seen from the experiment that the present invention can resist various backdoor poisoning attacks in the neural network model. Before formally training the neural network model, this method can filter out backdoor poisoning attacks and ensure the quantity of clean data for training, which is very meaningful for improving the security in the neural network model.
[0063] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A backdoor defense method based on filtering poisoned data with a combined backdoor trigger, comprising the following steps: Step 1: Back up the entire original data set; On the backup data set, perform preliminary screening of partial poisoned data on the data set to screen out a small number of poisoned samples from the data set; After training the proxy model on the entire data set for a small number of iterations, calculate the absolute value of the difference between the two largest predicted values when each image passes through the filtering model output to obtain the absolute difference; On the backup data set, screen out a part of the data with the largest absolute difference as the partial poisoned data set, and screen out a part of the data with the smallest absolute difference as the partial clean data set; Step 2: On the backup data set, add benign triggers to the partial poisoned data set and partial clean data set obtained in the above steps and modify the labels, then merge them with the remaining data to train the model, and after a certain number of training cycles, obtain the filtering model; Step 3: Add benign triggers to all the data in the backup data set, and then the filtering model makes predictions on it; Separate the backup data set into a poisoned data set and a clean data set through prediction; Step 4: Find the corresponding data in the original data set through the clean data set separated on the backup data set, and train the model with this part of the original clean data set to obtain the clean model.
2. The backdoor defense method for filtering poisoned data based on a combined backdoor trigger according to claim 1, wherein In Step 1, the calculation of the absolute difference is as follows: ΔTop2 diff = |max(f θ (x)) - 2 nd max(f θ (x))|, where f θ (x) is the output value corresponding to all labels of the input x of the model f θ , max(·), 2 nd max(·) represents taking the maximum value and the second largest value respectively, and ΔTop2 diff represents the absolute difference.
3. The backdoor defense method for filtering poisoned data based on a combined backdoor trigger according to claim 1, characterized in that In Step 2, the benign trigger is a 3*3 black and white pixel square in the upper left corner in the CIFAR10 data set, and a 14*14 black and white pixel square on the ImageNet-12 data set.
4. The backdoor defense method for filtering poisoned data based on a combined backdoor trigger according to claim 1, wherein In Step 2, the modified label of the partial poisoned data set is 10.
5. The backdoor defense method for filtering poisoned data based on a combined backdoor trigger according to claim 1, characterized in that, In Step 2, the modified label of the partial clean data set is 11.
6. The backdoor defense method for filtering poisoned data based on a combined backdoor trigger according to claim 1, wherein In Step 2, after adding the benign trigger to the partial poisoned data set and modifying the label to the new target class label, it will induce the filtering model after training. When the benign trigger and the attacker's trigger appear at the same time, induce the filtering model to predict the new target class.
7. The backdoor defense method for filtering poisoned data based on a combined backdoor trigger according to claim 1, wherein In Step 2, after adding the benign trigger to the partial clean data set and modifying the label to the benign target class label, it will induce the filtering model after training. When the benign trigger appears alone, induce the filtering model to predict the benign target class.
8. The backdoor defense method for filtering poisoned data based on a combined backdoor trigger according to claim 1, wherein In Step 3, if the prediction result is the new target class, it is considered poisoned data; if the prediction is the benign target class, it is considered clean data.
9. The backdoor defense method for filtering poisoned data based on a combined backdoor trigger according to claim 1, wherein In Step 2, the training cycle is 30 epochs.
10. The backdoor defense method for filtering poisoned data based on a combined backdoor trigger according to claim 1, wherein In Step 1, screen out 3% of the data with the largest absolute difference as the partial poisoned data set, and screen out 1% of the data with the smallest absolute difference as the partial clean data set.