Trigger generative model backdoor attack method based on low-frequency feature distribution
By implanting triggers with low-frequency feature distribution in the frequency domain and optimizing adversarial samples using particle swarm algorithm, the problem of insufficient covertness of triggers in the existing backdoor attack methods is solved, and higher covertness, attack effectiveness and robustness are achieved.
Patent Information
- Application Number
- CN202510273079.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-06
AI Technical Summary
The triggers in the existing backdoor attack methods are not concealed enough and are easily detected and cleaned, resulting in poor attack effectiveness.
A trigger generation model based on low-frequency feature distribution is adopted, triggers are implanted in the frequency domain, and the low-frequency part is optimized through the particle swarm algorithm to generate more concealed adversarial samples, and used to train the poisoning model.
It improves the concealment, attack effectiveness and robustness of backdoor attacks, and can maintain a high backdoor attack success rate and classification accuracy under low poisoning rates.
Smart Images

Figure CN120105412A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and deep learning security, and in particular to a trigger generation model backdoor attack method based on low-frequency feature distribution. Background Art
[0002] The widespread application of deep neural network (DNN) technology has made scenes such as autonomous driving, smart homes, and smart cities in movies gradually become a reality, creating a fully automated environment for society. However, the security and privacy issues brought about by the potential black box nature of DNN have also attracted widespread attention and become a new research hotspot. Among them, there are endless attack methods against DNN, and backdoor attacks are one of the new attack paradigms. After training on a poisoned dataset, a backdoor that can be activated by a trigger is implanted in the deep learning model, so that the model maintains high-precision work for normal inputs, and for inputs with triggers, the output is according to the target category specified by the attacker. Under this attack paradigm, deep learning models show great vulnerability.
[0003] The triggers of traditional backdoor attacks are usually visible, such as implanting specific patterns in images. Although such visible triggers can effectively activate backdoors, they are static and visible, and are easily reviewed and cleaned by the human eye, thus increasing the possibility of being defended. In response to the problem of visible backdoor triggers, researchers have proposed invisible trigger modes, such as mixing triggers with clean samples through hybrid strategies to reduce the transparency of triggers, using perturbations that are not easily perceived by humans, and using steganography and regularization to implant backdoors in images. These trigger generation modes all make minor modifications to the images, which improves the concealment of the backdoor.
[0004] In addition to the above-mentioned backdoor attack methods in the spatial domain, there are also some backdoor attack methods in the frequency domain. In order to solve the problem that the trigger model destroys the semantics of the poisonous pixels, thus causing the failure of attacking the dense prediction model, researchers proposed backdoor attack methods based on the frequency domain, such as injecting invisible frequencies to attack the FIBA paradigm, triggering perturbations in the frequency domain corresponding to small pixel perturbations scattered throughout the image to ensure that the image is not distorted, creating smooth backdoor triggers without high-frequency artifacts, and the detector tuned on the smooth trigger can be generalized to invisible weak smooth triggers and other backdoor attack methods.
[0005] The effectiveness of backdoor attacks is closely related to their triggering patterns, so it is crucial to find more covert triggering patterns. Currently, the triggering patterns of most attacks are designed based on heuristic methods, such as universal perturbations, or in a non-optimal way. However, these methods have some defects in terms of the effectiveness and robustness of the attacks. Summary of the invention
[0006] In view of the problem of insufficient concealment of the above triggers, the present invention provides a trigger generation model backdoor attack method based on low-frequency feature distribution. By implanting triggers in the frequency domain, the method can obtain poisoned samples with stronger concealment, attack effectiveness and robustness, and realize the control of the poisoned model. The technical solution of the present invention is as follows:
[0007] A trigger generation model backdoor attack method based on low-frequency feature distribution, comprising the following steps:
[0008] S1, trigger generation stage: the backdoor attacker uses the frequency domain trigger generation model to poison the clean samples, thereby obtaining adversarial samples with backdoor features;
[0009] S2, backdoor embedding stage: Use the poisoned dataset containing adversarial samples to train the model, and inject the backdoor into the convolutional layer of the victim model by mapping the adversarial samples and target labels;
[0010] S3, inference and prediction stage: put the test set into the victim model obtained in S2 for inference, and use the backdoor attack performance judgment index to measure the attack effect;
[0011] Furthermore, the method of using the frequency trigger generation model to poison the clean sample in step S1 specifically includes:
[0012] First, according to the poisoning rate α, select label L from the data set D which is not the target label L target The sample of vector is obtained. Carrier , dataset D except D Carrier The rest is D residue , then the attacker chooses to Carrier To poison, D, D Carrier , D residue The relationship between the three is as follows.
[0013] D=D Carrier ∪D residue
[0014]
[0015] Secondly, we use the clean model F that has been trained on the dataset D to find the vector sample x residue Then, through the threshold and sliding window method, we find the 3*128*128 area that has the greatest impact on classification under the feature map importance analysis. right Perform a two-dimensional discrete Fourier transform to obtain the transformed complex matrix. Then perform centralization to gather the low frequencies to the central position. This center point is consistent with the diagonal center point of the generated 3*3 binary mask matrix. Then, use the particle swarm algorithm to modify the appropriate low frequencies of the binary mask matrix to modify the low frequency part, inverse Fourier transform, and restore the image. After multiple rounds of iterations, the local optimal solution is obtained. D Carrier After the above processing, the poisoned data set Carrier is obtained. p The final training set is shown below.
[0016] D train =Carrier p ∪D residue
[0017] Furthermore, the method of training the model using the poisoned data set containing adversarial samples in step S2 specifically includes:
[0018] The training set D train Used in victim model F θ Perform standardized training and obtain the poisoning model F after training p Because the backdoor attacker's goal is to make Model F θ The correct label is predicted for the clean sample, and the label predefined by the attacker is predicted for the adversarial sample with trigger, so we need to minimize the loss function to ensure F p Performance on clean data, where represents the cross entropy loss, x represents the clean sample, T φ For the trigger generation model, the loss function is expressed as follows.
[0019]
[0020] In order to implement the backdoor attack, we need to minimize the loss function of the adversarial sample. The loss function is expressed as follows.
[0021]
[0022] Finally, the above is integrated into a constrained optimization problem and expressed as follows.
[0023]
[0024] Furthermore, the step S3 uses the backdoor attack performance judgment index to measure the attack effect, which specifically includes:
[0025] The following indicators are usually used to judge the quality of backdoor attack methods: accuracy (PA) and attack effectiveness (ASR). The characteristics of each indicator are introduced as follows.
[0026] Accuracy: Its purpose is to verify whether the prediction accuracy of the backdoor infected model is affected by the backdoor when testing clean data.
[0027] Attack effectiveness: The purpose of attack effectiveness is to quantify the likelihood that an attack sample containing a specific trigger will activate the target label.
[0028] Furthermore, the poisoning rate α and the carrier sample set D Carrier The selected strategies include:
[0029] The choice of poisoning rate has a huge impact on the effect of backdoor attacks. Among them, the attack success rate is higher: a higher poisoning rate means that more training samples are injected into the backdoor trigger, and the model is more likely to learn the backdoor pattern, thus showing a higher success rate in the backdoor attack. The attack success rate is lower: a lower poisoning rate means that there are fewer backdoor samples, and the model may not be able to fully learn the backdoor pattern, resulting in a decrease in the attack success rate. Therefore, when determining the poisoning rate, you can first use the data set of other backdoor research method experiments, and then fine-tune the poisoning rate according to the backdoor evaluation indicators to find a suitable poisoning rate.
[0030] Usually the carrier sample set D Carrier The choice of needs to be determined by the target label of the backdoor attack. After the target label is determined, the original samples with different target labels are randomly selected from the training data set, and these original samples are combined into the carrier sample set D Carrier .
[0031] Furthermore, the clean model F that has been trained on the dataset D is used to find the carrier sample x residue Methods for finding areas that have a significant impact on classification include:
[0032] The Grad-CAM method is used. By combining the feature maps of the last convolutional layer, Grad-CAM can generate a high-resolution heat map indicating the areas in the image that are most important for classification decisions.
[0033] Furthermore, the method using a threshold and a sliding window specifically includes:
[0034] Grad-CAM is used to find the area that has an important impact on classification, and then the threshold and sliding window method are used to find the 3*128*128 area that has the greatest impact on classification under the feature map importance analysis. Here, the threshold can be determined based on the statistical characteristics of the feature map score, that is, the threshold is set by weighting the mean and standard deviation. The selection of the threshold is conducive to finding the area that has a great impact on the classification result. The threshold calculation formula is as follows:
[0035] Threshold=η+c*σ
[0036] Where η is the mean of the feature map importance scores and σ is the standard deviation of the feature map importance scores.
[0037] c is an adjustable parameter that controls the offset of the threshold relative to the mean.
[0038] Furthermore, the two-dimensional discrete Fourier transform and inverse transform method specifically includes:
[0039] Assume that the grayscale image can be regarded as a H*W matrix, where H and W represent the height and width of the image respectively. The image is a signal f(p,q), where f(p,q) represents the pixel value of the spatial image at the coordinate point f(p,q). The image is converted from the spatial domain to the frequency domain using the discrete Fourier transform (DFT), and the inverse transform is represented by (IDFT). The mathematical expression is as follows:
[0040]
[0041] The coordinates of F(μ,v) are the frequency domain, which is the amplitude and phase of the frequency component at the (μ,v) coordinates.
[0042] The advantages and beneficial effects of the present invention are as follows:
[0043] Robustness: This invention is a backdoor attack method based on frequency domain injection. By using the idea of generative adversarial networks and the difference between the carrier image and the adversarial sample as a penalty term, the poisoning model is optimized while optimizing the trigger. The adversarial samples generated by this attack method can improve robustness while retaining the semantics of image pixels.
[0044] Concealment: The present invention can inject triggers without reducing perception capabilities, which can improve concealment.
[0045] Attack effectiveness: The present invention can maintain a high success rate of backdoor attacks while maintaining a poisoning rate lower than 0.01 without reducing the classification accuracy of clean samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is the main flow chart of the backdoor attack method of the trigger generation model based on low-frequency feature distribution of the present invention.
[0047] Figure 2 It is a diagram of the trigger generation algorithm based on the distribution of low-frequency features.
[0048] Figure 3 It is a flowchart for finding the optimal low-frequency offset using particle swarm algorithm for frequency domain trigger backdoor attack. DETAILED DESCRIPTION
[0050] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0051] The present invention provides a trigger generation model backdoor attack method based on low-frequency feature distribution, Figure 1 A flowchart of the backdoor attack method of the trigger generation model of low-frequency feature distribution is given, which includes the following steps:
[0052] S1, the backdoor attacker uses the frequency domain trigger generation model to poison the clean sample and obtain the adversarial sample with the backdoor feature introduced;
[0053] S2, trains the model with a poisoned dataset containing adversarial examples, and injects the backdoor into the convolutional layer of the victim model by mapping the adversarial examples and the target labels;
[0054] S3, the victim model obtained in S2 classifies the clean sample into the correct label, and the prediction result for the adversarial sample containing backdoor features becomes the target label;
[0055] For step S1, the specific implementation steps are as follows: First, the data set selected is CIFAR-10, which contains 60,000 (32*32) color images of ten categories, with 6,000 images in each category. The data set is divided into 50,000 training images and 10,000 test images. Secondly, determine that the target label is bird, and then select samples with labels other than bird from the CIFAR-10 training data set according to the poisoning rate of 0.01 to obtain the carrier sample set D Carrier The training data set excludes D Carrier The rest is D residue , then the attacker chooses to Carrier To poison.
[0056] Use other models F that have been trained on the CIFAR-10 training dataset other To find the vector sample x residue The region that has an important impact on classification is selected, and the threshold is determined by weighting the mean and standard deviation of the feature map score. The adjustable parameter is set to 0.2. Then, the sliding window method is used to find the 3*128*128 region that has the greatest impact on classification under the feature map importance analysis. right Perform a two-dimensional discrete Fourier transform to obtain the transformed complex matrix. Then perform a centralization process to gather the low frequencies to the central position. This center point is consistent with the diagonal center point of the generated 3*3 binary mask matrix. Use the particle swarm algorithm to modify the low frequencies determined by the binary mask matrix, and then modify and inverse Fourier transform the low frequencies to obtain the restored image, which is F other The probability of being classified to the target label is used as an indicator to judge the quality of this adversarial sample. After multiple rounds of iterations, a local optimal solution is obtained.
[0057] D Carrier After the above processing, the poisoned data set Carrier is obtained. p , and the final training set is represented as D train =Carrier p ∪D residue .
[0058] Regarding step S2, the specific embodiment steps are as follows: Use the poisoned training data set D containing adversarial samples train The model is trained, and then the backdoor is injected into the convolutional layer of the victim model by mapping the adversarial sample and the target label. After training, the poisoned model is obtained, which already contains the backdoor.
[0059] For step S3, the specific implementation steps are as follows: the CIFAR-10 test data set is used as the test data set of this experiment, and then the test data set is also used to generate adversarial samples according to the steps of S1, and the clean samples and adversarial samples of the test data set are put into the victim model obtained in S2 for classification tasks, and the accuracy and attack effectiveness are used as evaluation indicators to measure the performance of the backdoor attack. When studying the poisoning rate, the accuracy (PA) and attack effectiveness (ASR) results of the backdoor attack are shown in Table 1. Table 1 shows the impact of poisoning rate on attack performance.
Claims
1. A trigger generation model backdoor attack method based on low-frequency feature distribution, characterized in that: The following steps are involved: S1, trigger generation stage: the backdoor attacker uses the frequency domain trigger generation model to poison the clean samples, thereby obtaining adversarial samples with backdoor features; S2, backdoor embedding stage: Use the poisoned dataset containing adversarial samples to train the model, and inject the backdoor into the convolutional layer of the victim model by mapping the adversarial samples and target labels; S3, reasoning and prediction stage: put the test set into the victim model obtained in S2 for reasoning, and use the backdoor attack performance judgment index to measure the attack effect.
2. The trigger generation model backdoor attack method based on low-frequency feature distribution according to claim 1 is characterized in that: The method of using the frequency trigger generation model to poison the clean sample in step S1 specifically includes: First, according to the poisoning rate α, select label L from the dataset D which is not the target label L target The sample of vector is obtained. Carrier , dataset D except D Carrier The rest is D residue , then the attacker chooses to Carrier To poison, D, D Carrier , D residue The relationship between the three is as follows. D=D Carrier ∪D residue Secondly, we use the clean model F that has been trained on the dataset D to find the vector sample x residue Then, through the threshold and sliding window method, we find the 3*128*128 area that has the greatest impact on classification under the feature map importance analysis. right Perform a two-dimensional discrete Fourier transform to obtain the transformed complex matrix. Then perform centralization to gather the low frequencies to the central position. This center point is consistent with the diagonal center point of the generated 3*3 binary mask matrix. Then, use the particle swarm algorithm to modify the appropriate low frequencies of the binary mask matrix to modify the low frequency part, inverse Fourier transform, and restore the image. After multiple rounds of iterations, the local optimal solution is obtained. D Carrier After the above processing, the poisoned data set Carrier is obtained. p The final training set is shown below. D train =Carrier p ∪D residue。 3. The trigger generation model backdoor attack method based on low-frequency feature distribution according to claim 1 is characterized in that: The method for training a model using a poisoned data set containing adversarial samples in step S2 specifically includes: The training set D train Used in victim model F θ Perform standardized training and obtain the poisoning model F after training p Because the backdoor attacker's goal is to make Model F θ The correct label is predicted for the clean sample, and the label predefined by the attacker is predicted for the adversarial sample with trigger, so we need to minimize the loss function to ensure F p Performance on clean data, where represents the cross entropy loss, x represents the clean sample, T φ For the trigger generation model, the loss function is expressed as follows. In order to implement the backdoor attack, we need to minimize the loss function of the adversarial sample. The loss function is expressed as follows. Finally, the above is integrated into a constrained optimization problem and expressed as follows.
4. The trigger generation model backdoor attack method based on low-frequency feature distribution according to claim 3 is characterized in that: The attack effect measurement method using backdoor attack performance judgment indicators includes: The following indicators are usually used to judge the quality of backdoor attack methods: accuracy (PA) and attack effectiveness (ASR). The characteristics of each indicator are introduced as follows. Accuracy: Its purpose is to verify whether the prediction accuracy of the backdoor infected model is affected by the backdoor when testing clean data. Attack effectiveness: The purpose of attack effectiveness is to quantify the likelihood that an attack sample containing a specific trigger will activate the target label.
5. The trigger generation model backdoor attack method based on low-frequency feature distribution according to claim 2 is characterized in that: The poisoning rate α and the carrier sample set D Carrier The selected strategies include: The choice of poisoning rate has a huge impact on the effect of backdoor attacks. Among them, the attack success rate is higher: a higher poisoning rate means that more training samples are injected into the backdoor trigger, and the model is more likely to learn the backdoor pattern, thus showing a higher success rate in the backdoor attack. The attack success rate is lower: a lower poisoning rate means that there are fewer backdoor samples, and the model may not be able to fully learn the backdoor pattern, resulting in a decrease in the attack success rate. Therefore, when determining the poisoning rate, you can first use the data set of other backdoor research method experiments, and then fine-tune the poisoning rate according to the backdoor evaluation indicators to find a suitable poisoning rate. Usually the carrier sample set D Carrier The choice of needs to be determined by the target label of the backdoor attack. After the target label is determined, the original samples with different target labels are randomly selected from the training data set, and these original samples are combined into the carrier sample set D Carrier .
6. The trigger generation model backdoor attack method based on low-frequency feature distribution according to claim 2 is characterized in that: The clean model F that has been trained on the dataset D is used to find the vector sample x residue Methods for finding areas that have a significant impact on classification include: The Grad-CAM method is used. By combining the feature maps of the last convolutional layer, Grad-CAM can generate a high-resolution heat map indicating the areas in the image that are most important for classification decisions.
7. The trigger generation model backdoor attack method based on low-frequency feature distribution according to claim 2 is characterized in that: The method using a threshold and a sliding window specifically includes: Grad-CAM is used to find the area that has an important impact on classification, and then the threshold and sliding window method are used to find the 3*128*128 area that has the greatest impact on classification under the feature map importance analysis. Here, the threshold can be determined based on the statistical characteristics of the feature map score, that is, the mean or standard deviation to set the threshold. The selection of the threshold is conducive to finding the area that has a great impact on the classification result.
8. The trigger generation model backdoor attack method based on low-frequency feature distribution according to claim 2 is characterized in that: The two-dimensional discrete Fourier transform and inverse transform method specifically includes: Assume that the grayscale image can be regarded as a H*W matrix, where H and W represent the height and width of the image respectively. The image is a signal f(p, q), and f(p, q) represents the pixel value of the spatial image at the coordinate point f(p, q). The image is converted from the spatial domain to the frequency domain using the discrete Fourier transform (DFT), and the inverse transform is represented by (IDFT). The mathematical expression is as follows: The coordinates of F(μ, v) are the frequency domain, which is the amplitude and phase of the frequency component at the (μ, v) coordinates.