Hidden backdoor attack method based on dynamic mask and multilevel feature fusion
By using a technology of fusion of dynamic masks and multi-level features in deep neural networks, the high-frequency information of the trigger image is embedded in the multi-level wavelet transform high-frequency part of the clean image to generate visually invisible poisoned images, solving the problem of insufficient concealment and effectiveness of existing backdoor attack methods, and achieving efficient and concealed backdoor attack effects.
Patent Information
- Application Number
- CN202510496616.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing backdoor attack methods are insufficient in terms of concealment and effectiveness, they are easily identified and erased by defense measures, and have a low attack success rate.
The hidden backdoor attack method based on the fusion of dynamic masks and multi-level features is adopted. The high-frequency information of the flip-fiveness of the flip-five part of the clean image is embedded in the multi-level wavelet transform high-frequency part of the clean image, and a visually invisible poisoned image is generated, and weak and strong flip-fiveness strategies are used during the training and attack stages.
It significantly improves the concealment and effectiveness of backdoor attacks, and can bypass advanced defense methods such as Grad-CAM, Neural Cleanse, STRIP and Fine Pruning, without controlling the training process, and is suitable for more realistic scenarios.
Smart Images

Figure CN120032227A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep neural network security, and in particular to a covert backdoor attack method based on dynamic mask and multi-level feature fusion. Background Art
[0002] Deep neural networks (DNNs) have achieved remarkable results in image classification, speech recognition, natural language processing, and other fields, and are widely used in many scenarios in real life, such as smart homes, autonomous driving, and medical diagnosis. However, with the widespread application of deep neural networks, their security issues have become increasingly prominent. Backdoor attacks are a serious security threat. Attackers inject specific triggers into training data, causing the model to make incorrect predictions for inputs with triggers during the inference phase, while performing normally for clean samples.
[0003] Many studies have proposed a variety of backdoor attacks to enhance the effectiveness and concealment of backdoors. For example, attacks based on image mixing, attacks based on reflection, and attacks based on fixed image watermarks. Although these attack methods have a high success rate and are more concealed than the original methods, they can still be detected by the human eye and are easily identified and erased by defense measures. In order to enhance the concealment of the attack, the generation of invisible backdoor triggers has begun to be considered. For example, the backdoor attack based on image steganography considers writing a fixed text into the image in the form of deep learning steganography as a backdoor trigger; LIRA searches for invisible backdoor triggers in high-dimensional nonlinear parameter space; WaNet uses image distortion fields and uses distorted images as backdoor triggers. However, some of these hidden backdoor triggers require some strong assumptions, such as controlling the backdoor training process; some cannot bypass the most advanced defense methods, such as neural cleaning; some backdoor triggers are visually invisible, but their visual indicator parameters are still poor.
[0004] In addition, the above backdoor attack methods all require that the labels of the backdoor images be modified to the target labels. When subject to manual inspection, incorrectly labeled input label pairs (such as bird-cat) may arouse the auditor's suspicion, making the attack ineffective after being detected. Therefore, researchers proposed a clean label poisoning attack to ensure that the input label pairs are correct. Triggers are only inserted in the training images. This method without modifying the labels is more covert. Barn proposed setting a trigger so that the feature layer of the poisoned image after passing through the deep neural network is close to the target class. However, the attack success rate of this method is very low. Zhao et al. attacked the video classification task without modifying the labels. Ma et al. proposed using a GAN network to generate poisoned images, mixing the vector that generates the target class image with another random vector in a certain proportion, so that the generated image does not look like the target class, but can be classified into the target class. In addition, they also proposed using the method of generating adversarial samples (PGD) to confuse the classifier. It should be noted that the classifier that generates the adversarial samples is trained separately. These clean label backdoor attack methods have a common feature, which is to generate backdoor triggers from the perspective of the feature layer.
[0005] Therefore, there is a need for a covert backdoor attack method based on dynamic mask and multi-level feature fusion that can improve the concealment and effectiveness of backdoor attacks. Summary of the invention
[0006] The main purpose of the present invention is to provide a covert backdoor attack method based on dynamic mask and multi-level feature fusion to solve the problems of low concealment and effectiveness of backdoor attacks in the prior art.
[0007] To achieve the above object, the present invention provides a covert backdoor attack method based on dynamic mask and multi-level feature fusion, which specifically includes the following steps: S1, select the trigger diagram and scale the trigger diagram.
[0008] S2, performs three-level discrete wavelet transform on the clean image and wavelet transform on the trigger image to extract high-frequency information.
[0009] S3, embeds the high-frequency components of the trigger image into the high-frequency part of the multi-level wavelet transform of the clean image through linear combination, and generates the poisoned image using cubic inverse discrete wavelet transform.
[0010] S4, optimizes the poisoning strategy by using the backdoor random masking strategy and weak trigger training and strong trigger attack strategy.
[0011] Furthermore, step S1 specifically includes the following steps: S1.1, the size of the trigger image is 1 / 2 of the original clean image. After a discrete wavelet transform, the size of the high-frequency component of the first trigger image is 1 / 4 of the original clean image, and the size of the high-frequency component of the second-level wavelet transform of the original clean image matches.
[0012] S1.2, the size of the trigger image is 1 / 4 of the original clean image. After one discrete wavelet transform, the size of the high-frequency component of the second trigger image is 1 / 8 of the original clean image, matching the size of the high-frequency component of the third-level wavelet transform of the original clean image.
[0013] Furthermore, step S2 specifically includes the following steps: S2.1, for the clean image x c Perform three-level discrete wavelet transform to extract low-frequency component A step by step 1 , A 2 , A 3 and high frequency component H 1 , D 1 , V 1 ;H 2 , D 2 , V 2 ;H 3 , D 3 , V 3 : ; ; ; in, is the discrete wavelet transform.
[0014] S2.2, scale the trigger image into two pictures of different sizes, represented as: ; in, and is the scaling function; and The scaled image.
[0015] Then and Perform a wavelet transform:
[0016] ;
[0017] ;
[0018] in, and They are and The low-frequency component of and are respectively and high-frequency components of
[0019] Furthermore, step S3 specifically includes the following steps:
[0020] S3.1, fuse the high-frequency components of wavelet transforms at different scales:
[0021] ;
[0022] ;
[0023] Among them, and are embedding coefficients, is the second-level high-frequency component embedded with a backdoor, is the third-level high-frequency component embedded with a backdoor.
[0024] S3.2, generate the poisoned image P through three inverse discrete wavelet transforms:
[0025] ;
[0026] ;
[0027] ;
[0028] Among them, is the second-level low-frequency approximation component generated through the third-level wavelet inverse transform, is the first-level low-frequency approximation component generated through the second-level wavelet inverse transform; is the inverse discrete wavelet transform.
[0029] Furthermore, step S4 specifically includes the following steps:
[0030] S4.1, randomly pixel-mask the trigger image, and the masked area covers the sensitive areas in the original clean image where the pixel values are close to 0 or 255; in the training stage, expand the mask area, and the mask ratio is 70% - 90%, and in the inference stage, shrink the mask area, and the mask ratio is 5% - 20%.
[0031] S4.2, in the training stage, adopt the weak trigger mode, reduce the embedding coefficients α, β, and reduce the trigger strength to enhance the concealment; in the inference stage, switch to the strong trigger mode, increase the embedding coefficients α, β, and at the same time reduce the mask area to ensure the attack effectiveness of the trigger.
[0032] Furthermore, the value ranges of α and β are from 0.3 to 0.5, taking 0.3 in the training stage and 0.5 in the inference stage.
[0033] The present invention has the following beneficial effects: (1) To address the concealment problem of backdoor attacks, the present invention proposes a stealth backdoor attack method based on frequency domain transformation (WLAT). By using discrete wavelet transform, the high-frequency information of the trigger image is embedded into the high-frequency part of the multi-level wavelet transform of the clean image, generating a visually invisible poisoned image, which significantly improves the concealment of the trigger.
[0034] (2) To address the effectiveness of backdoor attacks, the present invention designs a strategy of weak trigger training and strong trigger attack. In the training phase, weaker high-frequency information embedding and a larger random mask coverage ratio are used, while in the attack phase, stronger high-frequency information embedding and a smaller random mask coverage ratio are used, which significantly improves the success rate of the attack.
[0035] (3) In response to the limitations of existing defense methods, this paper proposes a sample-specific trigger generation mechanism, which breaks through the assumption of fixed triggers in existing defense methods. It can effectively bypass advanced defense strategies such as Grad-CAM, Neural Cleanse, STRIP and Fine Pruning, and has stronger robustness.
[0036] (4) The present invention does not need to control the training process or access the loss function, and is applicable to more real-world scenarios, thus improving the practicality and wide applicability of backdoor attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the specific implementation of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the specific implementation or the prior art description. Obviously, the drawings described below are some implementations of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings: Figure 1 A flow chart of a concealed backdoor attack method based on dynamic mask and multi-level feature fusion of the present invention is shown.
[0038] Figure 2 A comparison chart of the PSNR entropy value results of the method proposed in the present invention and other methods is shown.
[0039] Figure 3 A comparison chart of the SSIM results of the method proposed in the present invention and other methods is shown.
[0040] Figure 4 A comparison chart of the GradCam results of the method proposed in the present invention and other methods is shown.
[0041] Figure 5A comparison chart of the STRIP entropy value results of the method proposed in the present invention and other methods is shown.
[0042] Figure 6 A comparison chart of pruning results between the method proposed in the present invention and other methods is shown. DETAILED DESCRIPTION
[0043] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0044] A covert backdoor attack method based on dynamic mask and multi-level feature fusion, specifically comprising the following steps: S1, trigger image selection and scaling: Select the trigger image and scale it. Randomly select an image as the trigger image and scale it to 1 / 2 and 1 / 4 of the original clean image size, ensuring that the high-frequency part after wavelet transform is different from the high-frequency part of the clean image. The semantic consistency of the trigger image is maintained by scaling, so that the scaled trigger image is 1 / 4 and 1 / 8 of the original image at the secondary and tertiary wavelet transforms, respectively.
[0045] S2, multi-level wavelet transform and high-frequency information extraction: perform three-level discrete wavelet transform on the clean image, perform wavelet transform on the trigger image, and extract high-frequency information. Perform three-level discrete wavelet transform on the clean image, extract low-frequency components step by step and decompose them to the third layer, and obtain the high-frequency components of the second layer ( ) and the third layer high frequency component ( ), and perform a first-level wavelet transform on two trigger images of different sizes to extract their high-frequency information ( )and( ), so as to be embedded into the second and third high-frequency components of the clean image to generate a covert poisoned image.
[0046] S3, high-frequency information fusion and poisoned image generation: The high-frequency components of the trigger image are embedded into the high-frequency part of the multi-level wavelet transform of the clean image through linear combination, and the poisoned image is generated using the cubic inverse discrete wavelet transform, which combines the low-frequency components of the clean image and the high-frequency components embedded in the backdoor to ensure the concealment and effectiveness of the poisoned image.
[0047] S4, optimizes the poisoning strategy by using the backdoor random masking strategy and weak trigger training and strong trigger attack strategy.
[0048] like Figure 1As shown, the present invention first performs discrete wavelet transform (W) on the clean image to obtain the low-frequency component A and the high-frequency components H, V, and D in the horizontal, vertical, and diagonal directions. The initial trigger image is scaled twice (R) and discrete wavelet transformed (W) to obtain two groups of components: one group is TA, TH, TD, and TV, and the other group is TA 2 , TH 2 , TD 2 , TV 2 .
[0049] Next, calculate the third-level high-frequency components, such as HT 2 =TH×α+H 3 ×(1-α) (shown in the red box), and calculate the secondary high-frequency components at the same time, such as HT=TH 2 ×β+H 2 ×(1-β) (shown in the green box). Then, through the inverse discrete wavelet transform (W - ) The processed components are synthesized into a preliminary poisoning image.
[0050] Finally, the initial poisoned image, the clean image and the masked trigger are added (+) to obtain the final poisoned image. The whole process achieves the hidden embedding of the backdoor in the image through operations such as discrete wavelet transform of the image, mixing of high-frequency components (using α and β parameters) and inverse transform.
[0051] Specifically, a trigger image is randomly selected from the image library, such as an image of a corgi's ear. The criterion for selecting a trigger image is that its high-frequency information is significantly different from that of the clean image. The trigger image should have rich high-frequency information so that it can effectively trigger the backdoor behavior after being embedded in the clean image.
[0052] Step S1 specifically includes the following steps: S1.1, the size of the trigger image is 1 / 2 of the original clean image. After a discrete wavelet transform, the size of the high-frequency component of the first trigger image is 1 / 4 of the original clean image, and the size of the high-frequency component of the second-level wavelet transform of the original clean image matches.
[0053] S1.2, the size of the trigger image is 1 / 4 of the original clean image. After one discrete wavelet transform, the size of the high-frequency component of the second trigger image is 1 / 8 of the original clean image, matching the size of the high-frequency component of the third-level wavelet transform of the original clean image.
[0054] Theoretically, as long as the high-frequency part of the selected trigger image after wavelet transformation is different from the high-frequency part of the clean image after wavelet transformation. Therefore, only one picture needs to be randomly selected to meet the required conditions. However, there are requirements for the size of the selected image. After a wavelet transformation, the image will be divided into four, and the length and width of the high-frequency classification will become 1 / 2 of the original. Therefore, if a high-frequency trigger is to be embedded in the secondary and tertiary wavelet transforms of the clean image after the three-level wavelet transformation, an image with a size of 1 / 2 of the original clean image and an image with a size of 1 / 4 of the original clean image are required, so that they become 1 / 4 and 1 / 8 of the size of the original clean image after wavelet transformation, which is exactly the same size as the secondary and tertiary wavelet transforms of the clean image. In order to enhance the effectiveness of the high-frequency component of the trigger image, the two selected trigger images should maintain the same semantics, otherwise it may reduce its availability. Therefore, the present invention finally chooses to scale a randomly selected image to keep its semantics unchanged.
[0055] Specifically, step S2 includes the following steps: S2.1, for the clean image x c Perform three-level discrete wavelet transform to extract low-frequency component A step by step 1 , A 2 , A 3 and high frequency component H 1 , D 1 , V 1 ;H 2 , D 2 , V 2 ;H 3 , D 3 , V 3 , obtain the high-frequency information of the deeper second and third layers to facilitate the embedding of more hidden backdoors: ; ; ; in, is the discrete wavelet transform.
[0056] S2.2, the trigger image ( ) to perform a wavelet transform. Note that the image scale will change after wavelet transform. In order to embed deeper and more hidden information, The backdoor is embedded in the high-frequency components of the second and third layers. We only need to extract the high-frequency information and scale the trigger image into two images of different sizes, expressed as: ; in, and is the scaling function; and The scaled image.
[0057] Then and Perform a wavelet transform: ; ; in, and They are and The low-frequency component of and They are and high frequency components.
[0058] Specifically, step S3 includes the following steps: S3.1, the high-frequency components of the backdoor trigger are embedded in a linear combination manner, the high-frequency components of wavelet transform at different scales are fused, and the high-frequency backdoor is embedded in the following expression: ; ; in, and is the embedding coefficient, and The value range is 0.3 to 0.5, with lower values in the training phase and higher values in the inference phase.
[0059] is the secondary high frequency component with the backdoor embedded. It is the third-level high-frequency component with the backdoor embedded.
[0060] WLAT dynamically adjusts the embedding coefficients through linear combination and , the trigger strength is differentially controlled in the training and inference stages, and a three-level discrete wavelet transform is used to hierarchically embed the trigger high-frequency information into the second-layer high-frequency component (H 2 / V 2 / D 2 ) and the third layer high frequency component (H 3 / V 3 / D 3 ) rather than embedding information in just a single high-frequency sub-band.
[0061] S3.2, generate the poisoned image P by three inverse discrete wavelet transforms: ; ; ; in, is the second-level low-frequency approximate component generated by the third-level inverse wavelet transform, is the first-level low-frequency approximate component generated by the second-level inverse wavelet transform; is the inverse discrete wavelet transform.
[0062] in, Utilizes a completely clean third-order low-frequency approximation component With the three-level high-frequency component embedded in the backdoor , The generated poisoned low frequency components are utilized and the secondary high frequency component with the backdoor embedded . The generated poisoned low frequency components are utilized With completely clean high frequency components .
[0063] The high-frequency noise of the present invention is extracted from another image, and the low-frequency components of the clean image are completely retained by three inverse discrete wavelet transforms (A 3 , A 22 , A 12 ), only high-frequency components are replaced to generate poisoned images.
[0064] Specifically, step S4 includes the following steps: S4.1, backdoor random masking strategy: random pixel masking is performed on the trigger image, and the masked area covers the sensitive areas with pixel values close to 0 or 255 in the original clean image; in the training stage, the masked area is expanded and the mask ratio is 70%-90%, and in the inference stage, the masked area is reduced and the mask ratio is 5%-20%.
[0065] S4.2, weak trigger training and strong trigger attack strategy: In the training phase, the weak trigger mode is adopted to embed the coefficients , Reduce, reduce the trigger strength to enhance concealment; in the inference phase, switch to strong trigger mode and embed the coefficient , Improve while reducing the mask area to ensure the attack effectiveness of the trigger.
[0066] Global trigger learning mechanism: The backdoor random masking strategy forces the model to learn global triggers distributed throughout the image rather than local regional features, avoiding the loss of concealment due to overfitting of local high-frequency information. Combined with multi-level wavelet transform, the high-frequency information of the trigger is embedded into the second and third high-frequency components of the clean image, generating a visually invisible but effective poisoned image.
[0067] Dataset selection: We conducted a large number of experiments on WLAT on three datasets, including Cifar10, Gtsrb, and ImageNet. Cifar10 contains 10 classes of natural images, including boats and horses, with a total of 50,000 training images and 10,000 test images. The image size is Gtsrb is a traffic sign dataset with 43 categories, nearly 40,000 training images and more than 10,000 test images. The image size is set to ImageNet is also a natural image dataset, including birds, dogs, etc. Due to its huge size, this paper only uses a subset of it. The subset used includes 100 categories of images, each with 60,000 training sets and 10,000 test sets. The image size is set to Three datasets of different scales are used to verify the wide applicability of WLAT.
[0068] Model selection: This paper experiments on three data sets on ResNet18 and RepVGG. ResNet18 is a classic classification model. The use of residual network modules greatly improves the model recognition accuracy. Many backdoor attack experiments use ResNet models for experiments. RepVGG is the latest deep convolutional neural network model VGG model. It combines the advantages of ResNet and VGG, applies the idea of residual network to VGG, and has higher accuracy for the same data set than the ordinary VGG network.
[0069] Evaluation indicators: This paper uses BadNets, Blend and WaNet to compare with the WLAT proposed in this paper. BadNets is the most classic defense method, Blend is a method that initially considers the concealment of triggers, and WaNet is the latest concealed trigger method. By comparing with the three methods, the superiority of WLAT is demonstrated in many aspects.
[0070] Defense resistance evaluation: The present invention uses Attack Success Rate (ASR) and Benign Accuracy (BA) to evaluate the effectiveness of different attacks. Specifically, ASR is defined as the ratio between successfully attacked poison samples and total poison samples. BA is defined as the accuracy of benign sample testing. In addition, the present invention uses peak signal-to-noise ratio PSNR and SSIM to evaluate concealment. Both indicators are used to measure the similarity between two images. Among them, PSNR is obtained by calculating the mean square error between two images. The mean square error is the average of the sum of the squares of the differences between two images. PSNR is the inverse logarithm of MSE multiplied by the maximum error, so it is usually expressed in decibels (db). SSIM refers to structural similarity, which compares the structural information of two images based on the idea of perceptual Hartmann coefficient to derive the similarity between them.
[0071] In order to further illustrate the advantages of the above method of the present invention, the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0072] Embodiment 1 Application of WLAT method on Cifar10 dataset On the Cifar10 dataset, the WLAT method selects a corgi ear image as a trigger and scales it to 1 / 2 and 1 / 4 of the original clean image size. The clean image is subjected to a three-level wavelet transform to extract its low-frequency and high-frequency components, and the scaled trigger image is subjected to a wavelet transform to extract its high-frequency information. The high-frequency information of the trigger image is embedded into the high-frequency components of the second and third level wavelet transforms of the clean image through linear combination, and the poisoned image is generated using the three-inverse discrete wavelet transform. In the training phase, weaker high-frequency information embedding (α=0.3, β=0.3) and a larger random mask coverage ratio are used; in the attack phase, stronger high-frequency information embedding (α=0.5, β=0.5) and a smaller random mask coverage ratio are used. Experimental results show that the attack success rate of WLAT on the Cifar10 dataset reaches 99.85%, and the benign accuracy decreases by less than 1%. The PSNR and SSIM values show that the poisoned image is almost visually indistinguishable from the original image.
[0073] Embodiment 2 Application of WLAT method on GTSRB dataset On the GTSRB dataset, the WLAT method also selects the corgi ear image as a trigger and performs wavelet transform and high-frequency information embedding in the same way. Weaker high-frequency information embedding and a larger random mask coverage ratio are used in the training phase, and stronger high-frequency information embedding and a smaller random mask coverage ratio are used in the attack phase. The experimental results show that the attack success rate of WLAT on the GTSRB dataset reaches 99.95%, and the benign accuracy decreases by less than 1%. The PSNR and SSIM values show that the poisoned images are almost visually indistinguishable from the original images.
[0074] Embodiment 3 Application of WLAT method on ImageNet dataset On a subset of the ImageNet dataset, the WLAT method selects a corgi ear image as a trigger and performs wavelet transform and high-frequency information embedding in the same way. Weaker high-frequency information embedding and a larger random mask coverage ratio are used in the training phase, while stronger high-frequency information embedding and a smaller random mask coverage ratio are used in the attack phase. Experimental results show that the attack success rate of WLAT on the ImageNet dataset reaches 99.90%, with a benign accuracy drop of less than 1%. The peak signal-to-noise ratio (PSNR) and the structural similarity index (SSIM) show that the poisoned image is almost visually indistinguishable from the original image.
[0075] It can be seen from the above that the new invisible backdoor attack method based on frequency domain transformation proposed in the present invention embeds the high-frequency information of the trigger image into the high-frequency part of the multi-level wavelet transform of the clean image through discrete wavelet transform. The generated poisoned image is almost indistinguishable from the original image visually, and has extremely high concealment; at the same time, the attack success rate of WLAT on Cifar10, GTSRB and ImageNet datasets is over 99%, and the impact on the accuracy of benign samples is minimal (the decrease is less than 1%), ensuring that the normal function of the model is not affected; in addition, WLAT can effectively bypass advanced defense strategies such as Grad-CAM, NeuralCleanse, STRIP and Fine Pruning, showing extremely strong robustness; compared with the existing technology, WLAT does not need to control the training process, is applicable to more real-life scenarios, and performs well on a variety of datasets and models, and has wide applicability.
[0076] PSNR is used to measure the degree of image distortion. The higher the value, the smaller the image distortion. Figure 2 The higher PSNR value is shown in the figure, indicating that the poisoned image after embedding the backdoor has a lower degree of distortion than the clean image. This means that the present invention has little impact on the image quality during the process of embedding the backdoor, and can effectively maintain the original visual effect of the image, providing strong support for the concealment of the backdoor.
[0077] SSIM is used to evaluate the structural similarity of images. The closer its value is to 1, the higher the structural similarity between images. Figure 3 The SSIM value of WLAT approaches 1, indicating that the poisoned image and the clean image are highly similar in structure. This further proves that the backdoor embedded in the present invention is extremely concealed, reducing the perceptible differences from the image structure level, making the poisoned image visually almost the same as the clean image, and it is difficult for the human eye or detection methods based on image structure differences to identify abnormalities.
[0078] like Figure 4 As shown in the figure, the heat map of the method (WLAT) provided by the present invention is highly similar to the heat map of the clean image, and the model decision focus area is consistent with the normal image (such as the target object itself), rather than fixed on the trigger. In contrast, the heat map of methods such as BadNets is abnormally concentrated in the visible trigger area. This shows that the backdoor trigger of WLAT is distributed in the high-frequency components of the global image, and no local detectable significant features are formed, thereby effectively circumventing the defense method based on saliency detection (such as GradCam), proving its concealment and anti-detection capabilities.
[0079] like Figure 5 As shown in the figure, the entropy of WLAT's poisoned samples is close to that of clean samples, while the entropy of poisoned samples of methods such as BadNets is significantly lower (the output is more biased towards the target label and has low randomness). STRIP defense determines whether a sample is poisoned by entropy. The higher the entropy value, the more random the output is, and the harder it is to detect as a poisoned sample. WLAT's high entropy value shows that it can bypass STRIP's detection logic, proving that it is highly resistant to defense methods based on output randomness analysis (such as STRIP).
[0080] like Figure 6 As shown in the figure, in the pruning defense, the attack success rate (ASR) of WLAT decreases the slowest with the increase of pruning rate, and the ASR begins to decrease when the pruning rate reaches 95%, while the ASR of other methods (such as BadNets and Blend) decreases significantly at low pruning rates. This shows that the backdoor features of WLAT are more dispersed in association with model neurons and are not easily removed by pruning operations, proving that it has excellent robustness against defense methods based on model pruning (such as Fine Pruning) and can maintain attack effectiveness after model compression.
[0081] As shown in Table 1, by comparing the performance of different methods in BA (success rate related indicators) and ASR (attack success rate), the advantages of the method of the present invention in attack effectiveness are intuitively demonstrated, and it is clear that it can achieve the attack goal more efficiently than other methods, providing data support for the practicality of the method. Clean is a clean model that has not been attacked by backdoors, BadNets is the most classic backdoor attack method, which embeds a visible white square trigger at a fixed position (such as the lower right corner) of the clean image, and modifies the label of the poisoned sample to the target label, Blend is a backdoor attack method based on image mixing, and WaNet is one of the latest hidden backdoor attack methods. It uses image distortion fields to generate triggers and embeds invisible backdoors in the image through deformation operations, but in some cases it still causes image distortion and poor visual indicators (such as SSIM).
[0082] The anomaly index in Table 2 reflects the concealment of the attack. Comparing the anomaly indexes of different methods can highlight the advantage of the method of the present invention in concealment, indicating that it can achieve the attack without significantly changing the visual effect of the image, reduce the risk of being detected, and enhance the concealment and security of the attack.
[0083] Table 3 By changing the high-frequency embedding amount and analyzing its impact on the attack success rate, the optimal high-frequency embedding amount can be determined, the attack effect can be optimized, and a basis can be provided for adjusting parameters in practical applications to achieve the best attack performance, thereby improving the operability and effectiveness of the method.
[0084] As shown in Table 4, analyzing the impact of different initial triggers on the attack success rate helps to screen out the optimal trigger, improve the reliability and stability of the attack, ensure efficient attacks in different scenarios, and enhance the adaptability and robustness of the method.
[0085] Table 1 Comparison of attack effectiveness results of the proposed method in terms of BA (%) and ASR (%)
[0086] Table 2 Comparison of abnormal index results of various methods
[0087] Table 3 The influence of high frequency embedding amount on attack success rate in the ablation experiment of the present invention
[0088] Table 4 The impact of the initial trigger selection on the attack success rate of the present invention
[0089] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by technicians in this technical field within the essential scope of the present invention should also fall within the protection scope of the present invention.
Claims
1. A hidden backdoor attack method based on dynamic mask and multi-level feature fusion, characterized in that: The specific steps include: S1, select a trigger graph and scale the trigger graph; S2, performs three-level discrete wavelet transform on the clean image and wavelet transform on the trigger image to extract high-frequency information; S3, embedding the high-frequency components of the trigger image into the high-frequency part of the multi-level wavelet transform of the clean image through linear combination, and generating the poisoned image using the cubic inverse discrete wavelet transform; S4, optimizes the poisoning strategy by using the backdoor random masking strategy and weak trigger training and strong trigger attack strategy.
2. According to claim 1, a hidden backdoor attack method based on dynamic mask and multi-level feature fusion is characterized in that: Step S1 specifically includes the following steps: S1.1, the size of the trigger image is 1 / 2 of the original clean image. After a discrete wavelet transform, the size of the high-frequency component of the first trigger image is 1 / 4 of the original clean image, and the size of the high-frequency component of the second-level wavelet transform of the original clean image matches; S1.2, the size of the trigger image is 1 / 4 of the original clean image. After one discrete wavelet transform, the size of the high-frequency component of the second trigger image is 1 / 8 of the original clean image, matching the size of the high-frequency component of the third-level wavelet transform of the original clean image.
3. According to claim 1, a hidden backdoor attack method based on dynamic mask and multi-level feature fusion is characterized in that: Step S2 specifically includes the following steps: S2.1, for the clean image x c Perform three-level discrete wavelet transform to extract low-frequency components A1, A2, A3 and high-frequency components H1, D1, V1; H2, D2, V2; H3, D3, V3 step by step: ; ; ; in, is discrete wavelet transform; S2.2, scale the trigger image into two pictures of different sizes, represented as: ; in, and is the scaling function; and is the scaled image; Then and Perform a wavelet transform: ; ; in, and They are and The low-frequency component of and They are and high frequency components.
4. According to claim 1, a hidden backdoor attack method based on dynamic mask and multi-level feature fusion is characterized in that: Step S3 specifically includes the following steps: S3.1, fusion of high-frequency components of wavelet transform at different scales: ; ; in, and is the embedding coefficient, is the secondary high frequency component with the backdoor embedded. It is the third-level high-frequency component with the backdoor embedded; S3.2, generate the poisoned image P by three inverse discrete wavelet transforms: ; ; ; in, is the second-level low-frequency approximate component generated by the third-level inverse wavelet transform, is the first-level low-frequency approximate component generated by the second-level inverse wavelet transform; is the inverse discrete wavelet transform.
5. According to claim 1, a hidden backdoor attack method based on dynamic mask and multi-level feature fusion is characterized in that: Step S4 specifically includes the following steps: S4.1, randomly mask the trigger image with pixels, and the masked area covers the sensitive area with pixel values close to 0 or 255 in the original clean image; in the training stage, the masked area is enlarged, and the mask ratio is 70%-90%, and in the inference stage, the masked area is reduced, and the mask ratio is 5%-20%; S4.2, in the training phase, a weak trigger mode is adopted to reduce the embedding coefficients α and β, and the trigger strength is reduced to enhance concealment; In the inference phase, we switch to the strong trigger mode, increase the embedding coefficients α and β, and reduce the mask area to ensure the effectiveness of the trigger attack.
6. According to claim 5, a hidden backdoor attack method based on dynamic mask and multi-level feature fusion is characterized in that: The value range of β is 0.3 to 0.5, with 0.3 in the training phase and 0.5 in the inference phase.