Adversarial sample detection method through feature attention and high-frequency information enhancement
Through local histogram equalization and high-frequency information extraction, combined with the attention mechanism of the convolutional block attention module, the problem of DNN being fragile to adversarial samples is solved, and the effectiveness and model independence of adversarial sample detection are achieved.
Patent Information
- Application Number
- CN202411798059.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Deep neural networks (DNNs) are fragile to adversarial samples, and existing detection methods usually rely on prior knowledge of the target model and may fail when the model is upgraded.
Local histogram equalization (LHE) is used to amplify the adversarial perturbation of the image, extract high-frequency information, and use the channel and spatial attention mechanism in the convolutional block attention module (CBAM) to extract the feature weights between the image and the high-pass information flow data to determine whether the input image is an adversarial sample.
This method can effectively detect adversarial samples without the need to understand prior knowledge of the target model, exhibits significant improvements under various attack strategies, and is applicable to different models and data sets.
Smart Images

Figure CN120014380A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of adversarial sample detection technology, and in particular to an adversarial sample detection method enhanced by feature attention and high-frequency information. Background Art
[0002] Deep neural networks (DNNs) have demonstrated great influence and achieved remarkable success in computer vision applications such as image classification, autonomous driving, and identity recognition. In particular, in the field of image classification, the performance of state-of-the-art neural networks even exceeds that of humans. However, recent studies have shown that DNNs are vulnerable to adversarial examples, which are generated by adding adversarial perturbations that are imperceptible to humans. These adversarial examples can easily lead to wrong predictions and pose a significant threat to the security of DNNs in real-world applications.
[0003] To address this challenge, many defense methods have been developed, which can be divided into three categories: adversarial training, adversarial purification, and adversarial detection. Adversarial training aims to enhance the robustness of the model by incorporating adversarial samples into the training process; adversarial purification aims to remove adversarial perturbations by preprocessing input samples. Unfortunately, adversarial training and purification techniques often lead to performance degradation in deep neural network (DNN) applications. In contrast, adversarial detection provides a more flexible solution. It actively distinguishes adversarial samples from original samples by adding a detection module before applying DNN.
[0004] Given these limitations, detection-based defense strategies have recently attracted widespread attention as a promising solution. Some studies on adversarial detection have attempted to exploit the feature differences between adversarial samples and original samples. For example, some scholars have used feature maps generated by the output of the hidden layer of a deep neural network (DNN) to distinguish adversarial samples. However, these methods perform poorly when facing attack methods with less perturbations such as DeepFool and C&W. To address this problem, some scholars have used pixel artifacts and confidence artifacts to identify adversarial samples, the latter of which combines data smoothing and attention mechanism techniques. These defense methods are closely related to the prior knowledge of the target model and are therefore model-specific, and the robustness improvement of one model cannot be easily transferred to other models. Other scholars have used steganalysis to enrich models to detect adversarial samples. However, most of these detectors have been shown to be unstable and prone to failure when facing stronger attacks and processing more complex datasets.
[0005] In summary, deep neural networks (DNNs) are vulnerable to adversarial examples, which can cause DNNs to make incorrect predictions. To protect these networks, many effective detection methods have been developed. Unfortunately, these methods usually rely on prior knowledge of the target model and may fail when the model is upgraded. Summary of the invention
[0006] To solve the problems existing in the prior art, the purpose of the present invention is to provide an adversarial sample detection method enhanced by feature attention and high-frequency information, which can protect pre-trained classifiers and resist various types of attacks without interacting with the target model.
[0007] To achieve the above object, the technical solution adopted by the present invention is: a method for detecting adversarial samples by enhancing feature attention and high-frequency information, comprising the following steps:
[0008] Step 1: Use local histogram equalization (LHE) to amplify the adversarial perturbation of the image to enhance the feature difference between the adversarial sample and the normal sample;
[0009] Step 2: extract high-frequency information from the enlarged image;
[0010] Step 3: Use the convolutional block attention module to extract the channel and spatial feature weights between the image and the high-pass information flow data using channel and spatial attention; input the output results into the same detection network to determine whether the input image is an adversarial sample.
[0011] As a further improvement of the present invention, the step 1 is specifically as follows:
[0012] First, the image is divided into multiple small pixel regions, then the grayscale histogram of each pixel region is calculated, histogram equalization is performed, and the pixel values are remapped according to the equalized histogram; after the processing is completed, the pixel regions are merged to form a complete image.
[0013] As a further improvement of the present invention, the step 2 is specifically as follows:
[0014] First, use bicubic interpolation to downsample the samples to 1 / 2 of the original size to obtain an image of size C×W / 2×H / 2 that removes high-frequency information while retaining low-frequency information; then, upsample the downsampled image by 2 times to restore the original size C×W×H to obtain the low-frequency information of the image; finally, subtract the low-frequency information from the original sample to obtain the high-frequency information; where C is the number of channels, W is the width, and H is the height.
[0015] As a further improvement of the present invention, in step 3, the convolution block attention module includes a channel attention module and a spatial attention module, wherein the channel attention module enhances the response to important features by adaptively adjusting the importance of each channel; and the spatial attention module optimizes the spatial representation of the feature map by paying attention to different positions in the image.
[0016] As a further improvement of the present invention, the step 3 is specifically as follows:
[0017] The input attention module is used to extract the two-stream attention parameters in the channel and spatial dimensions. Then, the ResNet network is used to calculate the adversarial score of the amplified perturbed image and its high-frequency information. Finally, the results are combined to predict whether the input image is an adversarial sample.
[0018] The beneficial effects of the present invention are:
[0019] 1. The defense method for detecting adversarial samples of the present invention does not rely on prior knowledge of the model or specific details of the target model, can effectively generalize the defense performance to deal with different attacks, and is well adapted to various models.
[0020] 2. The present invention uses local histogram equalization to amplify adversarial perturbations and utilizes the attention mechanism to extract high-frequency information to enhance the detection effect, all of which do not rely on prior knowledge of the target classifier.
[0021] 3. The detection method is evaluated on two real databases and its effectiveness is demonstrated. Compared with the previous state-of-the-art methods, the method of the present invention shows significant improvement under various attack strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 Schematic diagram of an original image and its corresponding adversarial sample in an embodiment of the present invention;
[0023] Figure 2 A schematic diagram of a method according to an embodiment of the present invention;
[0024] Figure 3 A schematic diagram of a process for extracting high-frequency information from an image in an embodiment of the present invention;
[0025] Figure 4 Schematic diagram of the structure of channel and spatial attention in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0027] Example
[0028] An adversarial sample detection method enhanced by feature attention and high-frequency information is used to detect adversarial samples without prior knowledge of the target image classification model. Figure 1 As shown, in Figure 1 In the figure, LHE refers to the image processed using local histogram equalization. High-pass refers to the extracted high-frequency information. It can be intuitively seen that the adversarial perturbation is amplified after the adversarial sample is subjected to local histogram equalization (LHE). In addition, it is observed that there is a significant difference in high-pass information between the adversarial sample and the original sample. Therefore, the core idea of this embodiment is that LHE can amplify adversarial perturbations, which mainly destroy the high-frequency information in the image. Based on this observation, LHE is applied to amplify adversarial perturbations, and the high-pass information in the image is used to detect adversarial samples. In order to improve the detection performance, channel and spatial attention mechanisms are used to effectively strengthen the intrinsic connection between the image and the high-pass information. In this way, the most critical information between the adversarial sample and the original sample can be easily identified and selected.
[0029] The present embodiment is further described below:
[0030] This embodiment uses a dual-stream architecture to detect adversarial samples. Figure 2 As shown in Figure 1, local histogram equalization is first applied to amplify the adversarial perturbation. Next, high-frequency information is extracted from the amplified image. Then, channel and spatial attention mechanisms are applied to the image and its high-pass information in the channel and spatial dimensions to extract attention weights. Finally, DNN* determines whether the input sample is an adversarial sample.
[0031] 1. Enhanced feature differences:
[0032] Local histogram equalization (LHE) is used to enhance feature differences. The process first divides the image into small 8×8 pixel regions, then calculates the grayscale histogram of each region, performs histogram equalization, and remaps the pixel values according to the equalized histogram. After the processing is completed, these regions are merged to form a complete image. Since adversarial perturbations are usually local and random, LHE can effectively amplify these perturbations, thereby enhancing the feature differences between adversarial samples and original samples, which helps improve the performance of the classifier.
[0033] 2. Extract high-frequency information from the image:
[0034] like Figure 3 As shown in Figure 2, the process of extracting high-frequency information is shown. Figure 3In the figure, (a) is the original image, (b) is the downsampled image, (c) is the low-frequency information, and (d) is the high-frequency information. First, the sample is downsampled to 1 / 2 of the original size using bicubic interpolation to obtain an image of size C×W / 2×H / 2. This step removes the high-frequency information while retaining the low-frequency information. Next, the downsampled image is upsampled by 2 times to restore its original size C×W×H. This process obtains the low-frequency information of the image. Finally, the high-frequency information is obtained by subtracting this low-frequency information from the original sample.
[0035] 3. Channel and spatial attention mechanism:
[0036] This embodiment uses a convolutional block attention module (CBAM), which provides a simple and powerful mechanism for enhancing the representation capability of a convolutional neural network (CNN). Channel and spatial attention are used to extract channel and spatial feature weights between images and high-pass information flow data. Figure 4 As shown in Figure 1, CBAM consists of two main modules: channel attention module and spatial attention module. The channel attention module enhances the response to important features by adaptively adjusting the importance of each channel, while the spatial attention module optimizes the spatial representation of the feature map by focusing on different locations in the image. This approach improves the performance of the model while maintaining computational efficiency.
[0037] 4. Detection module:
[0038] The input attention module is used to extract the dual-stream attention parameters in the channel and spatial dimensions. Next, the ResNet network is used to calculate the adversarial score of the amplified perturbed image and its high-frequency information. Finally, I combined these results to predict whether the input image is an adversarial sample. In order to further improve the method of this embodiment, a large number of experiments were carried out to evaluate the impact of different model parameter sizes on detection accuracy. Two data sets (CIFAR-10 and ImageNet) and three attack methods (PGD, DI-FGSM, and C&W) were specifically evaluated. In addition, three different ResNet models were considered: ResNet18, ResNet34, and ResNet50. The detection results are shown in Table 1. Considering the detection performance and computing resources, the ResNet18 model was selected as the final implementation solution.
[0039] Table 1 Comparison of the accuracy (%) of the method of this embodiment when using different ResNet models to resist various attacks (the target model of CIFAR-10 and ImageNet is VGG16)
[0040]
[0041] The present embodiment is further described below through experiments:
[0042] 1. Experimental setup:
[0043] Datasets: Extensive experiments were conducted on the CIFAR-10 and ImageNet datasets. CIFAR-10 contains 50,000 training images and 10,000 test images, each of which has a size of (3×32×32). ImageNet consists of 1,281,167 training images, 50,000 validation images, and 10,000 test images, all of which have a dimension of (3×224×224). Due to the large size of ImageNet, this example focuses on 10 specific categories in ILSVRC2012: flatworms, spiny lobsters, goldfish, roosters, American alligators, night snakes, harvester ants, snails, jellyfish, and turtles. Initially, each category consisted of 1,300 training images and 50 test images. To enhance the test set, 300 images were randomly selected from the training set of each category and redistributed to the test set, so that each category finally had 1,000 training images and 350 test images.
[0044] Classifier: This example uses VGG16, VGG1, and ResNet18 models as classifiers. These models are trained using the Adam optimizer with a batch size of 64, a learning rate of 0.0001, and more than 50 training rounds.
[0045] Comparison Methods: To demonstrate the superiority of the proposed architecture in this example, it is compared with several state-of-the-art adversarial sample detection methods, which serve as benchmarks. These methods include Local Intrinsic Dimension (LID), Mahalanobis Distance Adversarial Detection (MAD), Noise Feature Map and RGB Image Detection (AEND), Pixel Artifact and Confidence Artifact Detection (PACA), and Attention-Based Dual-Stream Detection (ADS).
[0046] Attack Techniques: Three attack techniques, namely PGD, DIFGSM and C&W, are used to evaluate the effectiveness of the method in this embodiment. For PGD and DIFGSM, the perturbation size (ε) is set to 4 / 255 on CIFAR-10 and ImageNet.
[0047] Performance Metrics: To evaluate the effectiveness of the proposed method, accuracy (Acc), precision (PRE), recall (REC) and area under the curve (AUC) were used to quantify the detection performance.
[0048] Details of model training: We use the Adamax optimizer to train our detector model, and the initial learning rate is set to 0.0001. The training process is carried out for 30 rounds, and the loss function uses cross entropy. The parameter settings of AEND, PACA, and ADS are consistent with the method of this embodiment. For LID and MAD, a logistic regression model is used to identify adversarial samples based on the extracted features.
[0049] 2. Performance evaluation on different datasets:
[0050] To demonstrate the effectiveness of the method in this embodiment, extensive experiments are conducted on the CIFAR-10 and ImageNet datasets.
[0051] Performance on the CIFAR-10 dataset: Table 2 shows the detection results of different methods on the CIFAR-10 dataset. It can be seen that LID performs poorly under various attacks. LID has a low recall rate, which means that many adversarial samples are incorrectly classified as original samples. MAD performs better than LID and shows stable performance when defending against various attacks. AEND performs well in resisting PGD and DIFGSM, but still encounters difficulties in defending C&W. ADS can effectively defend against various attacks. Among the benchmark methods, PACA performs best. It is worth noting that the method proposed in this embodiment outperforms PACA in most indicators when defending against various attacks. The adversarial perturbations generated by PGD and DIFGSM are larger than those generated by C&W. Therefore, PGD and DIFGSM destroy more high-frequency information, and the method of this embodiment can effectively identify these attacks. In addition, the method of this embodiment also performs well in defending against C&W attacks, and improves the AUC indicator by 3.4% compared with MAD.
[0052] Table 2 Comparison of the accuracy (ACC), precision (PRE), recall (REC), and AUC scores (%) of various adversarial detection methods when resisting different attacks. For the CIFAR-10 dataset, VGG16 is the target model. Adversarial samples are generated by performing various attacks on the target model.
[0053]
[0054] Performance on the ImageNet dataset: To further evaluate the effectiveness of the methods, extensive experiments were conducted on the large-scale ImageNet dataset. Table 3 shows the detection results of different methods on ImageNet. LID and MAD perform poorly when dealing with large-scale adversarial images. The defense performance of AEND on the ImageNet dataset is similar to its performance on CIFAR-10. Compared with the CIFAR-10 dataset, AEND's performance in defending PGD and DIFGSM has improved, but it is still ineffective in defending against C&W attacks. PACA and ADS perform similarly, and their detection performance has declined compared to CIFAR-10. It is worth noting that the method of this embodiment achieved the highest scores in all indicators when defending against PGD and DIFGSM. However, compared with ADS and PACA, the method of this embodiment is less effective in defending against C&W attacks.
[0055] Table 3 Comparison of the accuracy (ACC), precision (PRE), recall (REC), and AUC scores (%) of various adversarial detection methods when resisting different attacks. For the ImageNet dataset, VGG16 is the target model. Adversarial samples are generated by performing various attacks on the target model.
[0056]
[0057] 3. Cross-defense against different attack methods:
[0058] For actual deployment, it is impractical to constantly retrain the detector to defend against a constant stream of new attacks. Adversarial examples are generated by using PGD, DIFGSM, and C&W attacks on the VGG16 model. Table 4 shows the generalization ability of different detection methods.
[0059] Table 4. Comparison of generalization accuracy (ACC) of different detection methods on the CIFAR-10 dataset. Adversarial samples are generated by using PGD, DIFGSM, and C&W attacks on the VGG16 model.
[0060]
[0061] It can be seen that LID exhibits poor generalization performance regardless of the attack method used for training. MAD exhibits stable generalization performance, and the detection accuracy always remains at around 90%. When PGD or DIFGSM is used as a training attack, AEND and ADS perform poorly in defending against C&W attacks. When C&W is used as a training attack, PACA and ADS cannot effectively defend against PGD and DIFGSM. It is worth noting that when C&W is used as a training attack, the method of this embodiment exhibits the best generalization performance. However, when PGD or DIFGSM is used as a training attack, the method of this embodiment cannot effectively defend against C&W. This may be due to the fact that the perturbations generated by the C&W attack are too small.
[0062] 4. Protection of different models: This embodiment does not require prior knowledge of the model in the process of detecting adversarial samples. Therefore, this embodiment evaluates the generalization ability and transferability of the proposed method in protecting different models. Training samples with PGD and C&W attacks were generated using VGG16. All test adversarial samples with various attacks were generated using ResNet18. Table 5 shows the performance of different methods in protecting various models. LID and MAD performed poorly in defending against attacks generated by different models. When PGD (generated by VGG16) was used as a training attack, AEND had similar defense performance to ADS. When C&W (generated by VGG16) was used as a training attack, ADS could not protect different models. When PGD (generated by VGG16) was used as a training attack, PACA performed well in protecting different models. However, when C&W (generated by VGG16) was used as a training attack, it still had difficulty protecting different models. It can be seen that no matter whether C&W or PGD attack is used for training, the method proposed in this embodiment can consistently defend against adversarial samples generated by different models.
[0063] Table 5 Comparison of the accuracy (ACC) and AUC scores (%) of each detection method when protecting different models on the CIFAR-10 dataset. VGG16 is used to generate training samples with PGD or C&W attacks. ResNet18 is used to generate all test adversarial samples with various attacks.
[0064]
[0065] 5. Performance under different disturbance intensities:
[0066] In order to evaluate the ability of the method proposed in this embodiment to defend against adversarial samples with different perturbation strengths, experiments were conducted on the CIFAR-10 dataset using two attack methods, namely PGD and DIFGSM. Set ε = 1 / 255, ε = 2 / 255, ε = 4 / 255, ε = 6 / 255, and ε = 8 / 255 to generate adversarial samples. Table 6 shows the performance of different methods under various perturbation strengths. It can be seen that LID and MAD show stable defense performance under different perturbation strengths. However, their detection performance is not optimal. AEND, PACA, and ADS show similar performance trends when defending against different perturbation strengths.
[0067] When the adversarial perturbation is small, their defense performance is poor. However, as the perturbation intensity increases, their defense performance becomes stronger. It is worth noting that the method of this embodiment shows stable defense performance against various attacks at different perturbation intensities. In the method proposed in this embodiment, the accuracy index for defense performance at different perturbation intensities exceeds 93%.
[0068] Table 6 Comparison of the accuracy (%) of various adversarial detection methods for adversarial samples with different perturbation strengths on the CIFAR-10 dataset.
[0069]
[0070] 6. Ablation experiment:
[0071] In order to explore the impact of each component in the method of this embodiment, an ablation experiment was performed on the CIFAR-10 dataset to observe the effect by removing or replacing components. Table 7 shows the detection results under different settings. Our method performs best in different settings. Removing the preprocessing module (LHE) reduces the defense performance against DIFGSM and C&W attacks. Therefore, LHE effectively amplifies the adversarial perturbation. It is worth noting that the effectiveness of the high-frequency flow exceeds that of the image flow, which shows that it is indeed effective to detect adversarial samples using high-frequency information. Removing the attention mechanism causes a 1.5% drop in the defense performance against C&W attacks, which shows that the attention mechanism is also crucial.
[0072] Table 7 The accuracy of the proposed method and its four variants in defending against adversarial examples on the CIFAR-10 dataset.
[0073]
[0074] In summary, this embodiment proposes a novel adversarial sample detection method that can effectively protect pre-trained classifiers from various attack methods without interacting with the target model or acquiring its prior knowledge. First, local histogram equalization (LHE) is used to amplify adversarial perturbations, which can enhance the feature differences between adversarial samples and normal samples. Next, high-frequency information is extracted from the image. Then, channel and spatial attention mechanisms are used to extract channel and spatial feature weights between images and high-pass information flow data. The output results are input into the same detection network to determine whether the input image is an adversarial sample. The method of this embodiment is evaluated in five scenarios: performance on different data sets, cross-defense against different attack methods, protection against different models, effectiveness under different perturbation intensities, and ablation studies. The experimental results show that the method of this embodiment is highly stable and can provide effective protection against various attacks for different models.
[0075] The above-mentioned embodiments only express the specific implementation of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.
Claims
1. A method for detecting adversarial samples by enhancing feature attention and high-frequency information, characterized in that: The following steps are involved: Step 1: Use local histogram equalization (LHE) to amplify the adversarial perturbation of the image to enhance the feature difference between the adversarial sample and the normal sample; Step 2, extracting high-frequency information from the enlarged image; Step 3: Use the convolutional block attention module to extract the channel and spatial feature weights between the image and the high-pass information flow data using channel and spatial attention; The output is fed into the same detection network to determine whether the input image is an adversarial example.
2. The adversarial sample detection method based on feature attention and high-frequency information enhancement according to claim 1, characterized in that: The step 1 is specifically as follows: First, the image is divided into multiple small pixel regions, then the grayscale histogram of each pixel region is calculated, histogram equalization is performed, and the pixel values are remapped according to the equalized histogram; after the processing is completed, the pixel regions are merged to form a complete image.
3. The adversarial sample detection method based on feature attention and high-frequency information enhancement according to claim 2, characterized in that: The step 2 is specifically as follows: First, use bicubic interpolation to downsample the samples to 1 / 2 of the original size to obtain an image of size C×W / 2×H / 2 that removes high-frequency information while retaining low-frequency information; then, upsample the downsampled image by 2 times to restore the original size C×W×H to obtain the low-frequency information of the image; finally, subtract the low-frequency information from the original sample to obtain the high-frequency information; where C is the number of channels, W is the width, and H is the height.
4. The adversarial sample detection method based on feature attention and high-frequency information enhancement according to claim 3, characterized in that: In step 3, the convolution block attention module includes a channel attention module and a spatial attention module. The channel attention module enhances the response to important features by adaptively adjusting the importance of each channel; the spatial attention module optimizes the spatial representation of the feature map by paying attention to different positions in the image.
5. The adversarial sample detection method based on feature attention and high-frequency information enhancement according to claim 4, characterized in that: The step 3 is as follows: The input attention module is used to extract the two-stream attention parameters in the channel and spatial dimensions. Then, the ResNet network is used to calculate the adversarial score of the amplified perturbed image and its high-frequency information. Finally, the results are combined to predict whether the input image is an adversarial sample.
Citation Information
Patent Citations
An anti-attack defense method for a feature map attention mechanism and application
CN109948658A
Image processing method and device, storage medium and electronic equipment
CN114742738A
Adversarial sample detection method and system based on improved ViT loss distribution difference
CN118506105A