An Adversarial Attack Method Based on Targeted Data Augmentation

By constructing a class activation graph matrix of deep neural networks, the degree of pixel contribution is evaluated, and the adversarial samples with enhanced target data is generated, which solves the problem of insufficient migration in the prior art and achieves a high success rate of adversarial samples in black box attacks.

CN116433924BActive Publication Date: 2025-07-29NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310416256.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-18
Publication Date
2025-07-29
Estimated Expiration
2043-04-18

AI Technical Summary

Technical Problem

The existing adversarial sample generation method ignores the importance of different areas of the picture to the model prediction results during the data enhancement process, resulting in poor performance of some deep neural networks after robust training and reinforcement.

Method used

By constructing a class activation graph matrix of deep neural networks, we evaluate the contribution of each pixel in the input image to the model prediction results, mask important parts and fuse them with the input image, generate multiple enhanced images with different scales, and calculate the average gradient information to generate adversarial samples.

Benefits of technology

It significantly improves the migration of adversarial samples and improves the attack success rate in black box attack scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116433924B_ABST
    Figure CN116433924B_ABST
Patent Text Reader

Abstract

The present invention provides a targeted data augmentation-based adversarial attack method, belonging to the fields of deep learning and computer vision. The technical solution is as follows: First, generate the Class Activation Map (CAM) matrix of the input image under a specific classification by the Deep Neural Network (DNN), and mask some elements with larger values in the CAM matrix according to different ratios. Then, fuse the masked CAM matrix with the input image respectively to enhance the input image. After that, each fused image is copied multiple times, and the pixel values of each copied image are scaled according to different ratios. Subsequently, calculate the average gradient information of all generated images and update the momentum information. Finally, calculate the adversarial perturbation according to the momentum information and update the adversarial sample. Repeat the above steps T times until the final adversarial sample is generated. Compared with other invention techniques, the present invention performs targeted data augmentation on the input image based on the contribution degree of different regions in the input image to the DNN output result, and can significantly improve the transferability of the generated adversarial sample without reducing the success rate of white-box attacks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an adversarial attack method for deep neural networks, and particularly to an adversarial attack method based on targeted data augmentation. Background Art

[0002] Deep Neural Network (DNN) has been widely applied in many fields such as natural language processing, computer vision, and recommendation systems. However, it cannot be ignored that DNN is vulnerable to adversarial examples, that is, adding imperceptible perturbations to the original data can make the model make wrong predictions with high confidence. In addition, related research shows that adversarial examples have transferability, that is, adversarial examples generated for one model can, to a large extent, also deceive other models under the same training task. Based on the transferability of adversarial examples, an attacker can first generate adversarial examples based on a surrogate model and then transfer the generated adversarial examples to the target model, thereby implementing a black-box attack on the target model. In this scenario, improving the transferability of the generated adversarial examples can improve the success rate of the black-box attack. Data augmentation is a commonly used method to improve the transferability of the generated adversarial examples. However, existing data augmentation-based adversarial attack methods such as Admix (Xiaosen Wang, Xuanran He, Jingdong Wang, and Kun He. Admix: Enhancing the transferability of adversarial attacks. In proceedings of the IEEE / CVF International Conference on Computer Vision. 2021: 16158-16167), SCM-P (Donggon Jang, Sanghyeok Son, and Dae-Shik Kim. Strengthening the Transferability of Adversarial Examples Using Advanced Looking Ahead and Self-CutMix. In proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2022: 148-155) usually enhance the input image by randomly fusing the input image with itself or images of other classes. Such methods ignore the different importance degrees of different regions in the image for the model prediction results and have a certain randomness in the data augmentation process. Therefore, the transferability of the generated adversarial examples is poor for some DNNs strengthened by robust training. Summary of the Invention

[0003] To overcome the problem of insufficient transferability of adversarial examples generated by existing adversarial attack methods, the present invention provides an adversarial attack method based on targeted data augmentation. The method first constructs a Class Activation Map (CAM) matrix of the DNN for the input image under a specified class, and comprehensively evaluates the contribution degree of each pixel in the input image to the DNN prediction result; then, some pixels in the CAM matrix that contribute more to the output result of the DNN model are masked and fused with the input image to perform targeted data augmentation on the input image; after that, the images after targeted data augmentation are copied multiple times respectively, and the pixel values of each copy are scaled according to different ratios; finally, the adversarial perturbation is calculated based on the average gradient information of all the augmented images, so as to generate adversarial examples. The present invention performs targeted data augmentation on the input image based on the importance degree of different pixels in the input image to the DNN output result, and generates adversarial examples from the perspective of affecting the attention area of the DNN on the input image, which can significantly improve the transferability of the generated adversarial examples.

[0004] The technical solution adopted by the present invention to solve its technical problems: An adversarial example generation method based on targeted data augmentation, which is characterized by including the following steps:

[0005] Step 1: Initialize the parameters \(g_0 = 0\); \(x'_0 = x\); \(α = ∈ / T\), where \(x\) is the original image; \(x'_0\) is the initial adversarial example; \(α\) is the perturbation step size, and its value is a fixed constant; \(∈\) is the maximum perturbation value; \(T\) is the number of iterations; \(g_0\) is the initial momentum, and its value is a zero matrix with the same shape as the original image.

[0006] Step 2: During the calculation of the adversarial perturbation in the \(t\)-th round of iteration, for the input image \(x'\) t , calculate the CAM matrix of the DNN model for it under the specified class:

[0007]

[0008] In the formula, is the CAM matrix of the classifier \(f\) for the input image \(x'\) t under the class \(c\); \(A\) is the feature map output by the last convolutional layer of the classifier \(f\); \(A\) n is the data of the \(n\)-th channel in the feature layer \(A\); \(H(·)\) is a bilinear interpolation function used to convert the processing result for the feature layer \(A\) into the same shape as the input image \(x'\) t ; is the weight of the feature layer \(A\) n , and its calculation formula is as follows:

[0009]

[0010] In the formula, \(y\) cRepresents the predicted score of the DNN for class c; Represents the value corresponding to the coordinates (i, j) in the n-th channel of feature layer A; Z is equal to the product of the height and width of feature layer A.

[0011] Step 3: Generate the percentile set Q:

[0012]

[0013] In the formula, q i Represents the i-th percentile value, and its calculation method is as follows:

[0014]

[0015] Among them, σ is a positive integer.

[0016] Step 4: According to different percentiles q i in the set Q, mask the pixels in the CAM matrix that have a relatively high impact on the DNN prediction result to generate multiple masked CAM matrices Among them, The calculation method is as follows:

[0017]

[0018] In the formula, represents a matrix with the same shape as its input z, and all elements of this matrix are equal to the q i percentile of all elements of matrix z; Sign is the sign function; Min(z, 0) means replacing all elements in matrix z that are larger than 0 with 0.

[0019] Step 5: Respectively fuse the multiple generated masked CAM matrices with the input image x′ t in a certain proportion to generate multiple enhanced images Among them, The calculation method is as follows:

[0020]

[0021] In the formula, γ is the fusion coefficient; ξ is a random perturbation, and ξ ∈ [-∈, ∈], ∈ is a constant, and its value is equal to the maximum perturbation value of the generated adversarial sample.

[0022] Step 6: Copy each of the multiple enhanced images m2 times, and scale the pixel values of all images in the j-th (j ∈ [1, m2]) copy to 1 / 2 j-1 times of the original, and calculate the average gradient information of all the currently generated enhanced images

[0023]

[0024] Wherein, m1 is the size of the percentile set Q; m2 is the number of copies of the fused image ; J(·) is the cross-entropy loss function; x′ t is the input image; y is the label value corresponding to the image; θ is the parameter of the classifier f.

[0025] Step Seven, update the momentum information g t+1 :

[0026]

[0027] wherein, μ is a constant; g t is the momentum information in the previous iteration.

[0028] Step Eight, calculate the adversarial perturbation and update the adversarial sample x′ t+1 :

[0029] x′ t+1 = x′ t + α·Sign(g t+1 ) (9)

[0030] wherein, α is a constant; Sign is the sign function.

[0031] Step Nine, repeat Step Two to Step Eight for a total of T times to obtain the adversarial sample x′ for the original image x T .

[0032] The beneficial effects of the present invention are as follows: The method first evaluates the contribution degree of each pixel in the input image to the DNN prediction result by generating the CAM matrix of the input image under a specific category of the DNN output, and masks some pixels in the CAM matrix that have a greater impact on the DNN output result based on the evaluation result; then, fuses the masked CAM matrix with the input image to achieve targeted data augmentation; after that, copies the targeted augmented images multiple times, and scales the pixel values of each copy according to different ratios; finally, generates an adversarial sample based on the average gradient information of all the augmented images. This method performs targeted data augmentation on the input image based on the contribution degree of different pixels in the input image to the DNN prediction result. The adversarial sample generated by using this method has better transferability, that is, a higher attack success rate can be achieved in the black-box attack scenario. Description of the Drawings

[0033] Figure 1 is the implementation flowchart of the present invention.

[0034] Figure 2It is a comparison chart of the attack and transfer capabilities of the adversarial examples generated by the method of the present invention and other baseline algorithms. Detailed implementation manners

[0035] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0036] The background conditions and specific implementation manners of the present invention will be described in detail below with reference to the accompanying drawings and specific implementation examples. Here, the schematic diagrams and examples of the present invention are only for the explanation of the present invention and are not intended to limit the present invention.

[0037] The overall structural block diagram of an adversarial attack method based on target data augmentation according to the present invention is as Figure 1 shown, and specifically includes the following steps:

[0038] 1. Initialize the parameters: g0 = 0; x′0 = x; α = ∈ / T, where x is the original image; x′0 is the initial adversarial example; α is the perturbation step size, and its value is a fixed constant; ∈ = 16 is the maximum perturbation value; T = 10 is the number of iterations; g0 is the initial momentum, and its value is a zero matrix with the same shape as the original image.

[0039] 2. During the t-th round of iterative calculation of the adversarial perturbation, for the input image x′ t , calculate the CAM matrix of the DNN model for it under the specified category:

[0040]

[0041] In the formula, is the CAM matrix of the classifier f for the input image x′ t under the category c; A is the feature map output by the last convolutional layer of the classifier f; A n is the data of the n-th channel in the feature layer A; H(·) is the bilinear interpolation function used to convert the processing result for the feature layer A into the same shape as the input image x′ t ; is the weight of the feature layer A n , and its calculation formula is as follows:

[0042]

[0043] In the formula, yc represents the prediction score of the DNN for the category c; Denote the value corresponding to the coordinate (i, j) in the n-th channel of the feature layer A; Z is equal to the product of the height and width of the feature layer A.

[0044] 3. Set the parameter σ to 10, and generate the percentile set Q = {10, 20, 30, 40, 50, 60, 70, 80, 90} according to the following formula.

[0045]

[0046] 4. According to different percentiles q in the set Q i , mask some pixels in the CAM matrix that have a relatively high impact on the DNN prediction result; generate multiple masked CAM matrices Among them, The calculation method is as follows:

[0047]

[0048] In the formula, represents a matrix with the same shape as its input z, and all element values of this matrix are equal to the q i percentile of all element values of the matrix z; Sign is the sign function; Min(z, 0) means replacing all elements in the matrix z that are larger than 0 with 0.

[0049] 5. Respectively, fuse the multiple generated masked CAM matrices with the input image x′ t in a certain proportion to generate multiple enhanced images Among them, The calculation method is as follows:

[0050]

[0051] In the formula, γ = 0.6 is the fusion coefficient; ξ is a random perturbation, and ξ ∈ [-16, 16].

[0052] 6. Copy each of the multiple enhanced images five times. For the j-th (j ∈ [1, 5]) copy, scale all the pixel values of the images to 1 / 2 j-1 times of the original, and calculate the average gradient information of all the currently generated enhanced images

[0053]

[0054] In the formula, J(·) is the cross-entropy loss function; x′ t is the input image; y is the label value corresponding to the original image; θ is the parameter of the classifier f.

[0055] 7. Update the momentum information gt+1 :

[0056]

[0057] where μ = 1.0; g t is the momentum information in the previous iteration.

[0058] 8. Calculate the adversarial perturbation and update the adversarial sample x′ t+1 :

[0059] x′ t+1 = x′ t + α·Sign(g t+1 ) (8)

[0060] where α is a constant; Sign is the sign function.

[0061] 9. Repeat steps two to eight for a total of T times to obtain the adversarial sample x′ for the original image x T .

[0062] Based on 1000 images randomly selected from 1000 categories of the ImageNet dataset, the method of the present invention, Admix, and the SCM-P algorithm are respectively used to generate 1000 image adversarial samples for the Inception-v3 model. Then, the 1000 image adversarial samples generated by the three methods are respectively used to attack four commonly trained DNN models (Inception-v3, Inecption-v4, Inception-ResNet-v2, and ResNet-101, abbreviated as: Inc-v3, Inc-v4, IncRes-v2, and Res101 respectively) and three DNN models fortified by robust training (ens3-adv-Inception-v3, ens3-adv-Inception-v3, and ens-adv-Inception-ResNet-v2, abbreviated as: Inc-v3 ens3 、Inc-v3 ens4 、and IncRes-v2 ens ), and the attack success rates are as Figure 2 shown. It can be seen from Figure 2 that the white-box attack success rates of the three methods against the Inc-v3 model are all close to 100%; in the black-box attack scenarios against several other models, especially the three DNN models fortified by robust training, the attack success rate of the method of the present invention is significantly higher than the other two. It can be seen that the method of the present invention can significantly improve the transferability of the generated adversarial samples while not reducing the white-box attack success rate compared with the other two comparison algorithms.

Claims

1. An adversarial attack method based on targeted data augmentation, characterized in that It includes the following steps: Step 1, initialize the parameters: g0 = 0; x′0 = x; α = ∈ / T, where x is the original image; x′0 is the initial adversarial sample; α is the perturbation step size, whose value is a fixed constant; ∈ is the maximum perturbation value; T is the number of iterations; g0 is the initial momentum, whose value is a zero matrix with the same shape as the original image; Step 2. During the t-th round of iterative calculation of adversarial perturbations, for the input image x′ t , calculate the Class Activation Map (CAM) matrix of the DNN model for it under the specified class: Wherein, is the CAM matrix of the classifier f for the input image x' t under the category c; A is the feature map output by the last convolutional layer of the classifier f; A n is the data of the nth channel in the feature layer A; H(·) is a bilinear interpolation function used to convert the processing result for the feature layer A into the same shape as the input image x' t ; is the weight of the feature layer A n , and its calculation formula is as follows: where y c represents the prediction score of the DNN for class c; represents the value corresponding to the coordinates (i, j) in the n-th channel of feature layer A; Z is equal to the product of the height and width of feature layer A; Step 3, generate the percentile set Q: where q i represents the i-th percentile value, and its calculation method is as follows: where σ is a positive integer; Step 4: According to different percentiles q in set Q i , mask the part of pixels in the obtained CAM matrix that have a relatively high impact on the DNN prediction result to generate multiple masked CAM matrices Among them, The calculation method is as follows: In the formula, represents a matrix with the same shape as its input z, and all the element values of this matrix are equal to the q i percentile of all the element values of matrix z; Sign is the sign function; Min(z, 0) means replacing all the elements in matrix z that are greater than 0 with 0; Step 5. Respectively fuse the generated multiple masked CAM matrices with the input image x′ t in a certain proportion to generate multiple enhanced images wherein is calculated as follows: In the formula, γ is the fusion coefficient; ξ is a random perturbation and satisfies ξ ∈ [-∈, ∈], ∈ is a constant, and its value is equal to the maximum perturbation value of the generated adversarial sample; Step 6: Copy each of the multiple enhanced images m2 times, and scale the pixel values of all the images in the j-th (j ∈ [1, m2]) copy to 1 / 2 of the original value, and calculate the average gradient information of all the currently generated enhanced images j-1 times, and calculate the average gradient information of all the currently generated enhanced images Where m1 is the size of the percentile set Q; and m2 is the number of copies of the fused image ; J(·) is the cross-entropy loss function; x′ t is the input image; y is the label value corresponding to the image; θ is the parameter of the classifier f; Step Seven: Update the momentum information g t+1 : where μ is a constant; g t is the momentum information in the previous iteration; Step VIII: Calculate the adversarial perturbation and update the adversarial example x' t+1 : x′ t+1 = x′ t + α·Sign(g t+1 ) (9) where α is a constant; Sign is the sign function; Step Nine: Repeat Steps Two to Eight for a total of T times to obtain the adversarial example x′ for the input image x T .

Citation Information

Patent Citations

  • Intelligent migration confrontation method and system for electromagnetic signal identification under incomplete information

    CN115600083A

  • Method and system for defending against adversarial sample in image classification, and data processing terminal

    US20230022943A1