An Adversarial Attack Method Considering Both L2 Loss and L0 Loss

By taking into account the adversarial attack methods that take into account both L2 loss and L0 loss, the noise direction and dimension of the adversarial samples are optimized, and the problem of distortion generated by the adversarial samples in the prior art is solved, thereby achieving smaller L2 distortion and higher attack success rate.

CN115512190BActive Publication Date: 2025-06-27GUANGZHOU YIZHI TECHNOLOGY TRANSFER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211246656.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-12
Publication Date
2025-06-27
Estimated Expiration
2042-10-12

AI Technical Summary

Technical Problem

When generating adversarial samples, it is difficult to accurately estimate the gradient information of the black box model, resulting in greater distortion of the adversarial samples, and modifying the pixels involves the entire image, affecting the visual effect.

Method used

Adopting an adversarial attack method that takes into account both L2 loss and L0 loss is adopted. The initial adversarial noise direction is generated through the Sign-OPT attack, the noise dimension insignificance matrix is ​​calculated, and the threshold is optimized using binary search to obtain the final adversarial sample.

Benefits of technology

This method reduces L2 distortion and the number of attacked pixels in the adversarial sample, improves the attack success rate, and does not require repeated input of multiple model paths, saving computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115512190B_ABST
    Figure CN115512190B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image deep learning, and discloses an adversarial attack method that takes into account both L2 loss and L0 loss, including the following steps: First step: First, use Sign-OPT attack to generate the initial adversarial noise direction θ0 and distance λ0; Second step: Calculate the noise dimension unimportance matrix β; Third step: Use binary search to find a threshold t such that after setting the values higher than t in β to 0 and the values lower than t to 1, it still satisfies f(x0 + λ0θ0·β)!= y0, where x0 represents the original image and y0 represents the correct classification category of the neural network for x0, and perform dimension optimization on the initial adversarial noise direction θ0 to obtain the final noise direction θ' = β·θ0. This adversarial attack method that takes into account both L2 loss and L0 loss, compared with the prior art, does not require repeated input of multiple model paths multiple times, saves computing resources, and has fewer attacked pixels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image deep learning, and specifically to an adversarial attack method that takes into account both L2 loss and L0 loss. Background Art

[0002] With the wide application of deep learning in various fields, its security issues have gradually attracted people's attention. The deep learning convolutional neural network (CNN) will misjudge the input samples with added perturbations, and these samples with added perturbations are called adversarial samples, which have become a hot topic of research in recent years. Since adversarial samples pose a threat to the security of deep learning systems, especially the black-box attacks implemented by attackers without the need for model parameters are more serious, it is necessary to study adversarial samples for the proposal of defense strategies and the enhancement of model robustness.

[0003] At present, there are many studies on black-box attacks, but most black-box attacks generate global perturbations through gradient estimation to fool the target model. As we all know, for black-box models, especially hard-label black-box models, it is difficult for attackers to obtain the gradient information of the model. Attackers often estimate the gradient information by building surrogate models or transforming the search for adversarial samples into optimization problems. However, the gradient information estimated in this way is often not accurate enough. Therefore, the adversarial samples generated by gradient estimation often produce large distortions, and the modified pixels involve the entire image, thus affecting the visual effect of the adversarial samples.

[0004] The prior art is as follows:

[0005] Prior art solution 1: A method, system and terminal for generating adversarial samples by multi-path aggregation, 2022.

[0006] This invention discloses a method, system and terminal for generating adversarial samples by multi-path aggregation, belonging to the technical field of deep learning. Multiple model paths are established; random perturbation information is added to the original image respectively to obtain multiple first perturbed images; the original image is input into the first model path, and at the same time, multiple first perturbed images are input into other model paths respectively. The gradients of each neural network model are calculated, and the gradients of each neural network model are subjected to adaptive weight aggregation processing, and the image samples generated by each neural network model are updated according to the gradients obtained by the adaptive weight aggregation processing. This step is cycled multiple times, and the final adversarial sample is output. The first perturbed image of this invention synthesizes external perturbation factors and has strong generalization; through adaptive weight aggregation processing, various perturbation factors of the image are fitted, and the generalization of the adversarial sample is improved.

[0007] Prior art solution 2: A method, device, electronic device and storage medium for generating adversarial samples, 2022.

[0008] The invention relates to the field of artificial intelligence technology, and provides an adversarial sample generation method, device, electronic device and storage medium. The method includes: obtaining an original image; adding perturbation noise generated based on a multi-dimensional Gaussian distribution to the original image to obtain an adversarial sample of the original image. The adversarial sample gives an unexpected output of a preset image recognition model, and the original image can give a correct output of the preset image recognition model. By adding perturbation noise generated based on a multi-dimensional Gaussian distribution to the original image, the invention obtains an adversarial sample of the original image, can obtain the adversarial sample more efficiently and quickly, and further improves the robustness verification efficiency of the preset image recognition model.

[0009] Prior art solution 3: An adversarial attack generation method and system based on random change of image brightness, 2021.

[0010] The invention discloses an adversarial sample generation method and system based on random transformation of image brightness, which collects sample data for visual image classification and recognition, including input images and corresponding label data; constructs a deep neural network model for generating adversarial samples; performs data augmentation on the input image brightness through adversarial sample data, uses the momentum iterative FGSM image adversarial algorithm to solve the network model, searches for adversarial perturbations in the direction of the input gradient of the target loss function, and performs infinity norm limitation on the adversarial perturbations, and generates adversarial samples by maximizing the target loss function of the sample data on the network model. The invention introduces random transformation of image brightness into the adversarial attack, effectively eliminates overfitting in the process of generating adversarial samples, and improves the success rate and transferability of the adversarial sample attack.

[0011] Disadvantages of the prior art:

[0012] For the prior art solution 1, the disadvantages of this solution are: 1) The image needs to be input into multiple model channels repeatedly for many times, and the probability values and weights are calculated, which increases the memory occupancy and consumes a large amount of computing resources. 2) The number of models cannot cover all existing commonly used models, and there is a certain error between the probability value of the integrated model obtained by calculating the weights and the probability value of the real target model. 3) The generated noise covers the entire image and modifies many pixels.

[0013] For the prior art solution 2, the disadvantages of this solution are: 1) The initial perturbation noise based on the multi-dimensional Gaussian distribution will affect the visual effect of the final adversarial sample. 2) The generated noise covers the entire image and modifies many pixels.

[0014] For the existing technical solution 3, the disadvantages of this solution are as follows: 1) There are certain differences between the surrogate model and the specific target model, and the memory-based black-box attack based on migration is not ideal. 2) This method generates global adversarial perturbations, which are likely to generate noise in the smooth background area of the image, reducing the visual effect. Therefore, we propose an adversarial attack method that takes into account both L2 loss and L0 loss. Summary of the Invention

[0015] (1) Technical problems to be solved

[0016] In view of the deficiencies of the prior art, the present invention provides an adversarial attack method that takes into account both L2 loss and L0 loss, and solves the above problems.

[0017] (2) Technical solution

[0018] To achieve the above object, the present invention provides the following technical solution: An adversarial attack method that takes into account both L2 loss and L0 loss, comprising the following steps:

[0019] The first step: First, use the Sign-OPT attack to generate the initial adversarial noise direction θ0 and distance λ0;

[0020] The second step: Calculate the noise dimension unimportance matrix β;

[0021] The third step: Use binary search to find a threshold t such that after setting the values higher than t in β to 0 and the values lower than t to 1, it still satisfies f(x0 + λ0θ0·β) ≠ y0, where x0 represents the original image, y0 represents the correct classification category of the neural network for x0, and the initial adversarial noise direction θ0 is dimensionally optimized to obtain the final noise direction θ' = β·θ0.

[0022] Preferably, the specific steps of the first step are as follows:

[0023] S11: Randomly generate a large number of direction vectors θ, and then calculate the shortest distance g(θ) required to obtain the adversarial sample in each direction. The θ corresponding to the minimum g(θ) is an initial noise direction, g(θ) is the initial distance λ0, and the method for obtaining θ is as follows:

[0024]

[0025] S12: Obtain the update direction of θ such that the adversarial sample obtained in the new direction has vector The calculation of is carried out by means of symbolic gradient estimation, as follows:

[0026]

[0027] Where Q represents the number of random Gaussian samplings, and μ q represents the q-th Gaussian sampling vector, and sign(g(θ + εμ) - g(θ)) represents the sign gradient, and the calculation method is as follows:

[0028]

[0029] Repeat the process S2 to obtain the initial noise direction θ0 and the corresponding distance λ0.

[0030] Preferably, the specific steps in the second step are as follows:

[0031] S21: Randomly set a part of the dimensions of θ0 to zero to obtain θ0 * , θ0 * The obtaining method is as follows:

[0032]

[0033] ω i is a randomly distributed 0\1 matrix, and then calculate whether x0 + λ0θ0 * is an adversarial sample to obtain the sign matrix S i :

[0034]

[0035] Where R(·) represents flipping the elements in the 0\1 matrix, that is, changing 0 elements to 1 and 1 elements to 0;

[0036] S22: Calculate the sign matrix weight α i , multiply the sign matrix by a corresponding weight value, and the weight value calculation content is as follows:

[0037]

[0038] L2(·) represents calculating the L2 distance, where γ i reflects the information of the dimensions where θ0 is set to zero each time, and the calculation method is as follows:

[0039] γ i = R(ω i )·θ0;

[0040] S23: Calculate the noise dimension unimportance matrix, and the calculation method is as follows:

[0041]

[0042] Preferably, the specific content in the third step is as follows:

[0043] Through the binary search algorithm, find a threshold t such that t satisfies the following formula:

[0044]

[0045] Among them, Bin(β, ξ) means setting the values greater than ξ in β to 0 and the values less than ξ to 1;

[0046] First, take the initial upper bound high and lower bound low of the bisection method as maxβ and minβ respectively, and then the following can be obtained And determine whether f(x0 + λ0θ0Bin(β, ξ)) ≠ y0 holds. If it holds, take high = mid; if it does not hold, take low = mid, and repeat the search process until high - low > 10 -6 When outputting the threshold t = high, and after obtaining the threshold t, the final adversarial example x can be obtained;

[0047] x = x0 + λ0θ0Bin(β, t).

[0048] (III) Beneficial effects

[0049] Compared with the prior art, the present invention provides an adversarial attack method that takes into account both L2 loss and L0 loss, and has the following beneficial effects:

[0050] 1. This adversarial attack method that takes into account both L2 loss and L0 loss, compared with the prior art, does not need to repeatedly input multiple model paths many times, saves computing resources, and has fewer pixels under attack.

[0051] 2. This adversarial attack method that takes into account both L2 loss and L0 loss, compared with the prior art, the advantages of the present invention are that it optimizes both L2 loss and L0 loss.

[0052] 3. This adversarial attack method that takes into account both L2 loss and L0 loss, compared with the prior art, does not need to use a surrogate model and has fewer pixels under attack. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a schematic diagram of the process of generating initial noise;

[0054] Figure 2 It is a schematic diagram of using the bisection method to find the threshold t. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0056] Please refer toFigure 1-2 , an adversarial attack method that takes into account both L2 loss and L0 loss, includes the following steps:

[0057] (1) Process of generating initial noise

[0058] Our method optimizes the L2 distortion of adversarial samples and also optimizes the L0 distortion. Therefore, we first use the Sign-OPT attack method to generate an L2-optimized initial noise. This method regards the hard-label attack as the problem of finding the direction with the shortest distance to the decision boundary. The specific process is as follows:

[0059] ① Randomly generate a large number of direction vectors θ, and then calculate the shortest distance g(θ) required to obtain adversarial samples in each direction respectively. The θ corresponding to the minimum g(θ) is an initial noise direction, and g(θ) is the initial distance λ0. The method for obtaining θ is shown in Equation 1:

[0060]

[0061] ② Obtain the update direction of θ such that the adversarial samples obtained in the new direction have The vector is obtained by using the symbolic gradient estimation method, as shown in Equation 2:

[0062]

[0063] where Q represents the number of random Gaussian sampling times, and μ q represents the q-th Gaussian sampling vector. sign(g(θ + εμ) - g(θ)) represents the symbolic gradient, and the calculation method is shown in Equation 3:

[0064]

[0065] ③ Repeat process ②. Obtain the initial noise direction θ0 and the corresponding distance λ0.

[0066] The process of generating initial noise is shown in the appendix Figure 1 .

[0067] (2) Process of calculating the noise dimension unimportance matrix

[0068] Since very little model information can be obtained from the black-box hard-label model, the process of obtaining the initial noise in step 1 regards the hard-label attack as the problem of finding the direction with the shortest distance to the decision boundary. However, when obtaining θ and There are relatively large errors in the process. And in the later stage of this algorithm, the L2 loss of adversarial samples decreases very slowly. If more accurate values are desired, a large number of queries need to be added. Therefore, the present invention limits the number of queries in step 1, and only uses it to initially generate an initial noise direction optimized by L2 and the corresponding distance, and then we perform dimensional optimization on this basis. The specific process is as follows:

[0069] ① Randomly set a part of the dimensions of θ0 to zero to obtain θ0 * , θ0 * The obtaining method is as shown in formula 4:

[0070]

[0071] ω i is a randomly distributed 0\1 matrix. Then calculate whether x0 + λ0θ0 * is an adversarial sample to obtain the sign matrix S i :

[0072]

[0073] where R(·) represents flipping the elements in the 0\1 matrix, that is, changing 0 elements to 1 and 1 elements to 0.

[0074] ② Calculate the sign matrix weight α i . In order to reduce more L2 loss while optimizing the noise dimensions, we multiply the sign matrix by a corresponding weight value. This weight value is calculated using the maximum-minimum normalization method in formula 5,

[0075]

[0076] L2(·) represents calculating the L2 distance, where γ i reflects the information of the dimensions where θ0 is set to zero each time. The calculation method is as formula 6:

[0077] γ i = R(ω i )·θ0 (6)

[0078] ③ Calculate the unimportance matrix of noise dimensions, and the calculation method is as formula 7:

[0079]

[0080] (3) Noise dimension optimization process

[0081] Through steps (1) and (2), we obtain the initial noise direction θ0, distance λ0, and the noise dimension unimportance matrix β. The noise dimension unimportance matrix β reflects the importance degree of each dimension of the initial noise direction θ0. The larger the value, the lower the importance degree of that dimension, indicating a greater possibility that the noise dimension can be set to zero. In this patent, a binary search algorithm is used to find a threshold t such that t satisfies Equation 8:

[0082]

[0083] where Bin(β, ξ) means setting the values in β greater than ξ to 0 and the values less than ξ to 1.

[0084] We first take the initial upper bound (denoted by high) and lower bound (denoted by low) of the binary method as maxβ and minβ respectively. Then we can obtain and determine whether f(x0 + λ0θ0Bin(β, ξ)) ≠ y0 holds. If it holds, then take high = mid; if it does not hold, then low = mid, and repeat the search process until high - low > 10 -6 and output the threshold t = high. For the detailed process of obtaining the threshold t, see Appendix Figure 2 After obtaining the threshold t, the final adversarial example x can be obtained (see Equation 9).

[0085] x = x0 + λ0θ0Bin(β, t) (9)

[0086] Through a large number of experiments, the functions and significance of the present invention are verified. The proposed scheme of the present invention has been experimentally tested on the ImageNet - 1k, CIFAR10, and MNIST datasets. The performance comparison between the present invention and the existing black - box hard - label attack techniques is shown in Table 1. It can be seen that under the same query - number limit, our scheme can achieve a smaller L2 distortion and a higher attack success rate (within the allowable range of the same L2 distortion). In addition, the number of pixels modified by our attack method has been greatly reduced.

[0087] This scheme can generate adversarial examples with a smaller L2 distortion and fewer attacked pixels. Therefore, it is necessary to consider designing a defense scheme against this attack method. Due to the effectiveness, stealthiness of this scheme, and the fact that it does not require model gradient information, gradient shielding, slightly restricting the number of model queries, and adversarial training are not obstacles to the present invention.

[0088] This solution does not require estimating the gradient information of the image. By generating a matrix where the dimensions of the noise are unimportant, it guides the system to optimize the dimensions of the initial noise direction, resulting in the optimization of both the L2 loss and the L0 loss of the finally generated adversarial samples. Therefore, it has strong attack power and stealthiness. We suggest that the field of image deep learning can draw on this invention to better improve the robustness of the system.

[0089] Table 1 Comparison of L2 Losses under Different Datasets and Models

[0090]

[0091]

[0092] Note: SR is the attack success rate

[0093] Table 2 Overall Results of L2 and L0 Losses of this Patent

[0094]

[0095] Note: PP is the proportion of attacked pixels reduced on the premise of obtaining the attack results in Table 1.

[0096] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An adversarial attack method that takes into account both L2 loss and L0 loss, characterized in that, It includes the following steps: The first step: First, use the Sign-OPT attack to generate the initial adversarial noise direction θ0 and distance λ0; The second step: Calculate the noise dimension unimportance matrix β; The third step: Use binary search to find a threshold t such that after setting the values higher than t in β to 0 and the values lower than t to 1, it still satisfies f(x0 + λ0θ0·β) ≠ y0, where x0 represents the original image and y0 represents the correct classification category of the neural network for x0, and optimize the dimension of the initial adversarial noise direction θ0 to obtain the final noise direction θ' = β·θ0; The specific steps of the first step are as follows: S11: Randomly generate a large number of direction vectors θ, and then calculate the shortest distance g(θ) required to obtain the adversarial sample in each direction respectively. The θ corresponding to the minimum g(θ) is an initial noise direction, and g(θ) is the initial distance λ0. The method for obtaining θ is as follows: S12: Obtain the update direction of θ such that the adversarial sample obtained in the new direction has vector is obtained by sign gradient estimation as follows: where Q represents the number of random Gaussian samplings, and μ q represents the q-th Gaussian sampling vector, and sign(g(θ + εμ) - g(θ)) represents the sign gradient, and the calculation method is as follows: Repeat the process S12 to obtain the initial noise direction θ0 and the corresponding distance λ0; The specific content in the third step is as follows: Through the binary search algorithm, find a threshold t such that t satisfies the following formula: where Bin(β, ξ) means setting the values greater than ξ in β to 0 and the values less than ξ to 1; First, take the initial upper bound high and lower bound low of the bisection method as maxβ and minβ respectively, and then we can obtain and determine whether f(x0 + λ0θ0Bin(β, ξ)) ≠ y0 holds. If it holds, then take high = mid; if it does not hold, then take low = mid. Repeat the search process until high - low > 10 -6 At this time, output the threshold t = high. After obtaining the threshold t, the final adversarial example can be obtained x ; x = x0 + λ0θ0Bin(β, t).

2. The adversarial attack method that takes into account both L2 loss and L0 loss according to claim 1, wherein: The specific steps in the second step are as follows: S21: Randomly set a part of the dimensions of θ0 to obtain The obtaining method is as follows: ω i is a randomly distributed 0 / 1 matrix, and then calculate whether it is an adversarial sample to obtain the sign matrix S i : where R(·) means flipping the elements in the 0\1 matrix, that is, changing the 0 elements to 1 and the 1 elements to 0; S22: Calculate the weight α of the symbol matrix i , multiply the symbol matrix by a corresponding weight value, and the weight value calculation method is as follows: L2(·) represents calculating the L2 distance, where γ i reflects the information of the dimension where θ0 is set to zero each time, and the calculation method is as follows: γ i = R(ω i )·θ0; S23: Calculate the noise dimension unimportance matrix, and the calculation method is as follows:

Citation Information

Patent Citations

  • Target detection-oriented physical attack adversarial patch generation method and system

    CN113361604A

  • Gradient-based adversarial sample generation method and system

    CN114663665A