Adversarial Sample Generation Method Based on Average Gradient
By adopting an adversarial sample generation method based on average gradient in migration attacks, iteratively searches for the optimal integrated gradients, solving the problem of limited quality of adversarial sample generation in the prior art, and improving the attack success rate of other models.
Patent Information
- Application Number
- CN202311040941.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-08-17
AI Technical Summary
When generating adversarial samples, the existing migration attack methods directly average the output of multiple models without processing, resulting in limited quality of generating adversarial samples, affecting the success rate of attacks on other models.
Adversarial sample generation method based on average gradient is adopted to search for the optimal integration gradient through loop iteration, reducing the difference between the integration gradient and multiple individual gradients, thereby improving the quality and migration success rate of the adversarial sample.
By fusing the gradients of multiple models, an integration gradient with small variance is obtained, which can capture the update direction of adversarial samples more accurately, and improve the quality and migration success rate of generated adversarial samples.
Smart Images

Figure CN117010479B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of model attacks. More specifically, it relates to a method for generating adversarial samples based on average gradient. Background Art
[0002] In recent years, deep neural networks have been proven to be vulnerable to deliberately designed adversarial samples, which are usually generated by adding imperceptible perturbations to images. Although the generated adversarial samples have no visual differences, they can cause deep neural networks to misclassify. More seriously, the adversarial samples generated from one model also maintain a certain degree of aggressiveness against other models, even if there are significant differences in the structures of these two models. We call this kind of attack transfer attack. Since deep neural networks have been widely used in many real-world tasks, such as autonomous driving and face recognition, this phenomenon has attracted extensive attention in the industrial and academic fields.
[0003] Existing transfer attack methods can be divided into three categories: gradient-based attacks, input transformation attacks, and model ensemble attacks. Gradient-based attacks calculate the gradient of the loss function with respect to the input sample to obtain the gradient direction of the input sample, and then perturb the input sample according to the gradient direction. Input transformation attacks perform data augmentation on the input sample by methods such as flipping, cropping, and random recombination to reduce the fitting degree of the generated adversarial samples to the surrogate model. Model ensemble attacks achieve attacks by simultaneously integrating the outputs of multiple models to obtain a direction for generating adversarial samples suitable for multiple models.
[0004] The idea behind model ensemble attacks is that if an adversarial sample can be aggressive against multiple models simultaneously, then it may be able to misclassify more black-box models. Model ensemble attacks simultaneously focus on the outputs of multiple deep neural networks, fuse these outputs to find an ensemble output that can represent the optimization directions of multiple deep neural networks, and add perturbations to clean samples according to this ensemble output. This method can greatly improve the quality of adversarial samples. Although existing ensemble attacks have achieved very good results in generating adversarial samples, most model ensemble attacks still directly average the outputs of multiple models to obtain the ensemble result and implement model ensemble attacks by applying attacks to this ensemble result. However, there are very large differences in the structures between different models. Directly fusing the outputs of multiple models without any processing often limits the quality of the generated adversarial samples and affects the attack success rate when migrating to other models. Summary of the Invention
[0005] The object of the present invention is to overcome the deficiencies of the prior art and provide an adversarial sample generation method based on average gradient, which uses average gradient for model integration to reduce the difference between the integrated gradient and multiple individual gradients and improve the quality of generated adversarial samples.
[0006] To achieve the above object of the invention, the adversarial sample generation method based on average gradient of the present invention comprises the following steps:
[0007] S1: Initialize the iteration number t = 1 and the initial adversarial sample x 0 as the sample X to be attacked;
[0008] S2: Initialize the adversarial sample for this round of iteration
[0009] S3: Preset K target models as required, and input the initial adversarial sample into the K target models to obtain the corresponding model gradients g k , k = 1, 2,..., K;
[0010] S4: Search for the optimal integrated gradient, including the following steps:
[0011] S4.1: Initialize the iteration number m = 0 and the internal adversarial sample
[0012] S4.2: Randomly select a target model k * from the K target models;
[0013] S4.3: Input the internal adversarial sample into the target model k * to update the corresponding model gradient g k* ;
[0014] S4.4: Perform weighted averaging on the current K model gradients to obtain the internal integrated gradient
[0015]
[0016] where w k represents the preset weight, where 0 < w k < 1 and
[0017] S4.5: Update the internal adversarial sample using the following formula to obtain the updated internal adversarial sample
[0018]
[0019] Among them, clip() represents the clipping function, which restricts the generated adversarial sample within a circle centered on the clipped sample x with a preset perturbation threshold ε as the radius, and sign() represents the sign function;
[0020] S4.6: Determine whether m < M, where M represents the preset maximum number of iterations. If so, go to step S4.7; otherwise, go to step S4.8;
[0021] S4.7: Let m = m + 1, and return to step S4.2;
[0022] S4.8: Obtain the optimal integrated gradient for this round
[0023] S5: Update the adversarial sample using the following formula to obtain the updated adversarial sample x t :
[0024]
[0025] S6: Determine whether t < T, where T represents the preset maximum number of iterations. If so, go to step S7; otherwise, go to step S8;
[0026] S7: Let t = t + 1, and return to step S2;
[0027] S8: Use the adversarial sample x obtained in the last round T as the adversarial sample Y corresponding to the sample X to be attacked;
[0028] The above method is applied to the ImageNet - compatible dataset, which is 1000 pictures of different categories selected from the standard ImageNet dataset
[0029] The adversarial sample generation method based on the average gradient in the present invention iterates the adversarial sample multiple times. During each iteration, cyclic iteration is used to search for the optimal integrated gradient. During the process of searching for the optimal integrated gradient, a target model is randomly selected from the target models each time, and its gradient is updated based on the internal adversarial sample, then the new internal integrated gradient is obtained by weighted averaging, and then the internal adversarial sample is updated based on this internal integrated gradient. Such cycling is used to obtain the optimal integrated gradient for each round, so as to iteratively generate the adversarial sample of the sample to be attacked.
[0030] The model integration method adopted in the present invention is realized by fusing the gradients of multiple models. The gradient can intuitively represent the update direction of the adversarial sample during the generation process of the adversarial sample. The integrated gradient after fusion is continuously iteratively optimized in an inner loop by averaging the randomly updated model gradients and the remaining model gradients maintained in the memory. In this way, an integrated gradient with a smaller variance from the gradients of each model can be obtained, thereby capturing a more accurate update direction of the adversarial sample in the model integration attack, effectively reducing the variance between the integrated gradient and the gradients of each model in the model integration attack, improving the quality of the generated adversarial sample, and increasing the transfer success rate of the generated adversarial sample. Description of the Drawings
[0031] Figure 1 is the flowchart of the specific implementation manner of the adversarial sample generation method based on the average gradient of the present invention;
[0032] Figure 2 is the flowchart of searching for the optimal integrated gradient in the present invention. Specific Embodiment
[0033] The following describes the specific embodiments of the present invention with reference to the drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may obscure the main content of the present invention, these descriptions will be omitted here.
[0034] Embodiment
[0035] Figure 1 is the flowchart of the specific implementation manner of the adversarial sample generation method based on the average gradient of the present invention. As Figure 1 shown, the specific steps of the adversarial sample generation method based on the average gradient of the present invention include:
[0036] S101: Initialize the iteration number t = 1, and the initial adversarial sample x 0 is the sample X to be attacked.
[0037] S102: Initialize the adversarial sample for this round of iteration
[0038] That is to say, in the first iteration, the sample X to be attacked is used as the initial adversarial sample for this round Otherwise, the adversarial sample x obtained in the previous round of iteration t-1 is used as the initial adversarial sample for this round
[0039] S103: Obtain the model gradient:
[0040] According to actual needs, K target models are preset in advance, and the initial adversarial sample Input into K target models to obtain the corresponding model gradients g k , k = 1, 2, …, K.
[0041] S104: Search for the optimal integrated gradient:
[0042] In the present invention, an inner loop is used to search for the optimal integrated gradient in each round, reducing the difference between the integrated gradient and multiple individual gradients, and updating the adversarial examples in the direction of the integrated gradient obtained in the inner loop.
[0043] Figure 2 is the flow chart for searching the optimal integrated gradient in the present invention. As Figure 2 shown, the specific steps for searching the optimal integrated gradient in the present invention include:
[0044] S201: Initialize the iteration number m = 0 and initialize the internal adversarial examples
[0045] S202: Randomly select a target model:
[0046] Randomly select a target model k from the K target models * .
[0047] S203: Update the gradient of target model k * :
[0048] Input the internal adversarial examples into target model k * and update the corresponding model gradient
[0049] S204: Calculate the internal integrated gradient:
[0050] The gradient can more directly represent the update direction of the model or adversarial examples in different tasks. In the present invention, the internal integrated gradients of multiple models are obtained by weighted averaging the gradients of multiple models, and the gradients of the models are obtained through the backpropagation process of samples on the corresponding models. Therefore, the internal integrated gradient is obtained by weighted averaging the current K model gradients
[0051]
[0052] where w k represents a preset weight, where 0 < w k < 1 and
[0053] Compared with the existing integrated model predictions, logical outputs, and model losses, the gradient integration method proposed in the present invention can obtain a more accurate update direction of adversarial examples.
[0054] S205: Update the internal adversarial examples:
[0055] Move the internal adversarial examples in the opposite direction of their gradient descent direction, that is, update the internal adversarial examples using the following formula to obtain the updated internal adversarial examples
[0056]
[0057] where clip() represents the clipping function, that is, restricting the generated adversarial examples within a circle centered on the clipped sample x with a preset perturbation threshold ε as the radius, that is, restricting the perturbation size of the generated adversarial examples between [-ε, +ε], so that the difference between the adversarial examples and the original samples cannot be too large, and sign() represents the sign function.
[0058] S206: Judge whether m < M, where M represents the preset maximum number of iterations. If so, go to step S207; otherwise, go to step S208.
[0059] S207: Let m = m + 1, and return to step S202.
[0060] S208: Obtain the optimal integrated gradient:
[0061] Obtain the optimal integrated gradient for this round
[0062] S105: Update the adversarial examples:
[0063] Update the adversarial examples in the opposite direction of their gradient descent direction, that is, update the adversarial examples using the following formula to obtain the updated adversarial example x t :
[0064]
[0065] S106: Judge whether t < T, where T represents the preset maximum number of iterations. If so, go to step S107; otherwise, go to step S108.
[0066] S107: Let t = t + 1, and return to step S102.
[0067] S108: Obtain the adversarial examples:
[0068] Take the adversarial example x obtained in the last round T as the adversarial example Y corresponding to the sample X to be attacked.
[0069] To better illustrate the technical effects of the present invention, specific examples are used to experimentally verify the present invention. In this embodiment, different maximum attack step sizes are used to test the attack effect of adversarial samples on black-box models on the ImageNet-compatible dataset. The ImageNet-compatible dataset is 1000 pictures of different categories specially selected from the standard ImageNet dataset, and these pictures can be successfully classified on most models.
[0070] In this embodiment, the Ens (Ensemble Attack) algorithm and the SVRE (Stochastic Variance Reduced Ensemble Attack) algorithm are selected as comparison methods, and the Transfer Attack Success Rate is used as an indicator to evaluate the attack method proposed in this patent. The transfer attack success rate represents the attack success rate when the adversarial samples generated on the surrogate model are transferred to other black-box models. During the test, the transfer attack success rates of the method of the present invention and the existing methods on different baselines are tested on a variety of baseline methods and defense methods. Among them, the baseline methods include I-FGSM (Iterative Fast Gradient Sign Method), MIM (Momentum Iterative Method), TIM (Translation Invariant Method), TI-DIM (Diversity Input Method), SI-TI-DIM (Scale Invariance Method). In the defense method, the integrated defense model includes Inc-v3 ens3 、Inc-v3 ens4 、IncRes-v2 ensThere are three types (see the literature Tramèr, Florian, Kurakin A, Papernot N, et al. Ensemble Adversarial Training: Attacks and Defenses[J]. 2017. DOI: 10.48550 / arXiv.1705.07204.). The defense cleaning algorithms include NIPS-R3 (ranked third in the NIPS 2017 defense competition), RS (Randomized Smoothing), ComDefend (Compression Model to Defend Adversarial Examples), R&P (Random Resizing And Padding), and FD (Feature Distillation). Table 1 is a statistical table of the attack success rates of the present invention and comparative methods on the black-box model after defense training. Table 2 is a statistical table of the attack success rates of the present invention and comparative methods on the black-box model after defense cleaning.
[0071]
[0072]
[0073] Table 1
[0074]
[0075] Table 2
[0076] As shown in Table 1 and Table 2, the performance of the present invention on different baselines is better than that of the two existing comparative methods, indicating the effectiveness of the present invention.
[0077] Although the above-described illustrative specific embodiments of the present invention have been described to facilitate those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
Claims
1. An adversarial sample generation method based on average gradient, characterized in that, it includes the following steps: S1: Initialize the iteration number \(t = 1\) and the initial adversarial example \(x\) 0 as the sample to be attacked \(X\); S2: Initialize the adversarial example for this round of iteration S3: Preset K target models according to actual needs, and input the initial adversarial sample into the K target models to obtain the corresponding model gradients g k , where k = 1, 2, …, K; S4: Search for the optimal integrated gradient, including the following steps: S4.1: Initialize the number of iterations \(m = 0\) and initialize the internal adversarial examples S4.2: Randomly select a target model k from the K target models * ; S4.3: Input the internal adversarial examples into the target model k * and update the corresponding model gradients S4.4: Obtain the internal integrated gradient by weighted averaging the gradients of the current K models where, w k represents a preset weight, where 0 < w k < 1 and S4.5: Update the internal adversarial examples using the following formula to obtain the updated internal adversarial examples where clip() represents a clipping function that restricts the generated adversarial sample within a circle centered on the clipped sample x with a preset perturbation threshold ε as the radius, and sign() represents a sign function; S4.6: Determine whether m < M, where M represents the preset maximum number of iterations. If so, go to step S4.7; otherwise, go to step S4.8; S4.7: Let m = m + 1 and return to step S4.2; S4.8: Obtain the optimal integrated gradient for this round S5: Update the adversarial example using the following formula to obtain the updated adversarial example x t : S6: Determine whether t < T, where T represents the preset maximum number of iterations. If so, go to step S7; otherwise, go to step S8; S7: Let t = t + 1 and return to step S2; S8: Use the adversarial example \(x\) obtained in the last round T as the adversarial example \(Y\) corresponding to the sample \(X\) to be attacked; The above method is applied to the ImageNet - compatible dataset, which is 1000 pictures of different categories selected from the standard ImageNet dataset.