Adversarial sample generation method and system based on future gradient information, and storage medium
Through the adversarial sample generation method (LoFF-MI) based on future gradient information, multiple iterations acquire future gradient information and update momentum, solving the problem of failure to fully utilize future gradients in the existing technology, and improving the attack success rate and migration of adversarial samples.
Patent Information
- Application Number
- CN202510413711.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-22
AI Technical Summary
Existing adversarial sample generation methods fail to make full use of future gradient information, limiting the migration of generated adversarial samples.
Adversarial sample generation method (LoFF-MI) based on future gradient information is used to obtain the sum of gradients of sample points in future iterations through multiple iterations, update the momentum using future gradient information to generate the final adversarial sample.
It improves the attack success rate and migration of the adversarial samples, and achieves a more efficient attack effect.
Smart Images

Figure CN120356031A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sample generation, and in particular to an adversarial sample generation method, system and storage medium based on future gradient information. Background Art
[0002] Some scholars first generated adversarial samples through the Fast Gradient Sign Method (FGSM, an algorithm for generating adversarial samples). FGSM generates adversarial samples by adding a large amount of perturbation at one time along the direction of image gradient ascent. Subsequently, the iterative form of FGSM, the Iterative Fast Gradient Sign Method (I-FGSM), was proposed, which gradually adds perturbation through multiple iterative processes. I-FGSM has a powerful white-box attack ability, but due to its overfitting problem, it has a low black-box attack success rate. To solve this problem, the Momentum Iterative Fast Gradient Sign Method (MI), an iterative attack method based on momentum, was proposed. It replaces the gradient with momentum to avoid poor local minima, stabilizes the update direction, and effectively improves the black-box attack ability of adversarial samples. Thanks to the excellent effect and easy integration of MI, MI is widely used as a base method to improve the performance of existing adversarial attack methods.
[0003] However, MI still has deficiencies. For example, although MI introduces momentum to replace the gradient and effectively introduces historical gradient information, it does not consider future gradient information, which may limit the transferability of the adversarial samples it generates because this future gradient information is very likely to contain clues to improve transferability. Some existing works attempt to utilize future information. For example, some adversarial attack methods use Nesterov accelerated gradient and use the adversarial samples generated in the next iteration to generate gradients, but the utilization of future gradient information by these attack methods is limited and the role of future gradients is not fully exerted.
[0004] Therefore, we designed an adversarial sample generation method based on future gradient information to solve the above problems. Summary of the Invention
[0005] The object of the present invention is to solve the drawbacks existing in the prior art that most of the existing adversarial sample generation methods fail to fully utilize future gradient information, which limits the transferability of the generated adversarial samples and the utilization degree of future gradient information. Therefore, an adversarial sample generation method based on future gradient information is proposed. Taking the idea of looking from the future as the basis, the MI method is improved and called LoFF-MI. By continuously iterating forward multiple times using the existing attack method, multiple future iteration sample points are obtained. The future gradient information is obtained by calculating the sum of the gradients of these future iteration sample points. The future gradient information is used to assist in the generation of adversarial samples, thereby improving the transferability of adversarial samples.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] An adversarial sample generation method based on future gradient information, comprising the following steps:
[0008] Step 1, apply the momentum-based iterative attack method multiple times to reach the future iteration times, and generate adversarial samples generated after multiple future iterations;
[0009] Step 2, calculate the future gradient and obtain the sum of the gradients of the adversarial samples of multiple future iterations;
[0010] Step 3, use the future gradient to update the momentum. After multiple iterations, use the updated momentum to generate the final adversarial sample of this iteration.
[0011] As a further preferred solution of the present invention, in the process of applying the momentum-based iterative attack method multiple times to reach the future iteration times and generate adversarial samples generated after multiple future iterations, it includes:
[0012] Input the original image, the image classification model (classifier) and related parameters:
[0013] A classifier f with parameters θ and loss function J; an original image x with a true label y; a maximum perturbation ε and a total number of iteration rounds T; a future iteration number Q, a momentum decay factor μ; since generating adversarial samples requires multiple rounds of iteration, set the perturbation size for each step Initial adversarial sample The initial momentum g0 = 0. Before each round of iteration, clear the future gradient information.
[0014] As a further preferred solution of the present invention, the steps of applying the momentum-based iterative attack method multiple times, iterating multiple times and reaching the future iteration times, and generating adversarial samples generated after multiple future iterations are as follows:
[0015] (1a) Set the future iteration number Q and perform parameter initialization;
[0016] (1b) Apply the momentum-based iterative attack method A multiple times to obtain adversarial samples generated after multiple future iterations. The formula is as follows:
[0017]
[0018] Among them, represents the adversarial sample generated towards the i-th round of future iteration at the t-th round, represents the adversarial sample generated towards the (i + 1)-th round of future iteration at the t-th round; t is the current iteration round, with values ranging from 0 to T - 1, and i is the future iteration round, with values ranging from 0 to Q - 1; α is the perturbation size for each iteration, sign(·) is the sign function, and (A(x, f(x, θ), J) represents the gradient information generated by using the momentum-based iterative attack method A on the input sample x by the classifier f under the given model parameters θ and loss function J.
[0019] As a further preferred solution of the present invention, the future gradient is calculated, and the way to obtain the sum of the gradients of the adversarial samples for multiple future iterations is as follows:
[0020] (2a) The future gradient information G LoFF is obtained by dividing the gradient values of all sample points by the 1-norm and summing them. The formula is as follows:
[0021]
[0022] Among them, y is the correct label corresponding to the original image, is the gradient of the loss function with respect to the input image, and its value is calculated by the backpropagation algorithm in the convolutional neural network.
[0023] As a further preferred solution of the present invention, the future gradient is used to update the momentum. After multiple iterations, the updated momentum is used to generate the final adversarial sample for this iteration:
[0024] (3a) Update the momentum using the future gradient information. The formula is as follows:
[0025]
[0026] Among them, μ is the momentum decay factor, and g t is the momentum at the t-th round;
[0027] (3b) Use the updated momentum g t+1 and the adversarial sample output in the previous round to perform an update calculation to generate a new adversarial sample The formula is as follows:
[0028]
[0029] (3c) Perform the next round of iteration:
[0030] Repeat steps 1 - 3 in the new round of iteration until the total number of iteration rounds T is reached and stop, and output the final adversarial example.
[0031] An adversarial example generation system based on future gradient information, comprising:
[0032] An input module, configured to input an original image, an image classification model (classifier), and related parameters;
[0033] An iteration module, which applies an iterative attack method based on momentum, iterates multiple times and reaches the number of future iterations, and generates adversarial examples generated after multiple future iterations;
[0034] A calculation module, configured to calculate future gradients and obtain the sum of gradients of adversarial examples after multiple future iterations;
[0035] A generation module, which updates the momentum using the future gradients and generates the final adversarial example of this iteration using the updated momentum.
[0036] As a further preferred solution of the present invention, the information input by the input module includes:
[0037] A classifier f with parameters θ and a loss function J; an original image x with a true label y; a maximum perturbation ε and a total number of iteration rounds T; the number of future iterations Q, a momentum decay factor μ; set the perturbation size per step Initial adversarial example Initial momentum g0 = 0;
[0038] The iteration module performs multiple iterations in the following manner, reaches the number of future iterations, and generates adversarial examples generated after multiple future iterations:
[0039] (1a) Set the number of future iterations Q and perform parameter initialization;
[0040] (1b) Apply the iterative attack method A based on momentum multiple times to obtain adversarial examples generated after multiple future iterations. The formula is as follows:
[0041]
[0042] Wherein, represents the adversarial example generated towards the i-th round of future iteration at the t-th round, Denote the adversarial sample generated in the (t + 1)-th iteration towards the future at the t-th round; t is the current iteration round, taking values from 0 to T - 1, i is the future iteration round, taking values from 0 to Q - 1; α is the perturbation size for each iteration, sign(·) is the sign function, and (A(x, f(x, θ), J) represents that the classifier f generates gradient information for the input sample x using the momentum-based iterative attack method A under the given model parameters θ and loss function J.
[0043] As a further preferred solution of the present invention, the calculation module calculates the future gradient, and the method for obtaining the sum of the gradients of the adversarial samples for multiple future iterations is as follows:
[0044] (2a) The future gradient information G LoFF is obtained by dividing the gradient values of all sample points by the 1-norm and summing them, and the formula is as follows:
[0045]
[0046] where y is the correct label corresponding to the original image, is the gradient of the loss function with respect to the input image, and its value is calculated by the backpropagation algorithm in the convolutional neural network.
[0047] As a further preferred solution of the present invention, the generation module updates the momentum using the future gradient. After multiple iterations, the updated momentum is used to generate the final adversarial sample for this iteration:
[0048] (3a) Update the momentum using the future gradient information, and the formula is as follows:
[0049]
[0050] where μ is the momentum decay factor, and g t is the momentum at the t-th round;
[0051] (3b) Use the updated momentum g t+1 and the adversarial sample output in the previous round to perform an update calculation to generate a new adversarial sample The formula is as follows:
[0052]
[0053] (3c) Perform the next iteration:
[0054] Repeat steps 1 - 3 in the new iteration until the total number of iterations T is reached and stop, and output the final adversarial sample.
[0055] A computer-readable storage medium, characterized in that the computer-readable storage medium includes a stored computer program, wherein when the computer program is run by a processor, it controls the device where the storage medium is located to execute an adversarial sample generation method based on future gradient information
[0056] The adversarial attack method based on future gradient information proposed by the present invention adopts an inner and outer two-layer iterative manner. Each time an outer iteration is performed, several inner future iterations are required to obtain future gradient information. Compared with other gradient optimization-based adversarial attack methods, it makes full use of future gradient information, and the generated adversarial samples have higher attack success rates and transferability, thus achieving a further attack effect. Brief Description of the Drawings
[0057] Figure 1 It is a schematic flowchart of an adversarial sample generation method based on future gradient information proposed by the present invention. Detailed Embodiments
[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0059] The present invention uses the idea of Looking From the Future (LoFF) to improve the existing Momentum Iterative Fast Gradient Sign Method (MI), and proposes an adversarial sample generation method based on future gradient information (LoFF-MI), which overcomes the deficiency that traditional adversarial attack methods fail to make full use of future gradient information, and effectively improves the attack success rate and transferability of adversarial samples.
[0060] Embodiment 1:
[0061] Combined with Figure 1 , the present invention is an adversarial sample generation method based on future gradient information, including the following steps:
[0062] Step 1, first input the original image, the image classification model (classifier) and related parameters, including:
[0063] (1) A classifier f with parameters θ and a loss function J. The classifier f in this embodiment can be an image classification model based on a convolutional neural network such as VGG16 or ResNet50, or an image classification model based on a Transformer such as ViT.
[0064] (2) An original image x with a true label y;
[0065] (3) The maximum perturbation ε and the total number of iterations T;
[0066] (4) The number of future iterations Q, the momentum decay factor μ;
[0067] Since generating adversarial samples requires multiple rounds of iteration, set the perturbation size for each step Initial adversarial sample The initial momentum g0 = 0. The future gradient information G LoFF Is initially 0 and the future gradient information is cleared before each round of iteration.
[0068] Using the idea of looking ahead, apply the momentum-based iterative attack method multiple times to obtain the adversarial samples generated after multiple future iterations. The specific steps are as follows:
[0069] (1a) Set the number of future iterations Q and perform parameter initialization;
[0070] (1b) Apply the momentum-based iterative attack method A multiple times. A represents the momentum-based iterative attack method Momentum Iterative Fast Gradient Sign Method (MI), which replaces the gradient with momentum to avoid poor local minima, stabilize the update direction, and effectively improve the black-box attack ability of adversarial samples. By judging whether the future iteration number is reached, if the future iteration number is not reached, continue to apply the momentum-based iterative attack method A to generate adversarial samples until the future iteration number is reached to obtain the adversarial samples generated after multiple future iterations. The formula is as follows:
[0071]
[0072] Among them, Represents the adversarial sample generated towards the i-th future iteration at the t-th round, Represents the adversarial sample generated towards the (i + 1)-th future iteration at the t-th round; t is the current iteration round, taking values from 0 to T - 1, and i is the future iteration round, taking values from 0 to Q - 1; In particular, when i is 0, Is the initial adversarial sample at the t-th round of iteration α is the perturbation size for each iteration, sign(·) is the sign function, and (A(x, f(x, θ), J) represents the gradient information generated by using the momentum-based iterative attack method A on the input sample x by the classifier f under the given model parameters θ and loss function J. The gradient information is a data calculated by the backpropagation algorithm.
[0073] Step 2, calculate the future gradient, that is, the sum of the gradients of the adversarial samples for multiple future iterations. The specific steps are as follows:
[0074] (2a) Future gradient information G LoFF Obtained by dividing the gradient value of all sample points by the 1-norm and summing them, the formula is as follows:
[0075]
[0076] Where y is the true label corresponding to the original image, is the gradient of the loss function with respect to the input image, and its value can be calculated by the backpropagation algorithm in the convolutional neural network.
[0077] Step 3, update the momentum using the future gradient; use the updated momentum to generate the final adversarial sample for this iteration, the specific steps are as follows:
[0078] (3a) Update the momentum using the future gradient information, the formula is as follows:
[0079]
[0080] Where μ is the momentum decay factor, g t is the momentum at the t-th round, g t+1 represents updating the momentum using the future gradient information,
[0081] (3b) Use the updated momentum g t+1 and the adversarial sample output in the previous round to perform an update calculation to generate a new adversarial sample The formula is as follows:
[0082]
[0083] (3c) Perform the next round of iteration:
[0084] Repeat the above steps 1 - step 3 in the new round of iteration until the total number of iterations T is reached and stop, and output the final adversarial sample.
[0085] Example 2:
[0086] This example proposes an adversarial sample generation system based on future gradient information, and this system includes:
[0087] An input module, used to input the original image, the image classification model (classifier) and related parameters;
[0088] An iteration module, applying the iterative attack method based on momentum, iterating multiple times and reaching the number of future iterations to generate the adversarial samples generated after multiple future iterations;
[0089] A calculation module, used to calculate the future gradient and obtain the sum of the gradients of the adversarial samples for multiple future iterations;
[0090] The generation module uses future gradient to update the momentum, and uses the updated momentum to generate the final adversarial example for this iteration.
[0091] The above input module, iteration module, calculation module and generation module are used to execute the steps of the adversarial example generation method based on future gradient information in Embodiment 1.
[0092] Embodiment 3:
[0093] A computer-readable storage medium, which includes a stored computer program. When the computer program is run by a processor, it controls the device where the storage medium is located to execute the adversarial example generation method based on future gradient information as described above.
[0094] Embodiment 4:
[0095] In this embodiment, a comparative experiment is conducted in combination with the actual application situation to verify the effectiveness of applying the adversarial example generation method based on future gradient information in Embodiment 1 above.
[0096] First, for the dataset, we use 1000 images in the validation set of ILSVRC 2012 (a sub-dataset of ImageNet1K) as the dataset, where each class corresponds to one image. This dataset is widely used in experiments in the field of adversarial attacks.
[0097] For the models, we select 4 classification models based on convolutional neural networks, including Inception-v3 (Inc-v3), VGG16, ResNet50 (Res50) and DenseNet121 (Dense121), and 1 image classification model based on Transformer, namely ViT-Base (ViT-B). In terms of the black-box models selected for evaluation, we first select 3 classification models based on convolutional neural networks, including Inception-v4 (Inc-v4), Inception-Resnet-v2 (Inc-Res-v2) and ConvNeXt-Tiny (ConvNeXt-T), and 2 image classification models based on Transformer, including DeiT-Tiny (DeiT-T) and Swin-Tiny (Swin-T).
[0098] For the comparison methods, we selected MI and NI (an adversarial attack method based on Nesterov accelerated gradient) as the comparison methods. We used the attack success rate to evaluate the transferability and attack ability of adversarial examples. The attack success rate can be calculated by dividing the number of adversarial examples that cause the evaluation model to misclassify by the total number of the dataset, and its value ranges from 0 to 1. The higher the value of the attack success rate, the higher the transferability of the adversarial examples generated by the adversarial attack method and the better the attack performance. The experimental results are shown in Table 1. For the optimal results in the experiment, the font is bolded for easy observation.
[0099] Table 1 Attack success rates (%) of LoFF-MI and baseline methods against pre-trained models
[0100]
[0101] From the above comparison, it can be seen that the present invention improves the existing momentum-based iterative attack method by utilizing the idea of looking ahead. It adopts an inner and outer two-layer iterative manner, with the outer layer iterating T times, and each time the outer layer iterates, the inner layer iterates Q times in the future to obtain future gradient information, overcoming the deficiency that traditional adversarial attack methods fail to fully utilize future gradient information, effectively improving the attack success rate and transferability of adversarial examples, and thus achieving a further attack effect.
[0102] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, should be covered by the protection scope of the present invention.
Claims
1. An adversarial sample generation method based on future gradient information, characterized in that, It includes the following steps: Step 1: Apply the momentum-based iterative attack method, iterate multiple times and reach the future iteration times to generate adversarial examples produced after multiple future iterations; Step 2: Calculate the future gradients and obtain the sum of the gradients of the adversarial examples for multiple future iterations; Step 3: Use the future gradients to update the momentum. After multiple iterations, use the updated momentum to generate the final adversarial example for this iteration.
2. The adversarial sample generation method based on future gradient information according to claim 1, wherein When applying the momentum-based iterative attack method multiple times and reaching the future iteration times to generate adversarial examples produced after multiple future iterations, it includes: Input the original image, the image classification model (classifier) and related parameters: A classifier f with parameters θ and loss function J; an original image x with true label y; maximum perturbation ε and total number of iterations T; number of future iterations Q, momentum decay factor μ; set the perturbation size per step Initial adversarial sample Initial momentum g0 = 0, and clear the future gradient information before each iteration.
3. The adversarial sample generation method based on future gradient information according to claim 2, wherein, The steps of applying the momentum-based iterative attack method multiple times, iterating multiple times and reaching the future iteration times to generate adversarial examples produced after multiple future iterations are as follows: (1a) Set the future iteration times Q and perform parameter initialization; (1b) Apply the momentum-based iterative attack method A multiple times to obtain the adversarial examples produced after multiple future iterations. The formula is as follows: Among them, represents the adversarial sample generated by iterating i rounds into the future at the t-th round, represents the adversarial sample generated by iterating i + 1 rounds into the future at the t-th round; t is the current iteration round, taking values from 0 to T - 1, i is the future iteration round, taking values from 0 to Q - 1; α is the perturbation size for each iteration, sign(·) is the sign function, and (A(x, f(x, θ), J) represents the classifier f generating gradient information for the input sample x using the momentum-based iterative attack method A given the model parameters θ and the loss function J.
4. A method for generating adversarial samples based on future gradient information according to claim 1, wherein The way to calculate the future gradients and obtain the sum of the gradients of the adversarial examples for multiple future iterations is as follows: (2a) Future gradient information G LoFF Obtained by dividing the gradient values of all sample points by the 1-norm and summing them up, the formula is as follows: where y is the correct label corresponding to the original image, ▽ x J(x, y) is the gradient of the loss function with respect to the input image, and its value is calculated by the backpropagation algorithm in the convolutional neural network.
5. A method for generating adversarial samples based on future gradient information according to claim 1, characterized in that Use the future gradients to update the momentum. After multiple iterations, use the updated momentum to generate the final adversarial example for this iteration: (3a) Update the momentum using the future gradient information. The formula is as follows: where μ is the momentum decay factor, and g t is the momentum at the t-th round; (3b) Use the updated momentum g t+1 and the adversarial example output in the previous round to perform an update calculation to generate a new adversarial example The formula is as follows: (3c) Perform the next round of iteration: Repeat the above Steps 1 - 3 in the new round of iteration until the total number of iteration rounds T is reached and stop, and output the final adversarial example.
6. An adversarial sample generation system based on future gradient information, characterized in that, It includes: An input module for inputting the original image, the image classification model (classifier) and related parameters; An iteration module that applies the momentum-based iterative attack method, iterates multiple times and reaches the future iteration times to generate adversarial examples produced after multiple future iterations; A calculation module for calculating the future gradients and obtaining the sum of the gradients of the adversarial examples for multiple future iterations; A generation module that uses the future gradients to update the momentum and uses the updated momentum to generate the final adversarial example for this iteration.
7. The adversarial sample generation system based on future gradient information according to claim 6, wherein The information input by the input module includes: A classifier f with parameters θ and loss function J; an original image x with true label y; maximum perturbation ε and total number of iterations T; number of future iterations Q, momentum decay factor μ; set the perturbation size per step Initial adversarial sample Initial momentum g0 = 0; The iteration module performs multiple iterations in the following way, reaches the future iteration times, and generates adversarial examples produced after multiple future iterations: (1a) Set the future iteration times Q and perform parameter initialization; (1b) Apply the momentum-based iterative attack method A multiple times to obtain the adversarial examples produced after multiple future iterations. The formula is as follows: Among them, represents the adversarial example generated by iterating i rounds into the future at the t-th round, represents the adversarial example generated by iterating i + 1 rounds into the future at the t-th round; t is the current iteration round, taking values from 0 to T - 1, i is the future iteration round, taking values from 0 to Q - 1; α is the perturbation size for each iteration, sign(·) is the sign function, and (A(x, f(x, θ), J) represents the classifier f generating gradient information for the input sample x using the momentum-based iterative attack method A given the model parameters θ and the loss function J.
8. The adversarial sample generation system based on future gradient information according to claim 6, characterized in that, The way the calculation module calculates the future gradients and obtains the sum of the gradients of the adversarial examples for multiple future iterations is as follows: (2a) Future gradient information G LoFF Obtained by dividing the gradient values of all sample points by the 1-norm and summing them up, the formula is as follows: where y is the correct label corresponding to the original image, ▽ x J(x, y) is the gradient of the loss function with respect to the input image, and its value is calculated by the backpropagation algorithm in the convolutional neural network.
9. The adversarial sample generation system based on future gradient information according to claim 6, wherein The generation module uses the future gradients to update the momentum. After multiple iterations, use the updated momentum to generate the final adversarial example for this iteration: (3a) Update the momentum using the future gradient information. The formula is as follows: where μ is the momentum decay factor, and g t is the momentum at the t-th round; (3b) Use the updated momentum g t+1 and the adversarial example output in the previous round to perform an update calculation and generate a new adversarial example The formula is as follows: (3c) Perform the next round of iteration: Repeat Steps 1 - 3 in the new round of iteration until the total number of iteration rounds T is reached and stop, and output the final adversarial example.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program. When the computer program is run by a processor, it controls the device where the storage medium is located to execute the adversarial example generation method based on future gradient information according to any one of claims 1 to 5.