Contrast disturbance attack method based on global gradient iterative optimization
Through the global gradient iterative optimization method, efficient adversarial perturbation for intensive prediction tasks is generated, which solves the problem of untargeted perturbation and limited application scope in the existing technology, and achieves stronger aggressiveness and wider application scope.
Patent Information
- Application Number
- CN202510302621.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, the perturbation generated by the adversarial perturbation attack method for intensive prediction tasks is not targeted, and the scope of application of the attack is limited in the absence of label information.
Adversarial perturbation attack method based on global gradient iteration optimization is adopted, and the global gradient information of the initial adversarial sample is obtained through backpropagation, and new perturbations are generated through the spatial adaptive perturbation generation module, and combined with iterative gradient information is optimized until the number of iterations n is reached.
The generated adversarial perturbation is more targeted, more aggressive, and has an expanded scope of application. At the same time, the loss function is optimized to make the generated perturbation more smooth and effective.
Smart Images

Figure CN120147785A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural networks, and in particular to an adversarial perturbation attack method based on global gradient iterative optimization. Background Art
[0002] Existing mainstream dense prediction tasks, such as scene segmentation, depth estimation, etc., usually use deep neural networks to automatically extract features and achieve counting. However, deep neural networks are composed of multiple layers of non-linear activation functions, and even tiny perturbations may accumulate and amplify between layers, and this superposition effect causes small input changes to trigger significant fluctuations. Therefore, existing mainstream dense prediction tasks based on deep neural networks are vulnerable to adversarial attacks, that is, attackers introduce imperceptible perturbations into the input data, which will significantly affect the counting results.
[0003] Currently, there are few adversarial attack methods for dense prediction tasks, which can be divided into adversarial patch attacks and adversarial perturbation attacks. Adversarial patch attacks are not only easy to detect, but usually only affect specific local regions of the image and have limited impact on the whole. And adversarial perturbation attacks can achieve stronger attack performance by adding global perturbations to the input image. Existing adversarial perturbation attack methods for dense prediction tasks improve the attack performance by weight balancing and weight optimization of the perturbations. However, the perturbations generated by this method may be general and not targeted at specific target models, thus reducing the effectiveness of the attack. Therefore, gradient information can be used to generate more targeted and more aggressive perturbations. However, labeled samples are required in the process of generating perturbations using gradient information, but in real-world scenarios, the labeling information is often unavailable or requires a large cost, which limits the scope of application of the attack. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide an adversarial perturbation attack method based on global gradient iterative optimization. It solves the problems of non-targeted generation of perturbations and limited scope of application of the attack, and optimizes the loss function in the process of generating perturbations, making the generated adversarial perturbations more aggressive.
[0005] To solve the above technical problem, the technical solution adopted by the present invention is: an adversarial perturbation attack method based on global gradient iterative optimization, including the following steps:
[0006] Step 1: Obtain the original image and generate an initial adversarial sample;
[0007] Step 2: Obtain the global gradient information of the initial adversarial sample through backpropagation, and use the global gradient information as the total gradient information;
[0008] Step 3: Generate a new perturbation through the spatial adaptive perturbation generation module using the total gradient information and obtain a new adversarial sample;
[0009] Step 4: Generate global gradient information for the adversarial sample. At the same time, use the perturbation and the previous perturbation to obtain iterative gradient information, and add the two to get the total gradient information;
[0010] Step 5: Repeat Steps 3 - 4 until the number of iterations n is reached to obtain the final perturbation;
[0011] Step 6: Add the original image and the final perturbation to obtain the final adversarial sample and use it to attack the target model.
[0012] A further improvement of the technical solution of the present invention is that the specific steps of Step 1 are as follows:
[0013] Step 1.1: Randomly generate a perturbation P_pre of the same size as the original image and limit the perturbation range; wherein the perturbation is limited to be between -∈ and ∈;
[0014] Step 1.2: Add P_pre to the original image to generate an initial adversarial sample and limit the range of the initial adversarial sample, wherein the adversarial sample is limited to be between 0 and 1.
[0015] A further improvement of the technical solution of the present invention is that the specific steps of Step 2 are as follows:
[0016] Step 2.1: Input the adversarial sample into the target model to obtain a prediction result Y p , and input the prediction result Y p and the relevant label information Z into the global loss function L c to evaluate the gap between the prediction result of the target model and the label information, and obtain the global gradient information Grad_t of the adversarial sample with respect to the perturbation through L c ; the loss function L c uses the mean square error;
[0017] Step 2.2: Input the adversarial sample into a different model M_1 to obtain a prediction result Y 1 P , and input the prediction result Y 1 P and the relevant label information Z into the global loss function L c to evaluate the gap between the prediction result of the target model and the label information, and obtain the global gradient information Grad_1 of the adversarial sample with respect to the perturbation through L c ; the loss function L c uses the mean square error;
[0018] Step 2.3: Input the adversarial sample into a different model M_2 to obtain a prediction result Y 2 P , and input the prediction result Y 2 PInput the global loss function L with the relevant tag information Z c Evaluate the gap between the prediction result of the target model and the tag information through L c Obtain the global gradient information Grad_2 of the adversarial sample with respect to the perturbation; the loss function L c Uses the mean squared error;
[0019] Step 2.4: Utilize the corresponding gradient information Grad_t, Grad_1, and Grad_2 to obtain the global gradient information Grad_global, and take the global gradient information as the total gradient information grad_sum.
[0020] A further improvement of the technical solution of the present invention lies in: The specific steps of step 3 are as follows:
[0021] Step 3.1: Process the total gradient information into an unsigned magnitude matrix and normalize it;
[0022] Step 3.2: Multiply the magnitude matrix by the global parameter base_epsilon to generate a dynamic perturbation intensity matrix epsilon_matrix;
[0023] Step 3.3: Combine the dynamic perturbation intensity matrix with the gradient direction to generate a perturbation P;
[0024] Step 3.4: Add P to the original image to generate an adversarial sample and limit the range of the adversarial sample, where the adversarial sample is restricted between 0 and 1.
[0025] A further improvement of the technical solution of the present invention lies in: The specific steps of step 4 are as follows:
[0026] Step 4.1: Input the adversarial sample into the target model to obtain a prediction result Y p′ , input the prediction result Y p′ and the relevant tag information Z into the global loss function L c Evaluate the gap between the prediction result of the target model and the tag information through L c Obtain the global gradient information Grad_t of the adversarial sample with respect to the perturbation; the loss function L c Uses the mean squared error;
[0027] Step 4.2: Input the adversarial sample into a different model M_1 to obtain a prediction result Y 1 P′ , input the prediction result Y 1 P′ and the relevant tag information Z into the global loss function L c Evaluate the gap between the prediction result of the target model and the tag information through L cObtain the global gradient information Grad_1 of the adversarial sample with respect to the perturbation; the loss function L c Use the mean squared error;
[0028] Step 4.3: Input the adversarial sample into different models M_2 to obtain the prediction result Y 2 P′ , and input the prediction result Y 2 P′ and the relevant label information Z into the global loss function L c Evaluate the gap between the prediction result of the target model and the label information through L c Obtain the global gradient information Grad_2 of the adversarial sample with respect to the perturbation; the loss function L c Use the mean squared error;
[0029] Step 4.4: Use the corresponding gradient information Grad_t, Grad_1, and Grad_2 to obtain the global gradient information Grad_global;
[0030] Step 4.5: Input P_pre and P into the iterative loss function L p , evaluate the gap between the perturbation generated in the previous round and the perturbation generated in this round and obtain the iterative gradient information Grad_iter of the adversarial sample; the loss function L p Use the mean squared error;
[0031] Step 4.6: Add the global gradient information Grad_global and the iterative gradient information Grad_iter as the total gradient information Grad_sum;
[0032] Step 4.7: Update P to P_pre.
[0033] A further improvement of the technical solution of the present invention is that: the label information selection rules in Step 2 and Step 4 are as follows: if the original image has real label information, select the real label information as the label information; if the original image does not have real label information, use the label information replacement strategy to generate pseudo-label information as the label information; the label information replacement strategy rule is: input the original image into the pre-trained target model to obtain the prediction result, and use the prediction result as the label information, where the pre-trained target model refers to the model obtained by preliminary training on the data.
[0034] A further improvement of the technical solution of the present invention is that: the global gradient information Grad_global rule is as follows:
[0035] Grad_global = 0.5 * Grad_t + 0.25 * Grad_1 + 0.25 * Grad_2;
[0036] The total gradient information rule is as follows:
[0037] Grad_sum = Grad_global + λ * Grad_iter;
[0038] Among them, λ represents the hyperparameter that balances between two loss functions.
[0039] Due to the adoption of the above technical solution, the technical progress achieved by the present invention is as follows: By proposing to use the method of iteratively updating perturbations to generate adversarial samples for dense prediction models, applying greater perturbations in key gradient feature regions to improve the attack efficiency, while reducing the noise interference in non-critical regions, taking into account both the adversarial effect and concealment. Among them, the gradient information is obtained by adding three types of model gradient information, reducing the possibility of errors. At the same time, a loss function is designed. Based on calculating the loss function between the prediction result and the relevant label (i.e., the global loss function), a new loss function, namely the iterative process loss function, is added to record the iterative gradient information, which can make the generated perturbations smoother and the attack ability stronger. In addition, considering that the generation of traditional adversarial samples requires label information of images, but since the label information may not be visible to the attacker, the application scope of generating perturbations is limited. Therefore, the present invention develops a pseudo-label generation strategy by replacing annotation information, using the pseudo-label as the label information to generate adversarial samples, so as to expand the scope of the actual application scenarios of the proposed method. Description of the Drawings
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings;
[0041] Figure 1 is a schematic flowchart of the adversarial perturbation attack method based on global gradient iterative optimization of the present invention;
[0042] Figure 2 is an explanatory diagram of the content of the adversarial perturbation attack method based on global gradient iterative optimization of the present invention;
[0043] Figure 3 is a framework diagram of the relevant label selection rule proposed by the present invention;
[0044] Figure 4 is a module diagram of an adversarial perturbation attack system based on global gradient iterative optimization proposed by the present invention. Detailed Embodiment
[0045] The following further describes the present invention in detail with reference to embodiments:
[0046] As shown in Figure 4 the figure, it is a schematic structural diagram of an adversarial perturbation attack system based on global gradient iterative optimization required by the present invention, including an initial adversarial sample generation module, a global gradient information generation module, a spatial adaptive perturbation generation module, a perturbation range limitation module, and an iterative gradient information generation module connected in sequence. The iterative gradient information generation module is connected to the global gradient information generation module, and an attack performance evaluation module is arranged on the perturbation range limitation module. The specific functions of each module are as follows:
[0047] The initial adversarial sample generation module is used to generate an initial adversarial sample; the adversarial sample is obtained by adding an adversarial perturbation to the original image;
[0048] The global gradient information generation module is used to obtain the global gradient information of the adversarial sample with respect to the model by backpropagating the adversarial perturbation;
[0049] The spatial adaptive perturbation generation module is used to generate a spatial adaptive perturbation through the total gradient information;
[0050] The iterative gradient information generation module is used to obtain iterative gradient information; the iterative gradient information is obtained by calculating the loss function between the current perturbation and the previous round of perturbation;
[0051] The perturbation range limitation module is used to limit the generated adversarial perturbation within a specific range;
[0052] The attack performance evaluation module is used to input the final adversarial sample into the target model, record the prediction of the model for the adversarial sample, and compare it with the prediction of the model for the original sample.
[0053] As shown in Figure 1 the figure is a flowchart of an adversarial perturbation attack method based on global gradient iterative optimization implemented by means of the above system. As shown in Figure 2 the figure, taking a monocular depth estimation model as an example, the overall content of the perturbation attack method is described. The specific steps are as follows:
[0054] Step 1: Obtain the original image and generate an initial adversarial sample;
[0055] Step 1.1: Randomly generate a perturbation P_pre of the same size as the original image and limit the perturbation range;
[0056] Step 1.2: Add P_pre to the original image to generate an initial adversarial sample and limit the range of the initial adversarial sample;
[0057] The rules for limiting the perturbation range are as follows:
[0058] The perturbation is restricted to between -∈ and ∈. As ∈ increases, the generated perturbation becomes more aggressive but less concealed. ∈ = 0 indicates no attack.
[0059] The rules for restricting the range of adversarial samples are as follows:
[0060] The adversarial samples are restricted to between 0 and 1. Through normalization, it is easier to control and adjust the scale of the perturbation, so that the generated adversarial samples are kept within the effective range.
[0061] Step 2: Obtain the global gradient information of the initial adversarial sample through backpropagation, and use the global gradient information as the total gradient information;
[0062] Step 2.1: Input the adversarial sample into the target model (crowd counting model A) to obtain the prediction result Y p , and input the prediction result Y p and the relevant label information Z into the global loss function L c Evaluate the gap between the prediction result of the target model and the label information through L c to obtain the global gradient information Grad_t of the adversarial sample with respect to the perturbation; the loss function L c uses the mean squared error;
[0063] Step 2.2: Input the adversarial sample into a different model M_1 (crowd counting model B) to obtain the prediction result Y 1 P , and input the prediction result Y 1 P and the relevant label information Z into the global loss function L c Evaluate the gap between the prediction result of the target model and the label information through L c to obtain the global gradient information Grad_1 of the adversarial sample with respect to the perturbation; the loss function L c uses the mean squared error;
[0064] Step 2.3: Input the adversarial sample into a different model M_2 (crowd counting model C) to obtain the prediction result Y 2 P , and input the prediction result Y 2 P and the relevant label information Z into the global loss function L c Evaluate the gap between the prediction result of the target model and the label information through L c to obtain the global gradient information Grad_2 of the adversarial sample with respect to the perturbation; the loss function L c uses the mean squared error; where the crowd counting models A, B, and C are different crowd counting models, used to prevent a single model from having errors in the gradient.
[0065] Step 2.4: Use the corresponding gradient information Grad_t, Grad_1, and Grad_2 to obtain the global gradient information Grad_global, and use the global gradient information as the total gradient information grad_sum;
[0066] The relevant label information selection rules are as Figure 3 shown:
[0067] If the original image has true label information, select the true label information as the label information. If the original image does not have true label information, use the label information replacement strategy to generate pseudo-label information as the label information; the label information replacement strategy rules are as follows: Input the original image into the pre-trained target model to obtain the prediction result, and use the prediction result as the label information, where the pre-trained target model refers to the model obtained by preliminary training on the data.
[0068] The rules for the global gradient information Grad_sum are as follows:
[0069] Grad_global = 0.5 * Grad_t + 0.25 * Grad_1 + 0.25 * Grad_2;
[0070] Step 3: Generate a new perturbation through the spatial adaptive perturbation generation module using the total gradient information and obtain a new adversarial sample;
[0071] Step 3.1: Process the total gradient information into an unsigned magnitude matrix and normalize it;
[0072] Step 3.2: Multiply the magnitude matrix by the global parameter base_epsilon to generate a dynamic perturbation intensity matrix epsilon_matrix;
[0073] Step 3.3: Combine the dynamic perturbation intensity matrix with the gradient direction to generate a perturbation P;
[0074] Step 3.4: Add P to the original image to generate an adversarial sample and limit the range of the adversarial sample, where the adversarial sample is limited between 0 and 1;
[0075] Step 4: Generate global gradient information from the adversarial sample, and at the same time use the perturbation and the previous round of perturbation to obtain iterative gradient information, and add the two to obtain the total gradient information;
[0076] Step 4.1: Input the adversarial sample into the target model to obtain the prediction result Y p′ , and input the prediction result Y p′ and the relevant label information Z into the global loss function L c to evaluate the gap between the prediction result of the target model and the label information, and through L cObtain the global gradient information Grad_t of the adversarial sample with respect to the perturbation; the loss function L c Use the mean squared error;
[0077] Step 4.2: Input the adversarial sample into different models M_1 to obtain the prediction result Y 1 P′ , and input the prediction result Y 1 P′ and the relevant label information Z into the global loss function L c Evaluate the gap between the prediction result of the target model and the label information through L c Obtain the global gradient information Grad_1 of the adversarial sample with respect to the perturbation; the loss function L c Use the mean squared error;
[0078] Step 4.3: Input the adversarial sample into different models M_2 to obtain the prediction result Y 2 P′ , and input the prediction result Y 2 P′ and the relevant label information Z into the global loss function L c Evaluate the gap between the prediction result of the target model and the label information through L c Obtain the global gradient information Grad_2 of the adversarial sample with respect to the perturbation; the loss function L c Use the mean squared error;
[0079] Step 4.4: Utilize the corresponding gradient information Grad_t, Grad_1, and Grad_2 to obtain the global gradient information Grad_global;
[0080] Step 4.5: Input P_pre and P into the iterative loss function L p , evaluate the gap between the perturbation generated in the previous round and the perturbation generated in this round and obtain the iterative gradient information Grad_iter of the adversarial sample; the loss function L p Use the mean squared error;
[0081] Step 4.6: Add the global gradient information Grad_global and the iterative gradient information Grad_iter as the total gradient information Grad_sum;
[0082] Step 4.7: Update P to P_pre;
[0083] The selection rule for the relevant label information is as Figure 3 shown:
[0084] If the original image has true label information, select the true label information as the label information. If the original image does not have true label information, use the label information replacement strategy to generate pseudo-label information as the label information;
[0085] The rules of the label information replacement strategy are as follows:
[0086] Input the original image into the pre-trained target model to obtain the prediction result, and use the prediction result as the label information, where the pre-trained target model refers to the model obtained by preliminary training on the data;
[0087] The rules of the global gradient information Grad_global are as follows:
[0088] Grad_global = 0.5 * Grad_t + 0.25 * Grad_1 + 0.25 * Grad_2;
[0089] The rules of the total gradient information are as follows:
[0090] Grad_sum = Grad_global + λ * Grad_iter;
[0091] Among them, λ represents the hyperparameter for balancing the two loss functions;
[0092] Step 5: Repeat Step 3 - 4 until the iteration number n is reached to obtain the final perturbation. In this embodiment, n is specifically taken as 20;
[0093] Step 6: Add the original image and the final perturbation to obtain the final adversarial sample and use it to attack the target model.
[0094] The above-described embodiments are merely descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. An adversarial perturbation attack method based on global gradient iterative optimization, characterized in that: The steps include: Step 1: Get the original image and generate the initial adversarial sample; Step 2: Obtain the global gradient information of the initial adversarial sample through back propagation, and use the global gradient information as the total gradient information; Step 3: The total gradient information is used to generate new perturbations through the spatial adaptive perturbation generation module and obtain new adversarial samples; Step 4: Generate global gradient information from the adversarial sample, and use the perturbation and the previous perturbation to obtain iterative gradient information, and add the two to get the total gradient information; Step 5: Repeat steps 3-4 until the number of iterations n is reached to obtain the final perturbation; Step 6: Add the original image and the final perturbation to obtain the final adversarial sample and use it to attack the target model.
2. The adversarial perturbation attack method based on global gradient iterative optimization according to claim 1, characterized in that: Step 1 The specific steps are as follows: Step 1.1: Randomly generate a perturbation P_pre of the same size as the original image, limiting the perturbation range; the perturbation is limited to between -∈ and ∈; Step 1.2: Add P_pre to the original image to generate an initial adversarial sample and limit the range of the initial adversarial sample, where the adversarial sample is limited to between 0 and 1.
3. The adversarial perturbation attack method based on global gradient iterative optimization according to claim 1, characterized in that: Step 2 The specific steps are as follows: Step 2.1: Input the adversarial sample into the target model to obtain the prediction result Y p , the predicted result Y p And the relevant label information Z is input into the global loss function L c Evaluate the gap between the target model’s prediction results and label information, using L c Get the global gradient information Grad_t of the adversarial sample to the disturbance; the loss function L c Use mean square error; Step 2.2: Input the adversarial sample into different models M_1 to obtain the prediction result Y1 P , the predicted result Y1 P And the relevant label information Z is input into the global loss function L c Evaluate the gap between the target model’s prediction results and label information, using L c Get the global gradient information Grad_1 of the adversarial sample to the disturbance; the loss function L c Use mean square error; Step 2.3: Input the adversarial sample into different models M_2 to obtain the prediction result Y2 P , the predicted result Y2 P And the relevant label information Z is input into the global loss function L c Evaluate the gap between the target model’s prediction results and label information, using L c Get the global gradient information Grad_2 of the adversarial sample to the disturbance; the loss function L c Use mean square error; Step 2.4: Use the corresponding gradient information Grad_t, Grad_1 and Grad_2 to obtain the global gradient information Grad_global, and use the global gradient information as the total gradient information grad_sum.
4. The adversarial perturbation attack method based on global gradient iterative optimization according to claim 1, characterized in that: Step 3 The specific steps are as follows: Step 3.1: Process the total gradient information into an unsigned magnitude matrix and normalize it; Step 3.2: Multiply the amplitude matrix by the global parameter base_epsilon to generate the dynamic perturbation intensity matrix epsilon_matrix; Step 3.3: Combine the dynamic perturbation intensity matrix with the gradient direction to generate the perturbation P; Step 3.4: Add P to the original image to generate adversarial samples and limit the range of adversarial samples, where the adversarial samples are limited to between 0 and 1.
5. The adversarial perturbation attack method based on global gradient iterative optimization according to claim 1, characterized in that: Step 4 The specific steps are as follows: Step 4.1: Input the adversarial sample into the target model to obtain the prediction result Y p′ , the predicted result Y p′ And the relevant label information Z is input into the global loss function L c Evaluate the gap between the target model’s prediction results and label information, using L c Get the global gradient information Grad_t of the adversarial sample to the disturbance; the loss function L c Use mean square error; Step 4.2: Input the adversarial sample into different models M_1 to obtain the prediction result Y1 P′ , the predicted result Y1 P′ And the relevant label information Z is input into the global loss function L c Evaluate the gap between the target model’s prediction results and label information, using L c Get the global gradient information Grad_1 of the adversarial sample to the disturbance; the loss function L c Use mean square error; Step 4.3: Input the adversarial sample into different models M_2 to get the prediction results The predicted results And the relevant label information Z is input into the global loss function L c Evaluate the gap between the target model’s prediction results and label information, using L c Get the global gradient information Grad_2 of the adversarial sample to the disturbance; the loss function L c Use mean square error; Step 4.4: Use the corresponding gradient information Grad_t, Grad_1 and Grad_2 to obtain the global gradient information Grad_global; Step 4.5: Input P_pre and P into the iterative loss function L p , evaluate the gap between the perturbation generated in the previous round and the perturbation generated in this round and obtain the iterative gradient information Grad_iter of the adversarial sample; the loss function L p Use mean square error; Step 4.6: Add the global gradient information Grad_global and the iterative gradient information Grad_iter as the total gradient information Grad_sum; Step 4.7: Update P to P_pre.
6. The adversarial perturbation attack method based on global gradient iterative optimization according to claim 3 or 5, characterized in that: The label information selection rule in step 2 and step 4 is: if the original image has real label information, the real label information is selected as the label information; if the original image does not have real label information, the label information replacement strategy is used to generate pseudo label information as the label information; the label information replacement strategy rule is: input the original image into the pre-trained target model to obtain the prediction result, and use the prediction result as the label information, where the pre-trained target model refers to the model obtained by preliminary training on the data.
7. The adversarial perturbation attack method based on global gradient iterative optimization according to claim 6, characterized in that: The global gradient information Grad_global rules are as follows: Grad_global=0.5*Grad_t+0.25*Grad_1+0.25*Grad_2; The total gradient information rules are as follows: Grad_sum=Grad_global+λ*Grad_iter; Among them, λ represents the hyperparameter that balances the two loss functions.