Anti-attack method facing unknown transformation robustness

By using ERAA and FRAA methods in the process of adversarial sample generation, adversarial samples that are still valid after unknown transformation are generated, the problem of adversarial samples failing after unknown transformation in the prior art, and the robustness and wide applicability of adversarial samples are achieved.

CN119990248APending Publication Date: 2025-05-13HARBIN INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510054583.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing adversarial samples cannot effectively act on the target neural network after unknown transformations, and possible interference before input cannot be predicted during the generation of adversarial samples.

Method used

A robust adversarial attack method in the face of unknown transformations is proposed. The original sample data is processed through ERAA and FRAA methods to generate adversarial samples that are robust in both European and feature spaces. ERAA calculates the update direction through Gaussian sampling and cross-entropy loss in Euro-style space, and FRAA learns natural distribution through generators and discriminators in feature space to generate universal adversarial samples.

Benefits of technology

The robustness of adversarial samples after unknown transformation is achieved, the effectiveness can be maintained after unknown transformation, and it does not depend on a specific deep learning model, and can work on unknown models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990248A_ABST
    Figure CN119990248A_ABST
Patent Text Reader

Abstract

The invention provides a robust attack resisting method facing unknown transformation, and belongs to the technical field of computer vision. The method solves the problem that in the process of generating an adversarial sample, which interference will be experienced by the adversarial sample before the adversarial sample is input into a target neural network cannot be known. The method specifically comprises the following steps: step 1, preprocessing an input data set of a target deep neural network to be attacked to obtain an input sample; 2, dividing an input sample into original sample data x and an original sample data set S; 3, an ERAA method is used for processing the original sample data x to obtain a neighborhood robust adversarial sample x1, an FRAA method is used for processing the original sample data set S to obtain a neighborhood robust adversarial sample x2, generalization robustness of unknown transformation operation is achieved, ERAA is transformation robust adversarial attack in an Euclidean space, and FRAA is a transformation robust adversarial attack in an Euclidean space. And the FRAA is a robust adversarial attack which is pervasive to transformation in the feature space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a robust counterattack method facing unknown transformations, and belongs to the technical field of computer vision. Background Art

[0002] With the application of artificial intelligence in various fields, the exploitation of artificial intelligence vulnerabilities has become an issue worthy of in-depth study. The deep neural network of artificial intelligence statistically learns the optimal parameters for the task from a large amount of data. A large number of studies have shown that these neural networks are susceptible to adversarial samples. However, most existing adversarial samples need to be directly input into the target neural network to affect it, and when the adversarial samples are subject to certain interference, these adversarial samples will not work on the target neural network. In actual application scenarios, adversarial samples usually need to be transmitted in the network or resampled in the physical world. These actual scenarios will interfere with the adversarial samples, and these interferences are unknown in advance, that is, in the process of generating adversarial samples, it is impossible to know what kind of interference the adversarial samples will experience before entering the target neural network. The present invention proposes an adversarial attack method that is robust to unknown transformations, which can maintain the effectiveness of adversarial samples after unknown transformations. The present invention explores the formation mechanism of adversarial samples, which provides a research basis for the development of reliable artificial intelligence. Summary of the invention

[0003] In order to solve the problem that it is impossible to know what kind of interference the adversarial sample will experience before being input into the target neural network during the generation of the adversarial sample, the present invention proposes a robust adversarial attack method facing unknown transformations, which specifically includes:

[0004] Step 1: Preprocess the input data set of the target deep neural network to be attacked to obtain input samples;

[0005] Step 2: Divide the input sample into original sample data x and original sample data set S;

[0006] Step 3: Use the ERAA method to process the original sample data x to obtain a neighborhood-robust adversarial sample Use the FRAA method to process the original sample data set S to obtain neighborhood robust adversarial samples The generalized robustness of unknown transformation operations is achieved, where ERAA is a transformation robust adversarial attack in Euclidean space, and FRAA is a transformation universal robust adversarial attack in feature space.

[0007] Optionally, the classifier of the target deep neural network to be attacked in step 1 is based on the ResNet-34 model. The ResNet-34 model consists of r1 convolutional layers, r2 average pooling layers and r3 fully connected layers. r1, r2, and r3 are 34, 1, and 1 respectively. The convolution kernel size of the first convolutional layer is 7×7. The remaining 33 convolutional layers are divided into four layers. The number of residual blocks in each layer is 3, 4, 6, and 3 respectively. Each residual block contains 2 convolutional layers. The convolution kernel size of each convolutional layer is 3×3, the window size of the average pooling layer is 3×3, and the step size is 2.

[0008] Optionally, the neighborhood-robust adversarial sample obtained in step 3 The steps include:

[0009] Step 3.1.1: Get the initial adversarial sample point based on the original sample data x, and initialize the adversarial perturbation x of the initial adversarial sample point adv and the average update direction

[0010] Step 3.1.2: Collect initial adversarial sample points Gaussian sampling points on the θ neighborhood of ;

[0011] Step 3.1.3: Calculate each Gaussian sampling point The cross entropy loss L at θ ;

[0012] Step 3.1.4: Based on cross entropy loss L θ Calculate update direction And superimpose the calculation to get the average update direction of the sampling points

[0013] Step 3.1.5: Update the direction using the average Update the adversarial perturbation x adv , get the average update direction away from the classification at the sampling point;

[0014] Step 3.1.6: Constrain the adversarial perturbation strength x adv ∈[-∈ δ ,∈ δ ], output neighborhood-robust adversarial examples Among them, ∈ δ budget for disturbances;

[0015] Cross entropy loss L θ The calculation formula is:

[0016]

[0017] In formula (1), target is the target class label, model is the deep learning model, and θ is the i-th Gaussian sampling point;

[0018] The expression for updating the anti-perturbation is:

[0019]

[0020] In formula (2), n θ is the number of sampling points, lr is the learning rate, and sign is the sign function.

[0021] Optionally, in step 3.1.4, the cross entropy loss L θ Calculate update direction Specifically include:

[0022] Using cross entropy loss L θ Calculate the gradient to update the sampling point θ new , the direction of the Gaussian sampling point towards the gradient ascent is used as the update direction

[0023] Gradient update sampling point θ new The expression is:

[0024] θ new ←η·θ+μ·sign(▽ θ L θ )(3);

[0025] In formula (3), η and μ are equilibrium parameters.

[0026] Optionally, adversarial examples in step 3 The steps to obtain include:

[0027] Step 3.2.1: Obtain the initial adversarial sample point based on the original sample data set S and initialize the adversarial perturbation x of the initial adversarial sample point adv and the average update direction

[0028] Step 3.2.2: Build the generator G, take z, z~N(0,1) randomly sampled from the Gaussian distribution as the input of the generator, and get the output y;

[0029] Step 3.2.3: Build the discriminator D and randomly sample samples x from the dataset a , will output y and random sample x a Compare the composition data to calculate the resolution loss L of the data pair, and use the resolution loss to train the discriminator D;

[0030] Step 3.2.4: Use the classification loss of the target model C to perform gradient descent on z and adjust z to search for the most effective perturbation to attack the target model;

[0031] Step 3.2.5: Constrain the strength of the adversarial disturbance Output Neighborhood Robust Adversarial Examples

[0032] The beneficial effects of the present invention are:

[0033] (1) The present invention does not rely on training on known specific transformations and can generate robust adversarial samples even in the face of unknown transformations.

[0034] (2) The anti-interference ability of the adversarial samples generated by the present invention depends on the real data distribution in the feature space. Its robustness is not limited to a specific deep learning model and can be effective for unknown models. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A schematic diagram of a flow chart of a robust anti-attack method against unknown transformations provided by the present invention;

[0036] Figure 2 A flowchart of the transformation robust counterattack algorithm in Euclidean space provided by the present invention;

[0037] Figure 3 The flowchart of the robust counter-attack algorithm that is universally applicable to transformations in the feature space of the present invention system. DETAILED DESCRIPTION

[0038] Specific implementation method 1: Combination Figure 1 To illustrate this embodiment, existing robust attack methods, such as EOT, AI-PGD, etc., generate corresponding adversarial samples for specific transformations. Although these methods work for one or several known transformations, the effectiveness of their adversarial samples is greatly reduced under unknown transformations. These unknown transformations may be adding some noise, changing the contrast of the image, or performing affine operations such as rotation and scaling on the image. Figure 1 As shown, after these transformations, the adversarial samples may have undergone a small migration in the Euclidean space (such transformations are called Euclidean transformations in this embodiment) or a large migration (such transformations are called non-Euclidean transformations in this embodiment).

[0039] This embodiment first considers the robustness of adversarial samples under Euclidean transformation, that is, to make the adversarial samples maintain their effectiveness within the ∈-neighborhood in Euclidean space. To this end, this embodiment proposes a transformation robust adversarial attack (ERAA) in Euclidean space, and the steps include:

[0040] S1: Collect initial adversarial sample points Gaussian sampling points in the θ neighborhood of .

[0041] By sampling in the neighborhood of the initial adversarial sample point, the new adversarial perturbation calculated after sampling can better target the entire category of the initial classification rather than a specific sample, while avoiding the saddle point in the high-dimensional space that affects the iteration direction of the adversarial sample.

[0042] S2: Calculate each sampling point The cross entropy loss L at θ .

[0043] The cross entropy loss is used to calculate the gradient update sampling points. These points are points in the gradient direction that are more easily misclassified in the neighborhood. Updating the sampling points in the direction of the gradient increase can obtain an update direction with better effect under targeted attack.

[0044] S3: Use the gradient information of the new sampling point to calculate the update direction And superimpose and calculate the average update direction of the sampling points

[0045] use Update the adversarial perturbation x adv , and obtain the average update direction away from the classification, so as to obtain the robust adversarial perturbation for this category.

[0046] The ERAA method is based on the assumption that the disturbance caused by the transformation will not cause excessive deviation of the Euclidean distance in the sample space, that is, the transformed samples are still in the neighborhood of the original category. However, for the transformation of non-additive noise, it is meaningless to calculate the distance in the sample space because the disturbance caused by the transformation will cause a large deviation of the Euclidean distance in the sample space. Therefore, this method needs to be supplemented by the feature space method to achieve generalized robustness to unknown transformation operations.

[0047] In order to improve the robustness of adversarial samples under non-Euclidean transformations, this paper proposes a robust adversarial attack method against unknown transformations. Previous methods have focused on the study of affine transformations, which require information on various known transformations for data enhancement. Therefore, they lack generalization capabilities and cannot be well applied to the physical world.

[0048] The core problem of this method is how to discover a semantic feature space, which is defined in this embodiment as: a feature space in which all kinds of transformations of the original sample are clustered in their neighborhood because they have the same semantics. In this space, the ERAA method proposed in this article can be used to similarly find a flat local optimal value in the feature space to achieve the general robustness requirement. To this end, this embodiment first makes the following conjecture: All natural samples obey the natural distribution P, and the original samples still obey P under various transformations. The target model can fit the distribution P, so that the natural samples can still be recognized by the target model under various transformations, and the existing adversarial samples cannot be recognized by the target model after transformation because they are separated from P. Therefore, the FRAA method uses a generative adversarial network to learn the natural distribution P, and then searches for effective adversarial samples in P. The steps of the FRAA method include:

[0049] S4: First build a generator G;

[0050] The function of the generator is to project the random seed z that obeys the low-dimensional Gaussian distribution into the natural distribution P, and use the randomly sampled z from the Gaussian distribution as the input of the generator to obtain the output y;

[0051] S5: Build a discriminator D to assist in training the generator G

[0052] The role of the discriminator is to distinguish whether the sample comes from a natural distribution and randomly sample samples x from the dataset. a To form a data pair in y, first calculate the discrimination loss L of D for the data pair. Then use L to train the discriminator D.

[0053] S6: Query the target model to optimize the random seed z

[0054] Although random z can be mapped to high-likelihood samples in P under the action of G, these samples may not be identified as artificially set target categories by the target model C with high confidence. Therefore, it is necessary to adjust z to search for the most effective perturbation to attack the target model. Using the classification loss of C to perform gradient descent on z can simultaneously ensure that 1. the generated samples follow the natural distribution and 2. mislead the target model C with high confidence.

[0055] Specific implementation method 2: Combination Figure 2 and Figure 3 This embodiment is described. This embodiment tests the anti-attack method of a specific embodiment. Figure 2 and Figure 3 As shown, the test process includes:

[0056] S1: Preprocess the input data set of the target deep neural network to be attacked to obtain input samples.

[0057] S2: Divide the input data set into a training set and a test set, and use the test set to test the target system. In this test, the input data set uses the Cifar-10 data set, and the size of each image in the data set is 32×32×3. The classifier of the target system to be attacked is based on the ResNet-34 model. ResNet-34 consists of r1 convolutional layers, r2 average pooling layers, and r3 fully connected layers. The sizes of r1, r2, and r3 are 34, 1, and 1 respectively. The convolution kernel size of the first convolution layer is 7×7. The remaining 33 convolution layers are divided into four layers. The number of residual blocks in each layer is 3, 4, 6, and 3 respectively. Each residual block contains 2 convolutional layers, and the convolution kernel size of each convolution layer is 3×3. The window size of the average pooling layer is 3×3, and the stride is 2.

[0058] S3: Generate adversarial attack labels using the test set, and use the initial labels l of the samples in the dataset for untargeted attacks u , a targeted attack generates a label l of a different category from the initial label t In this test, the targeted attack labels are generated based on the initial labels. This classification task has 10 initial categories, and the targeted attack labels are generated based on the initial labels. t It can be expressed as:

[0059] l t =(l u +1)mod 10 (1);

[0060] S4: Use multiple adversarial attack methods: PGD, MIFGSM, DTA, GRA, SMIFGRM, PGN gradient-based white-box attack methods in the Transfer Attack adversarial robustness toolbox, DIM, TIM, BSR and other adversarial attack methods based on input transformation, as well as the ERAA and FRAA methods proposed in this invention, to generate adversarial samples under multiple sets of parameters. In this test, the strength constraints of the above attack methods are all based on the L2 norm, and the maximum number of iterations is set to β. In this implementation, β is 200. The above attack methods are applied to the test set under different parameters to generate adversarial samples.

[0061] Among them, black box attack and white box attack are as follows:

[0062] S401: White-box attacks have full access to the target model. One of the most important white-box attacks is gradient-based. FGSM is a common gradient-based attack algorithm that proves that the linear features of deep neural networks in high-dimensional space are sufficient to generate adversarial samples. It performs a one-step update:

[0063]

[0064] In formula (2), is the gradient of the loss function with respect to x, ∈ is the threshold of the adversarial perturbation, and sign is the sign function. PGD extends FGSM to an iterative version. It iteratively applies small-step gradient updates multiple times and clips the adversarial examples at the end of each step:

[0065]

[0066] In formula (3), Clip ∈ is a projection operation; B ∈ (x) is L with x as the center and radius ∈ p ball; α is the step size.

[0067] Optimization-based attacks aim to generate adversarial examples with minimal perturbations. Deepfool is an iterative attack method based on the idea of ​​hyperplane classification. In each iteration, the algorithm adds small perturbations to the image, gradually making the image cross the classification boundary until the image is misclassified. The final perturbation is the accumulation of perturbations in each iteration. Carlini & Wagner's method (C&W) is a powerful optimization-based method that adopts Lagrangian form and uses Adam for optimization, which can be written as:

[0068]

[0069] C&W is a very effective white-box attack method, but it lacks transferability to black-box models.

[0070] S402: Black Box Attack

[0071] Black-box attacks cannot access the parameters and gradients of the target model, and can generally be divided into transfer-based, score-based, and decision-based attacks. Transfer-based attacks use the source model to generate adversarial samples, and then use the adversarial transferability of the adversarial samples to transfer them to the target model. MIM improves transferability by adding a momentum term when generating adversarial samples. DIM proposes to improve the transferability of adversarial samples by increasing the diversity of inputs, applying random resizing and padding to the inputs with a given probability in each attack iteration, and inputting the outputs into the network for gradient calculation. In order to further improve the transferability on some defense models, Dong et al. proposed a translation-invariant attack (TI). This method reduces the computational complexity by convolving the untranslated gradient map with a predefined kernel.

[0072] Score-based attacks only have access to the output scores of the target model for each input. Attacks in this setting estimate the gradient of the target model using a gradient-free method through a set of queries. NES and SPSA use sampling methods to fully approximate the true gradient. Prior-guided Random Gradient-free (P-RGF) improves the accuracy of estimated gradients through transfer-based priors. NATTACK learns a probability density distribution centered on the input and samples from the distribution to generate adversarial examples.

[0073] Decision-based attacks are more challenging because the attacker can only obtain discrete hard-label predictions of the target model. Decision-based attacks, such as Boundary attacks and Evolutionary attacks, also play an important role in black-box attacks.

[0074] S5: After different transformations, the adversarial samples are input into the target system for classification to detect the effect of the adversarial attack. In this test, the transformations added to the adversarial samples can be divided into two categories. The first category is additive noise on the neighborhood, which includes brightness, contrast, Gaussian noise, and JPEG compression methods in this test, and the intensity parameters are set to [α1, α2, α3, α4]. The second category is non-additive noise represented by rotation changes, which includes scaling, rotation, random perspective, and mixed transformations based on affine transformations in this test. The intensity parameters are set to [γ1, γ2, γ3, γ4], where α1, α2, α3, and α4 are 0.5, 0.5, (convolution kernel size = 5, standard deviation σ = 1.0), and 0.5, respectively, and γ1, γ2, γ3, and γ4 are 1.25, 20°, (perspective intensity = 0.5, perspective probability p = 1.0), and (rotation angle = 5°, translation transformation ratio = 0.02, scaling factor = 0.75).

[0075] S6: After different transformations, adversarial samples need to be resized to meet the input requirements of the target model and reduce test errors.

[0076] S7: For Figure 3 The effectiveness of the different classification tasks mentioned in step S12 in the above test is tested respectively. In this test, the effectiveness of the untargeted attack is to distinguish the model classification result from the original label l u The accuracy (acc) is used to reflect the effectiveness of non-targeted attacks, and the change in accuracy before and after transformation reflects the robustness of adversarial samples. The effectiveness of targeted attacks is measured by the model classifying the results of adversarial samples after transformation operations as equivalent to the target category label l. t The attack success rate (ASR) is used to reflect the effectiveness of the targeted attack, and the change in the attack success rate after transformation reflects the robustness of the adversarial sample.

[0077] S8: The classification task results are obtained using the proposed test method and indicators. In this test, the attack target model used is ResNet-34, as shown in Tables 1, 2 and 3. Through experimental evaluation and model comparison, the adversarial sample robustness method is compared with two major categories of 9 adversarial attack methods. The robustness of the present invention reaches or even exceeds the performance of classical methods and the most advanced adversarial attacks. The present invention does not rely on training for known specific transformations, and can generate robust adversarial samples in the face of unknown transformations. The anti-interference ability of the adversarial samples generated by the present invention depends on the real data distribution in the feature space. Its robustness is not limited to a specific deep learning model and can be effective for unknown models.

[0078] Table 1

[0079]

[0080]

[0081] Table 2

[0082]

[0083] Table 3

[0084]

[0085] The above is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technician familiar with this profession can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement made to the above embodiments without departing from the content of the technical solution of the present invention, based on the technical essence of the present invention, within the spirit and principles of the present invention, still fall within the protection scope of the technical solution of the present invention.

Claims

1. A robust adversarial attack method against unknown transformations, characterized in that: The steps of the method for robustly confronting unknown transformations include: Step 1: Preprocess the input data set of the target deep neural network to be attacked to obtain input samples; Step 2: Divide the input sample into original sample data x and original sample data set S; Step 3: Use the ERAA method to process the original sample data x to obtain a neighborhood-robust adversarial sample Use the FRAA method to process the original sample data set S to obtain neighborhood robust adversarial samples The generalized robustness of unknown transformation operations is achieved, where ERAA is a transformation robust adversarial attack in Euclidean space, and FRAA is a transformation universal robust adversarial attack in feature space.

2. The method for robustly resisting attacks against unknown transformations according to claim 1, characterized in that: The classifier of the target deep neural network to be attacked in step 1 is based on the ResNet-34 model. The ResNet-34 model consists of r1 convolutional layers, r2 average pooling layers and r3 fully connected layers. r1, r2, and r3 are 34, 1, and 1 respectively. The convolution kernel size of the first convolutional layer is 7×7. The remaining 33 convolutional layers are divided into four layers. The number of residual blocks in each layer is 3, 4, 6, and 3 respectively. Each residual block contains 2 convolutional layers. The convolution kernel size of each convolutional layer is 3×3, the window size of the average pooling layer is 3×3, and the step size is 2.

3. The method for robustly resisting attacks against unknown transformations according to claim 1, characterized in that: In step 3, we get a neighborhood-robust adversarial sample. The steps include: Step 3.1.1: Get the initial adversarial sample point based on the original sample data x, and initialize the adversarial perturbation x of the initial adversarial sample point adv and the average update direction Step 3.1.2: Collect initial adversarial sample points Gaussian sampling points on the θ neighborhood of ; Step 3.1.3: Calculate each Gaussian sampling point The cross entropy loss L at θ ; Step 3.1.4: Based on cross entropy loss L θ Calculate update direction And superimpose the calculation to get the average update direction of the sampling points Step 3.1.5: Update the direction using the average Update the adversarial perturbation x adv , get the average update direction away from the classification at the sampling point; Step 3.1.6: Constrain the adversarial perturbation strength x adv ∈[-∈ δ ,∈ δ ], output neighborhood-robust adversarial examples Among them, ∈ δ budget for disturbances; Cross entropy loss L θ The calculation formula is: In formula (1), target is the target class label, model is the deep learning model, and θ is the i-th Gaussian sampling point; The expression for updating the anti-perturbation is: In formula (2), n θ is the number of sampling points, lr is the learning rate, and sign is the sign function.

4. The method for robustly resisting attacks against unknown transformations according to claim 3, characterized in that: In step 3.1.4, based on the cross entropy loss L θ Calculate update direction Specifically include: Using cross entropy loss L θ Calculate the gradient to update the sampling point θ new , the direction of the Gaussian sampling point towards the gradient rise is used as the update direction Gradient update sampling point θ new The expression is: In formula (3), η and μ are equilibrium parameters.

5. The method for robustly resisting attacks against unknown transformations according to claim 1, characterized in that: Adversarial examples in step 3 The steps to obtain include: Step 3.2.1: Obtain the initial adversarial sample point based on the original sample data set S and initialize the adversarial perturbation x of the initial adversarial sample point adv and the average update direction Step 3.2.2: Build the generator G, take z, z~N(0,1) randomly sampled from the Gaussian distribution as the input of the generator, and get the output y; Step 3.2.3: Build the discriminator D and randomly sample samples x from the dataset a , will output y and random sample x a Compare the composition data to calculate the resolution loss L of the data pair, and use the resolution loss to train the discriminator D; Step 3.2.4: Use the classification loss of the target model C to perform gradient descent on z and adjust z to search for the most effective perturbation to attack the target model; Step 3.2.5: Constrain the strength of the adversarial disturbance Output Neighborhood Robust Adversarial Examples