A transferable adversarial attack method based on generative adversarial networks
Patent Information
- Application Number
- CN202311328770.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-14
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-10-14
AI Technical Summary
它能够提高攻击的效果和泛化能力,但是容易过拟合于原始样本的空间流形导致迁移性变差和生成对抗样本的图像质量较低,通常缺乏逼真性和细节清晰度
[0010]相比于现有技术,本发明具有如下有益效果。
Smart Images

Figure CN117313107B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of image processing, specifically relating to a transferable adversarial attack method based on generative adversarial networks. Background Technology
[0002] Transferable adversarial attacks refer to the successful attack on one or more different neural network models by generating adversarial examples on one neural network. The attack uses a generative network to make small, targeted perturbations to the input data, causing the target neural network to make incorrect predictions or misclassifications. Transferable adversarial attacks are used to evaluate model robustness, advance adversarial security research, improve model security, and provide guidance for defending against adversarial attacks. They help reveal potential weaknesses in deep learning models, thereby promoting model development and improving model security.
[0003] The key to traditional transferable adversarial attack methods lies in generating adversarial examples and verifying their transferability across different models. Attackers analyze and improve attack methods, exploiting vulnerabilities or shared features in the source model to successfully attack other models, revealing model flaws and driving model improvement and enhancement. However, traditional transferable adversarial attack methods suffer from drawbacks such as insufficient diversity in generated examples, dependence on model structure, inefficiency, and high computational cost. These issues limit the generalization ability and practicality of the attack methods.
[0004] Attackers can generate adversarial examples by subtly perturbing the input. Using adversarial networks to generate adversarial examples, these examples closely resemble the original clean examples but can still fool deep neural networks, leading to incorrect judgments. The emergence of adversarial examples exacerbates the security issues of artificial intelligence, especially in sensitive areas involving personal and property safety, such as autonomous driving and facial recognition payment.
[0005] Transferable adversarial attacks based on generative adversarial networks (GANs) offer advantages such as higher quality generated samples, adaptive generation capabilities, wide applicability, and diverse attack strategies. While they can improve attack effectiveness and generalization ability, they are prone to overfitting to the spatial manifold of the original samples, leading to poor transferability and lower image quality in the generated adversarial examples, often lacking realism and detail. Summary of the Invention
[0006] This invention aims to address the shortcomings of existing technologies in adversarial example attacks. To this end, it employs a Generative Adversarial Network (GAN) as its basic framework and introduces an auxiliary network, the Auxiliary Generator, into the generator to improve the quality and diversity of generated adversarial examples. By introducing the Auxiliary Generator, this method aims to enhance the generator's performance and effectiveness. The Auxiliary Generator works in conjunction with the main generator to jointly generate adversarial examples. The introduction of the Auxiliary Generator increases the generator's generation capability, improving the realism and diversity of generated examples. By incorporating the Auxiliary Generator, this method can better capture data distribution and generate more deceptive and misleading adversarial examples. Furthermore, this invention adds a residual block model to enhance the network's feature extraction capabilities. The residual block model utilizes residual connections, which helps mitigate the gradient vanishing problem during training and improves the network's nonlinear modeling ability. By enhancing the network's feature extraction capabilities, it can better capture the correlation between the input image and auxiliary information, thereby generating more aggressive and deceptive adversarial examples. In addition, this invention introduces a pyramid segmentation attention mechanism in the bottleneck layer. This mechanism, through the use of pyramid segmentation attention, allows the generator to focus on different features at different scales. This enables it to better capture important regions and features in the image, improving the quality of generated samples and the effectiveness of attacks.
[0007] To achieve the above effects, the technical solution of this invention is as follows: An auxiliary generator generates pyramid features through a Gaussian pyramid module. This auxiliary generator can be used to generate auxiliary information during generator training, and its output can be used together with the output of the main generator to generate samples. The Gaussian pyramid module contains a series of convolutional operations, which can downsample the input features layer by layer to generate pyramid features. The pyramid features are flattened into a one-dimensional vector, and the flattened features are concatenated to obtain a large feature vector, which is input into a linear layer and activation function for mapping and nonlinear transformation. The mapped features are output as auxiliary information through the linear layer and the Sigmoid activation function.
[0008] The auxiliary information generated by the auxiliary generator is concatenated with the input image along the channel dimension, and then passed through the generator G. The input image passes through a series of convolutional layers and instance normalization layers to extract image features. After each convolutional layer, a non-linear transformation is performed using the LeakyReLU activation function. Residual blocks are used in the bottleneck layer to enhance the generation capability and depth. A pyramid segmentation attention mechanism is used to obtain a weighted feature map. An upsampling operation is performed in the deconvolutional layer to gradually restore the feature map to the size of the original image. Finally, the adversarial perturbation is output.
[0009] Gradient updates are performed on the generator and discriminator during training. The original samples are adjusted using generated perturbations, and pixel values are constrained to obtain adversarial examples. The discriminator network D is used to predict the generated adversarial examples, yielding the discriminator's prediction results. The overall loss of discriminator D is calculated, and gradient updates are performed on the discriminator. The adversarial loss, perturbation loss, and overall loss of the generator are calculated, and gradient updates are performed on the auxiliary generator AG and the generator G.
[0010] Compared with the prior art, the present invention has the following beneficial effects.
[0011] This invention introduces an auxiliary generator based on generative adversarial networks (GANs). This method enhances the generator's performance and improves the quality and diversity of generated adversarial examples. The auxiliary generator network generates supplementary information and concatenates it with the original image, providing the generator with richer input information and further improving the generation effect. A pyramid attention module is introduced, allowing the network to focus on important regions and features of the image at different scales. Through a series of neural network layers, a high-resolution, realistic image containing perturbations is generated. The improved generated adversarial examples have a higher success rate in black-box adversarial attacks and exhibit some transferability. Compared with existing black-box attack methods, this scheme provides an effective solution for adversarial attacks in black-box scenarios. Attached Figure Description
[0012] Figure 1 This is a flowchart illustrating the overall design of a transferable adversarial attack method based on generative adversarial networks.
[0013] Figure 2 This is a diagram of the auxiliary generator network structure.
[0014] Figure 3 This is a diagram of the generator network structure.
[0015] Figure 4 A structural diagram of the pyramid segmentation attention mechanism.
[0016] Figure 5 This is a diagram of the discriminator network structure. Detailed Implementation
[0017] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0018] like Figure 1 As shown, the specific operation process of a transferable adversarial attack method based on generative adversarial networks according to the present invention includes the following steps.
[0019] The original sample x is processed by the Gaussian pyramid module of the auxiliary generator to generate pyramid features. The Gaussian pyramid module contains a series of convolution operations that downsample the input features layer by layer to generate pyramid features. Figure 2 As shown, the pyramid features are flattened into one-dimensional vectors, and the flattened features are concatenated. When generating Gaussian pyramid features, the pyramid layers are iterated iteratively, and the corresponding convolutional layers are used to perform convolution operations on the input. For each pyramid layer i, the input image is convolved to obtain the pyramid feature X. i Its size is B i × C i B i = B / (2 i ), C i = C / (2 i Here, it is assumed that the pyramid layer indices start from 0, meaning the top layer is the original image. Flattening pyramid features: For each pyramid feature X... i Flatten it into a two-dimensional tensor of shape (batch_size, -1), and construct X. flat Finally, all the flattened pyramid features X flat The features are concatenated along dimension 1 to obtain a large feature vector, which is then input into a linear layer and activation function for mapping and nonlinear transformation. Finally, the mapped features are output as auxiliary information through a linear layer and a sigmoid activation function.
[0020] The auxiliary information generated by the auxiliary generator is concatenated with the input image along the channel dimension, and then the generator G generates adversarial perturbations. For example... Figure 3 As shown, the input image passes through a series of convolutional layers and instance normalization layers to extract image features. After each convolutional layer, a LeakyReLU activation function is used for non-linear transformation. Residual blocks are used in the bottleneck layer to enhance generation capability and depth. A weighted feature map is obtained through a pyramid segmentation attention mechanism, as shown below. Figure 4 As shown, adaptive average pooling is first used to pool the height and width of the input feature map to 1. Then, a convolutional layer is used to reduce the dimensionality of the feature map. The LeakyReLU activation function is applied to the dimensionality-reduced feature map to add non-linearity. The sigmoid activation function is then applied to the feature map after restoring the number of channels, restricting the feature map to the range [0, 1], resulting in a feature map with spatial attention weights. An upsampling operation is performed in the deconvolutional layer to gradually restore the feature map to the size of the original image. Finally, adversarial perturbation is output.
[0021] Gradient updates are performed on the generator and discriminator during training. The generated perturbations are used to adjust the original sample x, and pixel values are constrained to obtain adversarial examples. For example... Figure 5As shown, the discriminator network D is used to predict the generated adversarial examples, obtaining the discriminator's prediction results for the generated examples. The overall loss of the discriminator D is calculated, and its gradient is updated. The adversarial loss, perturbation loss, and overall loss of the generator are calculated, and the gradients of the auxiliary generator AG and generator G are updated. The adaptive learning rate optimization algorithm AMSGrad is used to update the parameters of the auxiliary generator AG and generator G, gradually optimizing them and approximating their ability to generate realistic examples.
[0022] The above description, in conjunction with the accompanying drawings, provides a detailed introduction to the method of the present invention. The specific embodiments described herein are merely illustrative and are intended to aid in understanding the method. Those skilled in the art will recognize that variations and modifications can be made to the specific embodiments and applications based on the principles of this invention; therefore, this invention should not be construed as limiting the scope of the invention.
Claims
1. A transferable adversarial attack method based on generative adversarial networks, characterized in that, Includes the following steps: Step 1: Generate auxiliary information from the original sample x using the auxiliary generator AG; the specific process of the auxiliary generator AG is as follows: The original sample x generates pyramid features through the Gaussian pyramid module of the auxiliary generator. The Gaussian pyramid module contains a series of convolution operations that downsample the input features layer by layer to generate pyramid features. The pyramid features are flattened into a one-dimensional vector and then concatenated to obtain a large feature vector, which is then input into a linear layer and activation function for mapping and nonlinear transformation. The mapped features are output as auxiliary information through the linear layer and the Sigmoid activation function. Step 2: The auxiliary information generated by the auxiliary generator is concatenated with the input image in the channel dimension, and the adversarial perturbation is generated by the generator G. The specific process of the generator G designed in this invention is as follows: The input image passes through a series of convolutional layers and instance normalization layers to extract image features. After each convolutional layer, a nonlinear transformation is performed by the ReLU activation function. Residual blocks are used in the bottleneck layer to enhance the generation capability and depth. A weighted feature map is obtained through the pyramid segmentation attention mechanism. An upsampling operation is performed in the deconvolution layer to gradually restore the feature map to the size of the original image. Finally, the adversarial perturbation is output. Step 3: During training, perform gradient updates on the generator and discriminator; adjust the original sample x using the generated perturbation and restrict the pixel values to obtain adversarial examples; use the discriminator network D to predict the generated adversarial examples to obtain the discriminator's prediction results for the generated samples; calculate the overall loss of the discriminator D, perform gradient updates on the discriminator, calculate the adversarial loss, perturbation loss, and overall loss of the generator, and perform gradient updates on the auxiliary generator AG and generator G; use the adaptive learning rate optimization algorithm AMSGrad to update the parameters of the auxiliary generator AG and generator G, so that they are gradually optimized and approach the ability to generate real samples; Step 4: Repeat the steps in Step 3 until the predetermined number of training iterations is reached or the stopping condition is met.
2. The transferable adversarial attack method based on generative adversarial networks according to claim 1, characterized in that: The auxiliary generator (AG) and the generator (G) are trained and optimized simultaneously in a generative adversarial network. The auxiliary generator (AG) defines a Gaussian pyramid consisting of multiple convolutional layers. In the linear layers, the amount of auxiliary information is controlled by defining the output dimension. When generating Gaussian pyramid features, the pyramid layers are iteratively increased, and the corresponding convolutional layers are used to convolve the input. For each pyramid layer i, the input image is convolved to obtain the pyramid feature X. i Its size is B i ×C i B i = B / (2 i ), C i = C / (2 i Here, it is assumed that the pyramid layer index starts from 0, meaning the top layer is the original image; Flattening pyramid features: For each pyramid feature X... i Flatten it into a two-dimensional tensor of shape (batch_size, -1), and construct X. flat Finally, all the flattened pyramid features X flat The pieces are stitched together on dimension 1.
3. The transferable adversarial attack method based on generative adversarial networks according to claim 1, characterized in that: After the original image passes through the auxiliary generator, the auxiliary output is concatenated with the original image in the channel dimension to obtain a dimensionally expanded input tensor for the generator G to use; here, bilinear interpolation is used to adjust the shape of the auxiliary output to be the same size as the original image.
4. The transferable adversarial attack method based on generative adversarial networks according to claim 1, characterized in that: The AMSGrad optimization algorithm is used when updating the gradients of the auxiliary generator AG and the generator G; the initial learning rate is set to 1×10. -4 β1=0.9, β2=0.999, and the learning rate decreases with the increase of the number of iterations thereafter; the AMSGrad optimization algorithm corrects the update of the second moment estimate v when updating the first moment estimate m and the second moment estimate v, changing the variable v to the maximum value of the gradient variance observed in history.
Citation Information
Patent Citations
Flow data generation method, device and system based on improved DCGAN model
CN112906019A
Adversarial sample generation method and system based on generative adversarial network
CN115641471A