A GAN-based black-box transferability adversarial attack method

By constructing a GAN-based neural network target model and introducing expanded convolutional residual blocks and pyramid segmentation attention mechanisms, the problems of limited generation capabilities and difficulty in scaling to real-world data sets in the existing technology are solved, and efficient generation and high success rate attacks of adversarial samples in complex tasks are achieved.

CN117057408BActive Publication Date: 2025-09-02XIAN UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310266763.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-17
Publication Date
2025-09-02
Estimated Expiration
2043-03-17

AI Technical Summary

Technical Problem

The existing GAN-based black box adversarial attack methods have problems such as limited generation capabilities and difficulty in scaling to real-world datasets when generating adversarial samples, and the success rate of black box attacks in complex tasks is not high.

Method used

Build a GAN-based neural network target model, design a black box attack scenario, use a generative adversarial network to generate adversarial samples, and introduce expanded convolutional residual blocks and pyramid segmentation attention mechanisms into the generator to enhance feature expression capabilities, and improve the migration and universality of attacks through proxy model distillation and data enhancement.

Benefits of technology

The adversarial sample generation efficiency and image quality are improved in complex real-life tasks, and the migrationability and attack success rate of adversarial samples are enhanced, especially on lung X-Ray images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117057408B_ABST
    Figure CN117057408B_ABST
Patent Text Reader

Abstract

In response to the problems of low attack success rate and low generation quality of existing adversarial methods in black box scenarios, the present invention discloses a black box transferable adversarial attack method based on GAN. First, a neural network target model is built, and the black box adversarial attack framework is used to train the proxy model to achieve transferable adversarial attack, thereby obtaining a more effective high black box attack success rate. Secondly, a GAN-based adversarial attack network is constructed, and both the generator G and the discriminator D adopt an end-to-end training method, using clean images and target categories as input to perform targeted adversarial attacks. A residual block based on dilated convolution and a lightweight and efficient pyramid segmentation attention module are designed in the generator to improve the model's multi-scale feature expression capability at a finer granularity. A discriminator with an auxiliary classifier is set to correctly classify the generated samples, and an attacker is added to perform adversarial training on the discriminator, thereby enhancing the attack capability of the adversarial samples and stabilizing the training process of GAN.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence security based on deep learning, and in particular to a GAN-based black box transferability counterattack method. Background Art

[0002] The development of neural networks has improved people's lives, but their inherent uninterpretability and fragility have raised questions about their security. In 2014, Goodfellow, Szegedy, and others discovered that deep neural network models are vulnerable to adversarial examples. These examples are generated by adding perturbations that are imperceptible to the human eye to clean input samples. The emergence of adversarial examples has raised security concerns for sensitive applications. After the discovery of adversarial examples that can mislead deep neural networks, various adversarial attack methods have been proposed. Adversarial attacks can be divided into white-box and black-box attacks based on the amount of information exposed to the attacker by the target model. White-box algorithms are easier and more efficient than black-box algorithms to generate adversarial perturbations because they can leverage the full knowledge of the target model, including model weights, architecture, and gradients. For example, the Fast Gradient Sign Method (FGSM) adds increments to the gradient to cause the model to misclassify samples. The Projected Gradient Descent (PGD) attack method performs multiple iterations during the gradient iteration process, controlling the perturbation within a specified range. The optimization-based iterative attack method (C&W) mainly fixes the network parameters during iterative training, treating the perturbation as the only parameter that needs to be trained, and adjusting the adversarial perturbation through backpropagation. However, due to privacy and security concerns, this attack scenario is generally not feasible in practical deployments. In more realistic adversarial scenarios, attackers can employ query-based black-box attacks. Although model information is hidden in the black-box attack, the attacker can query the model and observe the corresponding label predictions. However, this method is typically time-consuming and has a low success rate in most black-box attack scenarios. Therefore, most current black-box attack methods are based on the transferability of adversarial examples. This transferability of adversarial examples can be used to train alternative models to deceive unknown target models.

[0003] Some researchers have also used generative models, such as GANs, to generate adversarial perturbations or directly generate adversarial samples. Compared with gradient-based and optimization-based methods, generative models greatly reduce the time it takes to generate adversarial samples. However, existing methods have two significant drawbacks: 1) limited generation capabilities, i.e., they can only perform one specific target attack at a time, and different targets require retraining. 2) They are difficult to scale to real-world datasets. Most GAN-based adversarial attack methods have only been tested and evaluated on the MNIST and CIFAR-10 datasets and have achieved good results, but are not feasible for complex real-world tasks.

[0004] To address these existing challenges, we propose a GAN-based black-box transferable adversarial attack method. We construct a GAN network to generate adversarial examples for adversarial attacks. We also design black-box adversarial attack scenarios to increase the transferability and versatility of attack targets, improving the efficiency of adversarial example generation and image quality. Furthermore, we achieve high attack performance on the MNIST and CIFAR-10 datasets. Experiments on more realistic lung X-ray images demonstrate the effectiveness and feasibility of the proposed attack method. Summary of the Invention

[0005] The purpose of this invention is to provide a black-box transferable adversarial attack method based on GAN. First, a neural network target model is constructed, and a black-box attack scenario is designed to implement transferable adversarial attack. Second, adversarial samples are generated using a generative adversarial network, and a residual block based on dilated convolution and a pyramid segmentation attention mechanism are designed in the generator to enhance the feature expression capability. Finally, adversarial attacks are carried out on the target model using adversarial samples to identify and expose defects and security issues in the model, providing a reference solution for guiding the model to carry out targeted defense and enhance the adversarial robustness of the model.

[0006] The present invention provides a GAN-based black box transferability countermeasure attack method, which specifically includes the following steps:

[0007] (1) Construct a neural network target model. The specific implementation process is as follows:

[0008] Use the CheXNet model to build the target model T. The CheXNet model uses the DensNet121 network as its basic framework, replaces the 7×7 large convolution with a 3×3 small convolution to reduce the number of model parameters, and fully extracts the edge texture feature information in the image through dense connections;

[0009] Initialize the network weights using the weights from a pre-trained model on the ImageNet dataset and train the network end-to-end using the SGD+Momentum optimization algorithm with standard parameters.

[0010] At the end of the model, convolution is used instead of the fully connected layer, and the Sigmoid function is used to complete the final classification output of the model to achieve multi-label classification of images;

[0011] Continuously tune the parameters until the target model reaches the optimal accuracy and then save it.

[0012] (2) Design a black-box attack scenario and build a proxy model S to implement transferable adversarial attacks, which specifically includes the following steps:

[0013] Synthetic data: Map a batch of random noise Z to the required data X = VAE(Z). The goal of the generative model VAE is to synthesize data with a distribution close to the data required for target training. Input the synthesized training data X into the proxy model S and minimize the loss function to update the generative model. The generation loss is expressed as:

[0014]

[0015] Where d is the cross entropy loss function, S(X) is the data generated by the proxy model, Y is the random smoothing label, α is the hyperparameter for adjusting the regularization value, and L H is the information entropy loss;

[0016] Model distillation: To significantly improve the success rate of black-box attacks, when distilling the proxy model and the target model, we encourage the proxy model S and the target model T to have highly consistent decision boundaries to facilitate the training of the proxy model. Therefore, during the distillation process, we need to pay more attention to the two types of data. The final loss function consists of three parts. The loss function of the proxy model is defined as:

[0017]

[0018] Where: L dis represents the distillation loss between the target model and the proxy model, L bd represents the boundary support loss generated when there is decision disagreement between the proxy model S and the target model T, L adv Represents the adversarial sample support loss generated when data is easily transferred from the proxy model S to the target model T. β1 and β2 are used to control the proportion of the two loss functions;

[0019] Finally, an adversarial attack is carried out on the distilled and refined network.

[0020] (3) Generate adversarial samples using the GAN network to achieve a high black-box attack success rate and target transferability attack, specifically including the following steps:

[0021] Input the original sample x and the target category t into the generator G to generate the perturbation and then superimpose it on the original sample to generate the adversarial sample X pert And sent to the discriminator D;

[0022] The adversarial sample X generated by attacker a adv The original sample x is also fed into the discriminator. Since an auxiliary classifier is set in the discriminator, the discriminator can not only guide the training of the generator through the optimization function to make the generated adversarial samples indistinguishable from the real data, but also correctly classify the two types of adversarial samples.

[0023] Implement adversarial attack, with Xpert The input and output loss L target , which represents the distance of predicting the target class (targeted attack) as opposed to the distance of predicting the true class (untargeted attack).

[0024] (4) Generator structure design, the specific structure is as follows:

[0025] The ResNet-50 model is used as the main network of the generator. The residual block structure is used to simplify the deep learning process, enhance gradient propagation, and solve the degradation problem of deep neural networks.

[0026] Using a pre-trained encoder-decoder structure, the input image is encoded and mapped to the feature space, and the features are decoded and mapped back to the data space to complete data reconstruction, further learning the mapping relationship from input to feature space. In addition, dilated convolution is used in the feature generation block to effectively increase the receptive field of the convolution kernel, which can efficiently generate targeted adversarial perturbations when extracting features.

[0027] A lightweight and efficient pyramid segmentation attention module is introduced between the original sample input and the generator output. This attention module can fully extract the spatial information of multi-scale feature maps and realize cross-dimensional channel attention feature interaction, capturing the interdependence between long-range feature channels and improving network performance.

[0028] When using lung images for testing, due to the particularity of medical images, data augmentation methods are introduced as a mechanism into generative model training. On the one hand, diverse data augmentation methods can enrich the gradient flow information returned by the target model to increase data diversity. On the other hand, the introduction of data augmentation enables the generator to resist various data transformations to enhance the robustness of adversarial samples.

[0029] (5) Discriminator structure design, the specific structure is as follows:

[0030] Further improvements are made to the original GAN ​​by setting up an auxiliary classifier to obtain image classification capabilities to improve the performance of the original task. After adding the classifier to the discriminator, the discriminator can not only distinguish between real and fake images, but also distinguish between categories. Therefore, the discriminator loss consists of two parts: the discrimination loss and the classification loss. The classification loss is the cross entropy loss calculated by comparing the adversarial examples generated by the generator and the adversarial examples generated by the attacker with the true labels.

[0031] After the discriminator generates the adversarial loss, it is optimized and fed back to the generator network to guide the training of the generator to ensure that the generated adversarial samples are close to the data of the real image, thereby ensuring the authenticity of the adversarial samples.

[0032] (6) Test and evaluate the trained generator G. The generator G that has converged in training is allowed to generate perturbations on the test sets of different datasets to generate test adversarial samples. The test adversarial samples are allowed to attack the target classification network, and different target categories are set to perform targeted adversarial attacks. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The accompanying drawings are only used to more fully illustrate the process of the present invention and do not constitute a limitation to the scope of the present invention.

[0034] Figure 1 This is a basic flow chart of adversarial training in the present invention;

[0035] Figure 2 This is the overall architecture diagram for countering attacks in the present invention;

[0036] Figure 3 This is the overall architecture diagram of the agent model constructed in the present invention, where module (a) represents an efficient data synthesis method.

[0037] (b) Module representation distillation method for alternative models;

[0038] Figure 4 This is the pyramid segmentation attention module diagram introduced in the generator model of the present invention;

[0039] Figure 5 This is the SPC module diagram introduced in the pyramid segmentation attention module of the present invention;

[0040] Figure 6 This is a graph showing the FID score comparison experiment results in the present invention;

[0041] Figure 7 This is a comparison result of the success rate of the counterattack in the present invention;

[0042] Figure 8 This is the SSIM comparison experiment result diagram in the present invention. Specific implementation plan

[0043] In order to enable relevant personnel in this field to better understand the workflow of this method, the following will make a systematic and complete explanation of this method with reference to the accompanying drawings. Figure 2 shown.

[0044] Figure 1 The basic process of adversarial training in the present invention is depicted, and its main functions include:

[0045] Step 1: First, build the CheXNet model as the target model for transfer learning, using the DenseNet121 network as the basic framework. At the end of the model, use the convolutional layer instead of the fully connected layer, and use the 3×3 small convolution to replace the 7×7 large convolution to reduce the number of model parameters. Use the weights from the pre-trained model on the ImageNet dataset to initialize the network weights. Use the SGD+Momentum algorithm for iterative optimization, add the Sigmoid nonlinear activation function to achieve the final classification output of the model, train the model until convergence, and save the target model T.

[0046] Step 2: Use efficient data synthesis methods and alternative model distillation methods to train a proxy model S as an adversarial attack network to achieve black box transferability adversarial attack, and build the overall architecture of the proxy model as follows: Figure 3 As shown, the specific steps include:

[0047] First, a batch of random noise Z is mapped to the desired data X = VAE (Z). The goal is to synthesize the desired data with a distribution close to the target training data. The synthesized data X is input into the proxy model S to calculate the loss. To solve the problem of model collapse during training, the maximum information entropy and random label smoothing strategy are introduced. Minimizing the loss function is used to update the generative model. The final generative loss is expressed as:

[0048]

[0049] Where d is the cross entropy loss function, S(X) is the data synthesized by the generator input to the proxy model, Y is the random smoothing label, α is the hyperparameter for adjusting the regularization value, and L H is the information entropy loss;

[0050] The second step is to distill the proxy model and the target model, train the proxy model to imitate the target model, and minimize the distillation network:

[0051]

[0052] Where: d represents the cross entropy loss function, T(X) represents the output of the target model, and S(X) represents the output of the distilled proxy model;

[0053] In order to make the proxy model S and the target model T have highly consistent decision boundaries to promote the training of the proxy model, we need to pay more attention to two types of data during the distillation process. The first type refers to the data with decision differences between S and T. This type of data mainly exists between the decision boundaries of the target model and the proxy model. Giving this data more weight helps to bridge the gap between the two decision boundaries. Because of paying more attention to these samples, the boundary support loss L is introduced bd :

[0054]

[0055] Another important type is the adversarial samples generated during the adversarial attack. This type of data can be easily transferred from S to T. The existence of this type of data means that the decision boundary of S and T is close to it. Paying more attention to this type of data can ensure that S continues to move in the right direction close to the boundary of T. Therefore, the adversarial sample support loss L is introduced. adv :

[0056]

[0057] Where: Represents the adversarial sample, and the loss function of the final proxy model S is defined as:

[0058]

[0059] Where: β1 and β2 control the proportion of different loss functions;

[0060] By optimizing the distillation target of all training images, a proxy model S is obtained, whose characteristics are very close to the black-box target model, and then an adversarial attack is carried out on the distilled network;

[0061] Step 3: Input the original image x and the target class label t into the generator G. The generator G outputs the perturbation G(x,t). G(x,t) is clipped so that the range of G(x,t) is between (-c_treshold,c_treshold), where c_treshold is the set perturbation coefficient. The generated perturbation G(x,t) is then superimposed on the original sample x to obtain the adversarial sample X. pert =x+G(x,t), where the goal of the generator is not to directly generate adversarial samples, but to superimpose the generated perturbations on the original samples before outputting the adversarial samples. The purpose is to dynamically adjust the perturbation size to prevent excessive perturbations. The loss function of the generator includes the adversarial loss L generated by attacking the target model. target (pert) and the discriminant loss L generated when inputting the discriminator D (pert), specifically expressed as follows:

[0062]

[0063]

[0064] Where: X pert represents the adversarial sample generated by the generator, t is the category of the target attack, and maximizes L target (pert)+L D (pert)-L SMake the results of adversarial samples during the attack process closer to the expected value.

[0065] Step 4: Input the adversarial sample obtained in S3 into the discriminator D, which is used to distinguish the adversarial sample X pert And the original sample x. In order to further enhance the attack capability of the adversarial sample, we add attacker a to conduct adversarial training on the classification model. A robust discriminator helps to stabilize and accelerate the entire training. adv It is also input into the discriminator. Since the auxiliary classifier is set in the discriminator D, the discriminator D can also correctly classify the sample. Finally, there are two branches in the discriminator D: one is used to train the real image X real and the perturbation image X pert , and the other is to classify adversarial samples. The loss function of the discriminator consists of three parts: the cross entropy loss L used to distinguish real / perturbed images S , the classification loss L generated by the attacker and the generator C (adv) and L C (pert), defined as:

[0066]

[0067]

[0068]

[0069] Where: X real represents the real sample, X pert represents the adversarial sample generated by the generator, X adv represents the adversarial sample generated by attacker a, y represents the true label, and maximizes the loss function Ls+Lc(adv)+L C (pert) Make the generated image infinitely close to the real image to ensure the quality of the adversarial sample;

[0070] Step 5: Use the Adam method to optimize the loss function of the generator and discriminator, use backpropagation to modify the model weights, and continuously adjust the model parameters until the model reaches a convergence state and save it. The training of the generator G is completed;

[0071] In step 6, the trained generator G is tested by generating perturbations using test sets of different datasets to generate test adversarial samples. The test adversarial samples are input into the target classification network, and different target categories are set to perform targeted adversarial attacks.

[0072] Figure 4 The pyramid segmentation attention module introduced in the generator is shown, which mainly consists of the following four steps:

[0073] (1) First, the SPC module is used to segment the channel, and then multi-scale feature extraction is performed on the spatial information on each channel feature map to obtain the multi-scale feature map on the channel;

[0074] F=Cat([F0,F1,…,F N-1 ])

[0075] Where: Splitting and fusion module SPC is as follows Figure 5 As shown, in order to obtain different spatial resolutions and depths, the input feature map is divided into N groups at the channel level, represented as [X0, X1...., X N-1 ], each group performs convolution k of different scales i =

[0076] 2*(i+1)+1(i=0,1,...,N-1), so that a feature map containing a single type of convolution kernel can be obtained to extract the spatial information on each channel feature map. For each segmentation part, it can independently learn multi-scale spatial information and establish cross-channel interaction in a local manner. However, as the convolution kernel size increases, the computational complexity increases. Therefore, a multi-scale convolution kernel is used to group the features of each group, and the number of groups is The specific calculation method of the multi-scale feature extraction process is as follows:

[0077] F i =Conv(k i ×k i ,W i )(X i ),i=0,1,2…N-1

[0078] (2) The SEWeight module is used to extract the channel attention of feature maps of different scales, and the channel attention vector at each different scale is obtained. The vector of attention weight can be expressed as:

[0079] Z i =SEWeight(F i ),i=0,1,2…N-1

[0080] In order to achieve the interaction of attention information, the cross-dimensional vectors are fused without destroying the original channel attention vector, and the entire multi-scale channel attention vector is obtained in a serial manner. The entire multi-scale channel attention weight vector is:

[0081]

[0082] (3) Use the Softmax function to recalibrate the multi-scale channel attention vector to obtain the new attention weight after multi-scale channel interaction. The multi-scale channel weight after interaction is expressed as:

[0083]

[0084] (4) Perform element-wise multiplication on the recalibrated weights and the corresponding feature maps, and output a feature map after multi-scale feature information attention weighting. The specific calculation is as follows:

[0085] Out=Cat([Y0,Y1,…,Y N-1 ])

[0086] Where: Y i is to focus the multi-scale channel t t i The recalibrated weights and the corresponding scale F i The feature map of the multi-scale channel attention weights is obtained by multiplying the feature map of , and the multi-scale information representation capability of the feature map is richer.

[0087] Through the above operations, multi-scale spatial information and cross-channel attention can be integrated into each split feature block in the ResNet-50 network, which can produce better pixel-level attention, extract multi-scale spatial information at a more granular level, and capture the dependencies of long-range channels, thereby enhancing the feature extraction capability of the generator.

[0088] The advantages and feasibility of the present invention are illustrated below by comparing the experimental results.

[0089] (1) Table 1 shows the time required to generate adversarial samples by the existing adversarial attack methods and the attack method of the present invention. As shown in the table, the method BA-GAN proposed in the present invention improves the efficiency of sample generation.

[0090] Table 1 Time for attack methods to generate adversarial samples

[0091]

[0092] (2) Table 2 shows the success rates of targeted adversarial attacks using different target categories on the MNIST and CIFAR-10 datasets, respectively.

[0093] Table 2 Target attack success rate

[0094]

[0095] (3) Common GAN-based adversarial attack methods AdvGAN, AdvGAN++, Natural-GAN, and Rob-GAN are used to compare with the BA-GAN method proposed in this paper on the lung X-Ray image dataset.

[0096] Figure 6 A comparison chart of the FID scores of different adversarial attack methods is shown. FID is a metric used to evaluate the quality of image generation; smaller FID values ​​indicate a higher similarity between the generated image and the real image. The figure shows that our method has the smallest FID value, resulting in a more realistic adversarial example.

[0097] Figure 7 The figure plots the success rates of different adversarial attack methods as the number of iterations increases. The figure shows that the adversarial attack method BA-GAN proposed in this paper outperforms other mainstream adversarial attack strategies in terms of success rate, and can significantly improve the attack success rate under the black-box attack method.

[0098] Figure 8 The structural similarity results of different adversarial attack methods are plotted. A higher SSIM indicates that the generated adversarial sample has a higher similarity to the real image in terms of brightness, contrast, and structure. It can be seen from the figure that the SSIM value of the present invention is the largest, and the generated image is closer to the real image.

Claims

1. A GAN-based black-box transferability adversarial attack method, whose features include: (1) Use the CheXNet model to build the target model T. The CheXNet model uses the DensNet121 network as the basic skeleton. At the end of the model, the convolution layer is used instead of the fully connected layer. The 3×3 small convolution is used to replace the 7×7 large convolution to reduce the number of model parameters. The network weights are initialized using the weights from the pre-trained model on the ImageNet dataset. The SGD+Momentum algorithm is used for iterative optimization. The Sigmoid nonlinear activation function is added to achieve the final classification output of the model. The model is trained until convergence is reached and the target model T is saved. (2) Design a black-box attack scenario and build a proxy model S to implement transferable adversarial attacks. First, perform data synthesis. Set the target of the generative model VAE to a synthetic distribution close to the target training data X and input it into the proxy model S. Minimize the loss function to update the generative model. In order to solve the problem of model collapse during training, the maximum information entropy and random label smoothing strategy are introduced. The generation loss is expressed as: Where d is the cross entropy loss function, S(X) is the data generated by the proxy model, Y is the random smoothing label, α is the hyperparameter for adjusting the regularization value, and L H is the information entropy loss; Secondly, we use the model distillation method to train the proxy model to effectively imitate the target model, so that the proxy model S and the target model T have highly consistent decision boundaries to facilitate the training of the proxy model. The loss function of the proxy model is defined as: Where: L dis represents the distillation loss between the target model and the proxy model, L bd represents the boundary support loss caused by decision disagreement between the proxy model and the target model, L adv It represents the adversarial sample support loss generated when the adversarial sample is easily transferred from the proxy model S to the target model T. β1 and β2 are used to control the proportion of the two loss functions; (3) Constructing a GAN-based adversarial attack network to achieve target transferability adversarial attack and obtain a high black-box attack success rate; (4) Input the original image x and target category t into the generator G, superimpose high-dimensional noise to generate the adversarial perturbation G(x,t), and then transform X pert =x+G(x,t) and the original image x are fed into the discriminator D to be judged as the original input or the adversarial sample; (5) In order to enhance the attack capability of adversarial samples and stabilize the overall training process, attacker a is introduced into the discriminator for adversarial training, and an auxiliary classifier C is set in the discriminator D to achieve correct classification of samples; (6) After training the proxy model S and the generator G, use the adversarial sample X generated by the generator G pert Perform targeted attacks.

2. The GAN-based black-box transferability counterattack method according to claim 1, characterized in that: Using the AC-GAN discriminator, an auxiliary classifier is set up to distinguish between real images and perturbed images, and to correctly classify adversarial samples; the loss function of the discriminator D specifically includes three parts: the cross entropy loss L generated by distinguishing between real and perturbed images S , the loss L generated when classifying the adversarial samples generated by attacker a and the adversarial samples generated by generator G C (adv) and L C (pert), specifically expressed as follows: Where: X real represents the real sample, X pert represents the adversarial sample generated by the generator, X adv represents the adversarial sample generated by attacker a, y represents the true label, and maximizes the loss function Ls+Lc(adv)+L C (pert) makes the generated image infinitely close to the real image, ensuring the quality of the adversarial sample.

3. The GAN-based black-box transferability counterattack method according to claim 1, characterized in that: The generator G uses the ResNet-50 model as the basic skeleton, uses the encoding-decoding structure for feature extraction, and designs residual blocks, dilated convolutions, and pyramid segmentation attention mechanisms to enhance the generator's feature expression capabilities; the generator's loss function includes the adversarial loss L generated by attacking the target model. target (pert) and the discriminant loss L generated when inputting the discriminator D (pert), specifically expressed as follows: Where: X pert represents the adversarial sample generated by the generator, t is the category of the target attack, and maximizes L target (pert)+L D (pert)-L S Make the results of adversarial samples during the attack process closer to the expected value.

4. The GAN-based black-box transferability counterattack method according to claim 1, characterized in that: Step (6) should also include testing and evaluating the trained generator G, allowing the generator G that has converged in training to generate adversarial perturbations on the test sets of different datasets, thereby generating test adversarial samples, allowing the test adversarial samples to attack the target classification network, and setting different target categories to perform targeted adversarial attacks.

Citation Information

Patent Citations

  • Generative adversarial network-based adversarial attack sample generation method

    CN111275115A

  • Systems and methods for defense against adversarial attacks using feature scattering-based adversarial training

    US20210012188A1