An Adversarial Sample Generation Method Based on AdvGAN
By building an AdvGAN generative adversarial network, using noise generation and iterative optimization technologies, the problem of low quality of adversarial samples in the existing technology is solved, and high-quality adversarial samples are efficiently generated, which improves the attack success rate and migration.
Patent Information
- Application Number
- CN202211212575.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-29
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-09-29
AI Technical Summary
The quality of the adversarial samples generated by the prior art is low, can be recognized by the human eye, has low generation success rate, low attack mobility, large calculation volume and slow speed.
Build an adversarial network of AdvGAN, including generators, discriminators and distillation models, generate perturbation samples by adding random noise and Gaussian noise, set up adversarial loss functions and discriminators, and use the distillation model to iterate the generator and discriminator to optimize the total loss function to generate high-quality adversarial samples.
The speed and quality of the generation of adversarial samples is improved, the difficulty of human eye perception is reduced, the success rate of attack is increased, and the differences in the characteristics of generated adversarial samples and the real sample are reduced, which improves the migration of attacks.
Smart Images

Figure CN115510986B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning security, and more specifically, to an adversarial sample generation method based on AdvGAN. Background Art
[0002] As a part of the field of machine learning, deep learning has gradually become a research hotspot in the field of computer vision. With the continuous development of deep neural network models, more and more deep learning frameworks and tools have been developed. At the same time, with the continuous improvement of the performance of hardware such as GPU computing power required for training, the cost of training complex models has become lower and lower, making deep learning widely used in various fields. At the same time, due to the wide range of applications, security issues cannot be ignored. Common ones such as autonomous driving, face recognition, and logistics handling are the key objects of concern for computer vision task security issues. Many currently commonly used neural network structures are very vulnerable to some adversarial perturbations, such as GoogleNet, VGGNet, ResNet, and Inception vn, etc. Thus, the problem of adversarial attacks in deep learning is proposed. The essence of adversarial attacks is to add tiny perturbations to the input samples of deep learning models, etc., so as to generate adversarial samples that cause the model to make incorrect identifications. However, adversarial samples are not observable to humans. The generation of adversarial samples needs to meet an important criterion, that is, the perturbed samples should be indistinguishable from the original samples in the eyes of humans, but after being input into the model, the model will output an incorrect result, or even a specified incorrect result.
[0003] With the publication of a paper titled "Intriguing properties of neural networks" in 2014, adversarial attacks have officially become a research area in deep learning security. Research in this area focuses on proposing adversarial attack methods against existing deep learning models to prove the security issues of the models, as well as adversarial training for enhancing model robustness against security issues. How to effectively attack deep learning models is an important method for analyzing the security of deep learning models and improving model robustness. Traditional attack methods focus on adding perturbations to the original images. Common methods include single-step calculation or iterative calculation, such as the Fast Gradient Sign Method (FGSM), Iterative Fast Gradient Sign Method (I-FGSM), and infinity norm constraints. Pixel-level perturbations are applied to the original images through these methods, and then adversarial samples are generated. Traditional attack methods have the disadvantages of slow sample generation speed and large computational complexity, and most of them are white-box attacks. The process of generating samples requires obtaining the structure information, parameter content, etc. of the target model, resulting in relatively single applicability. Therefore, the newer direction of adversarial attacks turns to using neural networks, especially generative adversarial networks, for generating adversarial samples. However, there are also problems such as low quality of the generated adversarial samples, artificially recognizable added perturbations, low success rate of generating adversarial samples, and low transferability of the attacks.
[0004] The prior art discloses an adversarial sample generation method, device, equipment, and medium, including: training multiple different initial models using the training dataset of a black-box model to obtain multiple different surrogate models of the black-box model; generating adversarial samples corresponding to each original sample using the multiple surrogate models respectively; and determining the final adversarial sample corresponding to the original sample based on the multiple adversarial samples generated by the multiple surrogate models corresponding to each original sample. This application uses multiple different surrogate models and cannot guarantee the quality of the adversarial samples. Summary of the Invention
[0005] To overcome the defect of low quality of adversarial samples generated by the above prior art, the present invention provides an adversarial sample generation method based on AdvGAN, which can generate high-quality adversarial samples, reduce the perception degree of the human eye, and improve the success rate of generating adversarial samples.
[0006] To solve the above technical problems, the technical solution of the present invention is as follows:
[0007] The present invention provides an adversarial sample generation method based on AdvGAN, including:
[0008] S1: Construct an AdvGAN generative adversarial network, including a generator, a discriminator, and a distillation model;
[0009] S2: Obtain real samples, and add random noise to the real samples to obtain noise-enhanced samples;
[0010] S3: Input the noise-enhanced samples into the generator, calculate the perturbation using the gradient descent method, and combine it with the noise-enhanced samples to generate perturbed samples;
[0011] S4: Input the perturbed samples and real samples into the discriminator, set the adversarial loss function as the objective function; by optimizing the objective function, obtain the preliminarily trained generator and discriminator;
[0012] S5: Select Gaussian noise, input the Gaussian noise and real samples into the preliminarily trained generator, and generate adversarial samples according to the generator loss;
[0013] S6: Input the adversarial samples and real samples into the preliminarily trained discriminator, distinguish the adversarial samples and real samples, and calculate the discriminator loss;
[0014] S7: Calculate the norm distance between the adversarial samples and real samples, compare it with the preset perturbation range, and set the BIM perturbation loss constraint or hinge loss constraint according to the comparison result;
[0015] S8: Update the parameters of the preliminarily trained generator and discriminator according to the generator loss and discriminator loss, and obtain the preliminarily optimized generator and discriminator;
[0016] S9: Input the adversarial samples and perturbed samples into the distillation model, set the adversarial loss function of the distillation model, and use the dynamic distillation method to iteratively optimize the distillation model with the preliminarily optimized generator and discriminator;
[0017] S10: Establish the total loss function according to the adversarial loss function, BIM perturbation loss constraint, hinge loss constraint, and adversarial loss function; by optimizing the total loss function, obtain the finally optimized distillation model, generator, and discriminator, and form the optimized AdvGAN generative adversarial network;
[0018] S11: Use the optimized AdvGAN generative adversarial network to generate the final adversarial samples.
[0019] Preferably, the specific method of step S2 is:
[0020] S2.1: Obtain real samples \(x\in X\), and the noise distribution of the real samples is \(N(0,\sum * )\); where \(X\) represents the real sample dataset, and \(\sum * \) represents the noise matrix;
[0021] S2.2: For each real sample \(x\), select random noise \(n\in N\), add noise constraints and dispersion constraints to generate noise-enhanced samples \(x'\), and obtain the noise-enhanced dataset \(X'\);
[0022] The noise constraint is as follows:
[0023]
[0024] The dispersion constraint is as follows:
[0025]
[0026] where X enc,n (x, x′) represents the noise constraint between the true sample x and the noise-enhanced sample x′, and f w (*) represents the feature extraction operation, argmin represents the function to obtain the value of x′ when minimizing ||f w+n (x) - f w (x′)||2; ||*||2 represents the operation of obtaining the L2 norm; KL(p x′ (x′)||q X (x)) represents the KL divergence between the true sample x and the noise-enhanced sample x′, p(x′) represents the probability density function of the noise-enhanced sample x′, and q(x) represents the probability density function of the true sample x.
[0027] The KL divergence is used to judge the difference in data distributions between the true sample and the noise-enhanced sample, so that the data distribution of the noise-enhanced sample does not differ too much from the data distribution of the original true sample, and they are consistent in data features.
[0028] Preferably, the specific method of step S3 is as follows:
[0029] Randomly select a noise-enhanced sample x′ and input it into the generator G. On the premise of maintaining the data features of the noise-enhanced sample, use the gradient descent method to iteratively calculate the perturbation δ adv :
[0030]
[0031]
[0032]
[0033]
[0034] where g c represents the gradient symbol, represents the expected perturbation at the t-th iteration, w represents the gradient parameter, y represents the feature label, represents the gradient momentum at the t-th iteration; represents the gradient momentum at the (t + 1)-th iteration, μ represents the decay factor, a represents the sign coefficient, and [-ε, ε] represents the preset perturbation range. Denote the expected perturbation of the (t + 1)-th iteration; Denote the perturbation of the (t + 1)-th iteration;
[0035] When and the difference is within the preset perturbation range, then is used as the perturbation δ adv ; Combine the perturbation δ adv with the noise-enhanced sample x′ to generate a perturbed sample
[0036] Preferably, in the step S3, after generating the perturbed sample, the perturbed sample needs to be screened. The specific method is as follows:
[0037] Calculate the L2 norm distance between the perturbed sample and the real sample x:
[0038]
[0039] In the formula, L p represents the L2 norm distance between the perturbed sample and the real sample;
[0040] Set a perturbation threshold ∈. If L p >|∈|, then discard the perturbed sample.
[0041] Preferably, in the step S4, the adversarial loss function is:
[0042]
[0043] In the formula, L GAN represents the adversarial loss function, f L (*) represents a function that satisfies 1-Lipschitz continuity, that is
[0044] Preferably, the specific method of the step S5 is:
[0045] Randomly select m real samples x, select Gaussian noise η, input the real samples and Gaussian noise into the preliminarily trained generator, initialize the parameters of the preliminarily trained generator, and generate an adversarial sample x adv , and calculate the generator loss:
[0046]
[0047] In the formula, represents the generator loss, G(*) represents the preliminarily trained generator, D(*) represents the preliminarily trained discriminator, and N′ represents the Gaussian noise distribution.
[0048] Preferably, in the step S6, the discriminator loss is:
[0049]
[0050] In the formula, represents the discriminator loss, G(*) represents the preliminarily trained generator, D(*) represents the preliminarily trained discriminator, and N′ represents the Gaussian noise distribution.
[0051] Preferably, the specific method of the step S7 is:
[0052] Calculate the L2 norm distance between the adversarial sample x adv and the real sample x:
[0053]
[0054] In the formula, represents the L2 norm distance between the adversarial sample and the real sample;
[0055] Compare with the set perturbation threshold ∈. If then set the BIM perturbation loss constraint:
[0056]
[0057] In the formula, L B represents the BIM perturbation loss; m represents the number of randomly selected real samples, k represents the number of adversarial samples, and σ represents the first hyperparameter;
[0058] If then set the hinge loss constraint:
[0059] L hinge = E x max(0, ||G(x)||2 - c)
[0060] In the formula, L hinge represents the hinge loss, and c represents the upper limit value.
[0061] Preferably, the specific method of the step S9 is:
[0062] S9.1: Input the adversarial sample and the perturbed sample into the distillation model;
[0063] S9.2: Set the distillation model f with the black-box model as the target model, set the target attack category t or do not set the attack category. The adversarial loss function of the distillation model f is:
[0064]
[0065] In the formula, Denotes the adversarial loss of the distillation model f, l f Denotes the first cross-entropy loss;
[0066] S9.3: Update the preliminarily trained generator G of the i-th iteration i-1 and the preliminarily trained discriminator D i using the distillation model f of the (i - 1)-th iteration; i The update method is:
[0067]
[0068] wherein, denotes the adversarial loss of the distillation model of the (i - 1)-th iteration; τ, denotes the second and third hyperparameters;
[0069] S9.4: Feedback and update the distillation model f of the i-th iteration i using the preliminarily trained generator G of the i-th iteration; i The update method is:
[0070]
[0071] wherein, denotes the second cross-entropy loss, f(x adv ) denotes the output of the distillation model for the adversarial sample x adv , denotes the output of the distillation model for the perturbed sample , b(x adv ) denotes the query result of the adversarial sample x adv from the black-box model, denotes the query result of the perturbed sample from the black-box model.
[0072] Preferably, the total loss function is:
[0073]
[0074] wherein, α, β, γ denote the first, second, and third proportionality coefficients.
[0075] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0076] The present invention first adds random noise to real samples to obtain noise-enhanced samples; inputs the noise-enhanced samples into a generator to generate perturbed samples; sets an adversarial loss function, and uses a discriminator to distinguish between the perturbed samples and the real samples, enabling the initially trained generator and discriminator to learn the distribution characteristics of the real samples and the perturbations; then adds Gaussian noise to the real samples and inputs them into the initially trained generator to generate adversarial samples; uses the initially trained discriminator to distinguish between the adversarial samples and the real samples, and sets a BIM perturbation loss constraint or a hinge loss constraint according to the difference between the adversarial samples and the real samples to limit the perturbation size of the adversarial samples, improve the quality of the adversarial samples without affecting the attack effect, and reduce the human perception thereof, thereby obtaining an initially optimized generator and discriminator; finally, sets a distillation model to iteratively optimize with the initially optimized generator and discriminator, enabling the initially optimized generator and discriminator to better learn the distribution characteristics of the adversarial samples, and making the distillation model closer to the target black-box model; by optimizing the total loss function, an optimized AdvGAN generative adversarial network is obtained, and the final adversarial samples of the specified target attack category or the non-target attack category are generated using the same. The present invention improves the generation speed of adversarial samples, improves the quality of the generated adversarial samples; reduces the difference between the distribution characteristics of the adversarial samples and the distribution characteristics of the real samples, increases the difficulty of human identification, and at the same time improves the attack success rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 FIG. is a flowchart of a method for generating adversarial samples based on AdvGAN according to Embodiment 1.
[0078] Figure 2 FIG. is a schematic structural diagram of an AdvGAN generative adversarial network according to Embodiment 2.
[0079] Figure 3 FIG. is a schematic structural diagram of the generator G according to Embodiment 2.
[0080] Figure 4 FIG. is a schematic structural diagram of the discriminator D according to Embodiment 2. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0081] The drawings are only for illustrative purposes and should not be construed as limiting the present patent;
[0082] For better illustration of this embodiment, some components in the drawings are omitted, enlarged or reduced, and do not represent the dimensions of the actual product;
[0083] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0084] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.
[0085] Example 1
[0086] This example provides an adversarial sample generation method based on AdvGAN, as Figure 1 shown, including:
[0087] S1: Construct an AdvGAN generative adversarial network, including a generator, a discriminator, and a distillation model;
[0088] S2: Obtain real samples and add random noise to the real samples to obtain noise-enhanced samples;
[0089] S3: Input the noise-enhanced samples into the generator, calculate the perturbation using the gradient descent method, and combine it with the noise-enhanced samples to generate perturbed samples;
[0090] S4: Input the perturbed samples and real samples into the discriminator, set the adversarial loss function as the objective function; by optimizing the objective function, obtain the preliminarily trained generator and discriminator;
[0091] S5: Select Gaussian noise, input the Gaussian noise and real samples into the preliminarily trained generator, and generate adversarial samples according to the generator loss;
[0092] S6: Input the adversarial samples and real samples into the preliminarily trained discriminator, distinguish the adversarial samples from the real samples, and calculate the discriminator loss;
[0093] S7: Calculate the norm distance between the adversarial samples and real samples, compare it with the preset perturbation range, and set the BIM perturbation loss constraint or hinge loss constraint according to the comparison result;
[0094] S8: Update the parameters of the preliminarily trained generator and discriminator according to the generator loss and discriminator loss, and obtain the preliminarily optimized generator and discriminator;
[0095] S9: Input the adversarial samples and perturbed samples into the distillation model, set the adversarial loss function of the distillation model, and use the dynamic distillation method to iteratively optimize the distillation model with the preliminarily optimized generator and discriminator;
[0096] S10: Establish a total loss function according to the adversarial loss function, BIM perturbation loss constraint, hinge loss constraint, and adversarial loss function; by optimizing the total loss function, obtain the finally optimized distillation model, generator, and discriminator, and form an optimized AdvGAN generative adversarial network;
[0097] S11: Use the optimized AdvGAN generative adversarial network to generate the final adversarial samples.
[0098] In the specific implementation process, in this embodiment, random noise is first added to the real samples to obtain noise-enhanced samples; the noise-enhanced samples are input into the generator to generate perturbed samples; an adversarial loss function is set, and the discriminator is used to distinguish between the perturbed samples and the real samples, so that the initially trained generator and discriminator learn the distribution characteristics of the real samples and the perturbations; then, Gaussian noise is added to the real samples and input into the initially trained generator to generate adversarial samples; the initially trained discriminator is used to distinguish between the adversarial samples and the real samples, and according to the difference between the adversarial samples and the real samples, a BIM perturbation loss constraint or a hinge loss constraint is set to limit the perturbation size of the adversarial samples, obtaining the initially optimized generator and discriminator; finally, a distillation model is set to iteratively optimize with the initially optimized generator and discriminator, so that the initially optimized generator and discriminator can better learn the distribution characteristics of the adversarial samples, and the distillation model is closer to the target black-box model; by optimizing the total loss function, an optimized AdvGAN generative adversarial network is obtained, and the final adversarial samples of the specified target attack category or non-target attack category are generated using it. The present invention improves the generation speed of adversarial samples, improves the quality of the generated adversarial samples; reduces the difference between the distribution characteristics of the adversarial samples and the distribution characteristics of the real samples, increases the difficulty of human recognition, and at the same time improves the attack success rate.
[0099] Embodiment 2
[0100] This embodiment provides an adversarial sample generation method based on AdvGAN, as Figure 2 shown, including:
[0101] S1: Construct an AdvGAN generative adversarial network, including a generator, a discriminator, and a distillation model;
[0102] S2: Obtain real samples, and add random noise to the real samples to obtain noise-enhanced samples; the specific method is:
[0103] S2.1: Obtain real samples \(x\in X\), and the noise distribution of the real samples is \(N(0,\sum * ); where \(X\) represents the real sample data set, and \(\sum * represents the noise matrix;
[0104] S2.2: For each real sample \(x\), select random noise \(n\in N\), add noise constraints and dispersion constraints to generate noise-enhanced samples \(x'\), and obtain a noise-enhanced data set \(X'\);
[0105] The noise constraint is:
[0106]
[0107] The dispersion constraint is:
[0108]
[0109] wherein, X enc,n (x, x′) represents the noise constraint between the real sample x and the noise-enhanced sample x′, and f w (*) represents the feature extraction operation, and argmin represents the value function of x′ when obtaining the minimum value of ||f w+n (x) - f w (x′)||₂; ||*||₂ represents the operation of obtaining the L2 norm; KL(p X′ (x′)||q X (x)) represents the KL divergence between the real sample x and the noise-enhanced sample x′, p(x′) represents the probability density function of the noise-enhanced sample x′, and q(x) represents the probability density function of the real sample x.
[0110] The KL divergence is used to judge the difference in data distributions between the real sample and the noise-enhanced sample, so that the data distribution of the noise-enhanced sample will not differ too much from the data distribution of the original real sample, and they remain consistent in data features.
[0111] S3: Input the noise-enhanced sample into the generator, and use the gradient descent method to calculate the perturbation and combine it with the noise-enhanced sample to generate a perturbed sample; the specific method is as follows:
[0112] As Figure 3 shown, it is the structural schematic diagram of the generator G. The generator G includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a first residual block, a second residual block, a third residual block, a fourth residual block, a fourth convolutional layer, a fifth convolutional layer, and a sixth convolutional layer connected in sequence. The activation functions of the first to sixth convolutional layers are all ReLU functions; randomly select the noise-enhanced sample x′ and input it into the generator G. On the premise of maintaining the data features of the noise-enhanced sample, use the gradient descent method to iteratively calculate the perturbation δ adv :
[0113]
[0114]
[0115]
[0116]
[0117] wherein, g c represents the gradient symbol, represents the expected perturbation at the t-th iteration, w represents the gradient parameter, y represents the feature label, represents the gradient momentum at the t-th iteration; denotes the gradient momentum of the (t + 1)-th iteration, μ denotes the decay factor, a denotes the sign coefficient, [-ε, ε] denotes the preset perturbation range, denotes the expected perturbation of the (t + 1)-th iteration; denotes the perturbation of the (t + 1)-th iteration;
[0118] When and the difference of is within the preset perturbation range, then is used as the perturbation δ adv ; The perturbation δ adv is combined with the noise-enhanced sample x′ to generate a perturbed sample
[0119] After generating the perturbed sample, it is also necessary to screen the perturbed sample:
[0120] Calculate the L2 norm distance between the perturbed sample and the real sample x:
[0121]
[0122] In the formula, L p denotes the L2 norm distance between the perturbed sample and the real sample;
[0123] Set the perturbation threshold ∈, if L p >|∈|, then discard the perturbed sample.
[0124] S4: Input the perturbed sample and the real sample into the discriminator, and set the adversarial loss function as the objective function; By optimizing the objective function, obtain the preliminarily trained generator and discriminator; The adversarial loss function is:
[0125]
[0126] In the formula, L GAN denotes the adversarial loss function, f L (*) denotes a function that satisfies 1-Lipschitz continuity, that is
[0127] As Figure 4 shown, it is the structural schematic diagram of the discriminator D; The discriminator includes a seventh convolutional layer, an eighth convolutional layer, a ninth convolutional layer, and a fully connected layer connected in sequence; The activation functions of the seventh to ninth convolutional layers are LeakyReLU functions, and the activation function of the fully connected layer is the Softmax function.
[0128] S5: Select Gaussian noise, input the Gaussian noise and the real sample into the preliminarily trained generator, and generate adversarial samples according to the generator loss; The specific method is:
[0129] Randomly select m real samples x, select Gaussian noise η, input the real samples and Gaussian noise into the pre-trained generator, initialize the parameters of the pre-trained generator, and generate adversarial samples x adv , and calculate the generator loss:
[0130]
[0131] where, denotes the generator loss, G(*) denotes the pre-trained generator, D(*) denotes the pre-trained discriminator, and N′ denotes the Gaussian noise distribution.
[0132] S6: Input the adversarial samples and real samples into the pre-trained discriminator to distinguish between adversarial samples and real samples, and calculate the discriminator loss; the discriminator loss is:
[0133]
[0134] where, denotes the discriminator loss, G(*) denotes the pre-trained generator, D(*) denotes the pre-trained discriminator, and N′ denotes the Gaussian noise distribution.
[0135] S7: Calculate the norm distance between the adversarial samples and real samples, compare it with the preset perturbation range, and set the BIM perturbation loss constraint or hinge loss constraint according to the comparison result; the specific method is:
[0136] Calculate the L2 norm distance between the adversarial sample x adv and the real sample x:
[0137]
[0138] where, denotes the L2 norm distance between the adversarial sample and the real sample;
[0139] Compare with the set perturbation threshold ∈. If then set the BIM perturbation loss constraint:
[0140]
[0141] where, L B denotes the BIM perturbation loss; m denotes the number of randomly selected real samples, k denotes the number of adversarial samples, and σ denotes the first hyperparameter;
[0142] If then set the hinge loss constraint:
[0143] L hinge = E xmax(0, ||G(x)||₂ - c)
[0144] In the formula, L hinge represents the hinge loss, and c represents the upper limit value.
[0145] S8: According to the generator loss and discriminator loss, update the parameters of the preliminarily trained generator and discriminator to obtain a preliminarily optimized generator and discriminator;
[0146] S9: Input the adversarial samples and perturbation samples into the distillation model, set the adversarial loss function of the distillation model, and use the dynamic distillation method to iteratively optimize the distillation model with the preliminarily optimized generator and discriminator; The specific method is:
[0147] S9.1: Input the adversarial samples and perturbation samples into the distillation model;
[0148] S9.2: Set the distillation model f with the black-box model as the target model, set the target attack class t or not set the attack class, and the adversarial loss function of the distillation model f is:
[0149]
[0150] In the formula, represents the adversarial loss of the distillation model f, and l f represents the first cross-entropy loss;
[0151] S9.3: Use the distillation model f of the (i - 1)-th iteration i-1 to update the generator G of the preliminarily trained i-th iteration i and the discriminator D of the preliminarily trained i-th iteration i ; The update method is:
[0152]
[0153] In the formula, represents the adversarial loss of the distillation model of the (i - 1)-th iteration; τ, represents the second and third hyperparameters;
[0154] S9.4: Use the generator G of the preliminarily trained i-th iteration i to feedback and update the distillation model f of the i-th iteration i ; The update method is:
[0155]
[0156] In the formula, represents the second cross-entropy loss, and f(x adv ) represents the output of the distillation model for the adversarial sample x adv ; Denote the output of the distilled model for the perturbed sample , and \(b(x adv )\) represents the query result of the adversarial sample \(x adv \) from the black-box model. Denote the query result of the perturbed sample from the black-box model.
[0157] S10: Establish a total loss function according to the adversarial loss function, BIM perturbation loss constraint, hinge loss constraint, and adversarial loss function; by optimizing the total loss function, obtain the finally optimized distilled model, generator, and discriminator, and form an optimized AdvGAN generative adversarial network; the total loss function is:
[0158]
[0159] In the formula, \(\alpha\), \(\beta\), and \(\gamma\) represent the first, second, and third proportionality coefficients.
[0160] S11: Use the optimized AdvGAN generative adversarial network to generate the final adversarial sample.
[0161] In this embodiment, a feed-forward generative adversarial network is used. Each time after training, multiple adversarial samples are generated; the speed of generating adversarial samples is improved, and the calculation time is reduced; the generator is optimized through the feedback of the distilled model, so that the generator finally learns the distribution of adversarial samples similar to the real samples. By directly generating perturbations by inputting real samples and combining them into adversarial samples with attack targets, the success rate in attacks targeting black-box models is improved; by setting the BIM perturbation loss constraint and hinge loss constraint to constrain the generation of perturbations, the quality of adversarial samples is improved without affecting the attack effect, and the human perception of them is reduced.
[0162] Embodiment 3
[0163] This embodiment provides an adversarial sample generation method based on AdvGAN. Based on constructing an AdvGAN generative adversarial network, two stages of training processes are set, including a normal training stage and an adversarial training stage; the AdvGAN generative adversarial network includes a generator, a discriminator, and a distilled model;
[0164] In the normal training stage, the generator and the discriminator are trained to learn the data distribution of real samples; use the publicly available dataset MNIST as the real sample dataset \(X\), the real sample \(x\in X\), and the noise distribution of the real sample is \(N(0,\sum * )\). Input the real sample into the feature extractor \(f w \) to extract the distribution features. The normal training stage includes the following steps:
[0165] 1) Select a real sample x from the real sample dataset X, and select a random noise n from the noise distribution N, where n ∈ N, and ∑ * = β1F -1 (w) represents the noise matrix;
[0166] 2) For each real sample x, select a random noise n until the entire real sample dataset X is injected with random noise to generate a noise-augmented sample x'; by adding the noise constraint:
[0167]
[0168] In the formula, X enc,n (x, x') represents the noise constraint between the real sample x and the noise-augmented sample x', f w (*) represents the feature extraction operation, argmin represents the function to obtain the value of x' when minimizing ||f w+n (x) - f w (x')||2; ||*||2 represents the operation of obtaining the L2 norm;
[0169] 3) Use the KL divergence to judge the difference in the data distributions of the real samples and the noise-augmented samples, so that the data distribution of the noise-augmented samples and the noise distribution of the real samples do not differ too much, that is, the data features are kept consistent:
[0170]
[0171] In the formula, L(p X′ (x')||q X (x)) represents the KL divergence between the real sample x and the noise-augmented sample x', p(x') represents the probability density function of the noise-augmented sample x', and q(x) represents the probability density function of the real sample x;
[0172] 4) Initialize the gradient parameter w0, the expected perturbation Randomly initialize the perturbation δ0;
[0173] 5) Randomly select a noise-augmented sample x' from the noise-augmented dataset X' and input it into the generator G, and use the gradient descent method to iterate while maintaining the data features to generate a perturbation δ adv :
[0174]
[0175]
[0176]
[0177]
[0178] In the formula, g c represents the gradient symbol, represents the expected perturbation at the t-th iteration, w represents the gradient parameter, y represents the feature label, represents the gradient momentum at the t-th iteration; represents the gradient momentum at the (t + 1)-th iteration, μ represents the decay factor, a represents the sign coefficient, [-ε, ε] represents the preset perturbation range, represents the expected perturbation at the (t + 1)-th iteration; represents the perturbation at the (t + 1)-th iteration;
[0179] When and the difference of is within the preset perturbation range, then is used as the perturbation δ adv ;
[0180] 6) Combine the perturbation δ adv with the noise-enhanced sample x′ to generate a perturbed sample
[0181] 7) Calculate the gradient penalty term to avoid the phenomenon of gradient explosion in the network:
[0182]
[0183] 8) Use regularization to update the weight parameters:
[0184]
[0185] 9) Calculate the L2 norm distance between the perturbed sample and the real sample x, which is used to determine whether the perturbation exceeds the range:
[0186]
[0187] In the formula, L p represents the L2 norm distance between the perturbed sample and the real sample;
[0188] Set the perturbation threshold ∈, if L p >|∈|, then discard the perturbed sample.
[0189] 10) Set the adversarial loss function:
[0190]
[0191] In the formula, L GAN represents the adversarial loss function, f L (*) represents a function that satisfies 1-Lipschitz continuity, that is
[0192] After each round of training, update the parameters of the generator G and the discriminator D to obtain a preliminarily trained generator and discriminator that have learned the distribution of real samples and the perturbation range, and the normal training phase ends.
[0193] The adversarial training phase includes the following steps:
[0194] 1) Randomly select m real samples x from the real sample dataset X, and select Gaussian noise η ∈ N′(0, ε);
[0195] 2) Use the preliminarily trained generator G and discriminator D as initializations, and initialize θ G as the parameters of the generator, θ D as the discriminator parameters, initialize α G , α D as the learning rates of the generator and discriminator respectively, and the weight parameter w;
[0196] 3) Input the real samples and Gaussian noise into the preliminarily trained generator to generate adversarial samples x adv ;
[0197] 4) Input the adversarial samples x adv and the real samples into the discriminator D to enable the discriminator D to distinguish the difference between the adversarial samples x adv and the real samples x. Calculate the loss of the discriminator
[0198]
[0199] where, represents the discriminator loss, G(*) represents the preliminarily trained generator, D(*) represents the preliminarily trained discriminator, and N′ represents the Gaussian noise distribution.
[0200] 5) At the same time, calculate the loss of the generator:
[0201]
[0202] where, represents the generator loss, G(*) represents the preliminarily trained generator, D(*) represents the preliminarily trained discriminator, and N′ represents the Gaussian noise distribution.
[0203] 6) Calculate the parameters θ D and θ G of the discriminator and the generator for optimizing the generator and the discriminator:
[0204]
[0205]
[0206] 7) Update the weight parameter:
[0207] w ← w + α * ADAM(w, θ D , θ G )
[0208] 8) Calculate the adversarial example x adv The L2 norm distance between the adversarial example x and the real sample x:
[0209]
[0210] In the formula, represents the L2 norm distance between the adversarial example and the real sample.
[0211] 9) Set the perturbation threshold ∈. If it indicates that the perturbation is too small, then set the BIM perturbation loss constraint:
[0212]
[0213] In the formula, L B represents the BIM perturbation loss; m represents the number of randomly selected real samples, k represents the number of adversarial examples, and σ represents the first hyperparameter;
[0214] If it indicates that the perturbation is too large, then set the hinge loss constraint:
[0215] L hinge = E x max(0, ||G(x)||2 - c)
[0216] In the formula, L hinge represents the hinge loss, and c represents the upper limit value.
[0217] 10) Input the adversarial example and the perturbed sample into the distillation model. The adversarial loss function of the distillation model f is:
[0218]
[0219] In the formula, represents the adversarial loss of the distillation model f, and l f represents the first cross-entropy loss;
[0220] Use the method of dynamic distillation to optimize the generator and discriminator. At the same time, it also makes the distillation model closer to the target black-box model, making the adversarial examples generated by the distillation model closer to the attack target category, or the adversarial examples for the black-box non-class attack on the target model. Specifically:
[0221] A) Set the target attack category t or do not set the attack category;
[0222] B) Use the distillation model f of the (i - 1)-th iterationi-1 Update the generator G of the preliminary training in the i-th iteration i and the discriminator D of the preliminary training i ; First, set the initial weights of the generator G i to be the same as those of G in the previous step i-1 , and then update the generator and discriminator in the i-th iteration through the parameters of the distillation model f i-1 in the corresponding (i - 1)-th iteration. The update method is as follows:
[0223]
[0224] In the formula, represents the adversarial loss of the distillation model in the (i - 1)-th iteration; τ, represents the second and third hyperparameters; in this embodiment, τ, are respectively set to 0.5;
[0225] C) Use the generator G of the preliminary training in the i-th iteration i to feedback and update the distillation model f in the i-th iteration i ; First, use the distillation model f i-1 to initialize f i , and then update the distillation model f i by using the query result set of the black-box model for the generated adversarial samples and the features extracted from the original samples in the generator G . The update method is as follows: i In the formula,
[0226]
[0227] In the formula, represents the second cross-entropy loss, f(x adv ) represents the output of the distillation model for the adversarial sample x adv , represents the output of the distillation model for the perturbed sample , b(x adv ) represents the query result of the adversarial sample x adv from the black-box model, represents the query result of the perturbed sample from the black-box model.
[0228] D) Optimize through the real samples with the given attack target to obtain a distillation model f that is very close to the black-box model b, and the generated final adversarial samples can effectively attack the black-box model;
[0229] 11) Set the total loss function as:
[0230]
[0231] In the formula, α, β, and γ represent the first, second, and third proportionality coefficients. α + β + γ = 1. In this embodiment, they are respectively set to 0.3, 0.3, and 0.4.
[0232] Repeat the above steps to optimize the total loss function, and obtain the finally optimized distillation model, generator, and discriminator, which constitute the optimized AdvGAN generative adversarial network.
[0233] 12) Use the optimized AdvGAN generative adversarial network to generate the final adversarial samples.
[0234] This embodiment proposes an adversarial sample generation method based on AdvGAN. Through the coordinated optimization of the distillation model, generator, and discriminator, and training in two stages, it can more effectively extract the features of the picture. When generating adversarial samples of the target class, it can retain the main features of the picture and generate adversarial samples with higher quality and more in line with human vision.
[0235] The same or similar reference numerals correspond to the same or similar components;
[0236] The terms used to describe the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of this patent;
[0237] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. An adversarial sample generation method based on AdvGAN, characterized in that Including: S1: Construct an AdvGAN generative adversarial network, including a generator, a discriminator, and a distillation model; S2: Obtain real samples, and add random noise to the real samples to obtain noise-enhanced samples; S3: Input the noise-enhanced samples into the generator, calculate the perturbation using the gradient descent method, and combine it with the noise-enhanced samples to generate perturbed samples; S4: Input the perturbed samples and real samples into the discriminator, and set the adversarial loss function as the objective function; by optimizing the objective function, obtain the preliminarily trained generator and discriminator; S5: Select Gaussian noise, input the Gaussian noise and real samples into the preliminarily trained generator, and generate adversarial samples according to the generator loss; S6: Input the adversarial samples and real samples into the preliminarily trained discriminator, distinguish the adversarial samples from the real samples, and calculate the discriminator loss; S7: Calculate the norm distance between the adversarial samples and real samples, compare it with the preset perturbation range, and set the BIM perturbation loss constraint or hinge loss constraint according to the comparison result; S8: Update the parameters of the preliminarily trained generator and discriminator according to the generator loss and discriminator loss, and obtain the preliminarily optimized generator and discriminator; S9: Input the adversarial samples and perturbed samples into the distillation model, set the adversarial loss function of the distillation model, and use the dynamic distillation method to iteratively optimize the distillation model with the preliminarily optimized generator and discriminator; S10: Establish the total loss function according to the adversarial loss function, BIM perturbation loss constraint or hinge loss constraint, and adversarial loss function; by optimizing the total loss function, obtain the finally optimized distillation model, generator, and discriminator, and form the optimized AdvGAN generative adversarial network; S11: Use the optimized AdvGAN generative adversarial network to generate the final adversarial samples.
2. The adversarial sample generation method based on AdvGAN according to claim 1, wherein The specific method of step S2 is: S2.1: Obtain a real sample \(x\in X\), where the noise distribution of the real sample is \(N(0,\sum * ); Here, \(X\) represents the real sample dataset, and \(\sum * represents the noise matrix; S2.2: For each real sample x, select random noise n ∈ N, add noise constraint and dispersion constraint to generate the noise-enhanced sample x′, and obtain the noise-enhanced dataset X′; The noise constraint is: The dispersion constraint is: Where X enc,n (x, x′) represents the noise constraint between the real sample x and the noise-enhanced sample x′, and f w (*) represents the feature extraction operation, argmin represents the value function of x′ when obtaining the minimum value of ||f w+n (x) - f w (x′)||2; ||*||2 represents the operation of obtaining the L2 norm; KL(p X′ (x′)||q X (x)) represents the KL divergence between the real sample x and the noise-enhanced sample x′, p(x′) represents the probability density function of the noise-enhanced sample x′, and q(x) represents the probability density function of the real sample x.
3. The adversarial sample generation method based on AdvGAN according to claim 2, wherein The specific method of step S3 is: Randomly select the noise-enhanced sample x′ and input it into the generator G. On the premise of maintaining the data characteristics of the noise-enhanced sample, use the gradient descent method to iteratively calculate the perturbation δ adv : where g c represents the gradient symbol, represents the expected perturbation at the t-th iteration, w represents the gradient parameter, y represents the feature label, represents the gradient momentum at the t-th iteration; represents the gradient momentum at the (t + 1)-th iteration, μ represents the decay factor, a represents the sign coefficient, [-ε, ε] represents the preset perturbation range, represents the expected perturbation at the (t + 1)-th iteration; represents the perturbation at the (t + 1)-th iteration; When and are within the preset perturbation range, is used as the perturbation δ adv ; The perturbation δ adv is combined with the noise-enhanced sample x′ to generate a perturbed sample 4. The adversarial sample generation method based on AdvGAN according to claim 3, wherein, In step S3, after generating the perturbed samples, the perturbed samples also need to be screened. The specific method is: Calculating perturbed samples L2-norm distance from the true sample x: In the formula, represents the L2 norm distance between the perturbed sample and the true sample; Set the perturbation threshold ∈. If L p >|∈|, then discard the perturbed sample.
5. The adversarial sample generation method based on AdvGAN according to claim 4, wherein, In step S4, the adversarial loss function is: where L GAN represents the adversarial loss function, and f L (*) represents a function that satisfies 1-Lipschitz continuity.
6. The adversarial sample generation method based on AdvGAN according to claim 5, wherein, The specific method of step S5 is: Randomly select m real samples x, select Gaussian noise η, input the real samples and Gaussian noise into the pre-trained generator, initialize the parameters of the pre-trained generator, and generate adversarial samples x adv , and calculate the generator loss: Wherein, represents the generator loss, G(*) represents the pre-trained generator, D(*) represents the pre-trained discriminator, and N′ represents the Gaussian noise distribution.
7. The method for generating adversarial samples based on AdvGAN according to claim 6, wherein In step S6, the discriminator loss is: In the formula, represents the discriminator loss, G(*) represents the preliminarily trained generator, D(*) represents the preliminarily trained discriminator, and N′ represents the Gaussian noise distribution.
8. The adversarial sample generation method based on AdvGAN according to claim 7, wherein The specific method of step S7 is: Calculate the adversarial example x adv The L2 norm distance from the true example x: In the formula, represents the L2 norm distance between the adversarial sample and the real sample; Compare with the set perturbation threshold ∈. If then set the BIM perturbation loss constraint: where L B represents the BIM perturbation loss; m represents the number of randomly selected true samples, k represents the number of adversarial samples, and σ represents the first hyperparameter; If then set the hinge loss constraint: L hinge = E x max(0, ||G(x)||2 - c) where L hinge represents the hinge loss, and c represents the upper limit value.
9. The adversarial sample generation method based on AdvGAN according to claim 8, wherein The specific method of step S9 is: S9.1: Input the adversarial samples and perturbed samples into the distillation model; S9.2: Set the distillation model f with the black-box model as the target model, set the target attack class t or not set the attack class, and the adversarial loss function of the distillation model f is: In the formula, represents the adversarial loss of the distilled model f, and l f represents the first cross-entropy loss; S9.3: Update the generator G and discriminator D that are preliminarily trained in the i-th iteration by using the distillation model f of the (i-1)-th iteration. The update method is as follows: i-1 Update the generator G that is preliminarily trained in the i-th iteration i and the discriminator D that is preliminarily trained i ; The update method is: In the formula, represents the adversarial loss of the distilled model in the (i - 1)-th iteration; τ, represents the second and third hyperparameters; S9.4: Use the generator G initially trained in the i-th iteration i Feed back and update the distillation model f in the i-th iteration i ; The update method is as follows: Wherein, represents the second cross-entropy loss, f(x adv ) represents the output of the distilled model for the adversarial sample x adv . represents the output of the distilled model for the perturbed sample , b(x adv ) represents the query result of the adversarial sample x adv from the black-box model, represents the query result of the perturbed sample from the black-box model.
10. The adversarial sample generation method based on AdvGAN according to claim 9, wherein, The total loss function is: In the formula, α, β, γ represent the first, second, and third proportionality coefficients.
Citation Information
Patent Citations
Deep learning adversarial sample generation method based on second-order method
CN111325324A
GAN-based medical diagnosis model anti-attack method
CN113178255A