A method for improving the attack and defense performance of adversarial examples based on the GRAN architecture
Through the adversarial sample generation and repair methods under the GRAN architecture, the problem of lag in the existing defense method is solved, the defense capability of the target model against the adversarial sample is improved, and the stronger robustness is achieved.
Patent Information
- Application Number
- CN202210876725.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-25
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-07-25
AI Technical Summary
The existing adversarial sample defense methods have lag compared to attack methods, making it difficult to effectively resist new adversarial sample attacks.
Design the GRAN network architecture, put adversarial sample generation and repair under the same architecture, and continuously adversarial sample attacks and defenses are carried out through generators and repairers to improve their offensive and defensive performance.
Improves the robustness of the target model against samples, allowing it to more effectively resist stronger adversarial sample attacks.
Smart Images

Figure CN115272793B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence security, and particularly relates to a method for improving the attack and defense performance of adversarial samples based on the GRAN architecture. Background Art
[0002] In recent years, with the continuous maturity of machine learning theory and technology, artificial intelligence technology has been widely applied in many fields such as computer vision, natural language processing, and speech recognition, and has profoundly affected people's daily lives. However, further research has found that there are various security problems in artificial intelligence technology itself, and adversarial samples are currently a research hotspot in the field of artificial intelligence security. By adding subtle perturbations to normal samples, adversarial samples can make the target model return incorrect results. In safety-critical fields such as autonomous driving, industrial control, and intelligent security, the existence of adversarial samples hinders the further application of artificial intelligence technology in these fields. Therefore, it is necessary to conduct in-depth research on the defense methods of adversarial samples.
[0003] Since most current artificial intelligence models have poor interpretability, it is difficult to completely defend against adversarial samples from a mechanism perspective. Effective defense methods mostly retrain the target model based on the generated adversarial samples to improve the robustness of the target model against adversarial samples, that is, the adversarial training method; or learn the pattern of adversarial perturbations from adversarial samples and remove them to restore the adversarial samples to the original samples, that is, the adversarial sample repair method. These methods all perform corresponding defenses after generating adversarial samples, and the defense ability is restricted by the attack ability of existing adversarial samples, and they perform poorly when facing stronger new adversarial sample attacks.
[0004] Reference 1 (Jiang Lingyun, Qiao Kai, Qin Ruoxi, et al. Cycle-Consistent Adversarial GAN: The Integration of Adversarial Attack and Defense [J]. Security and Communication Networks, 2020, v2020) proposed an adversarial sample generation and repair framework based on a cycle generative adversarial network to address the problem of the lagging performance of existing adversarial sample defense methods, and studied adversarial sample attack and defense in the same framework. However, this method uses two cycle generative adversarial networks to generate and repair adversarial samples respectively. The adversarial samples are generated by a neural network, and no defense research is conducted on the adversarial samples generated by traditional methods. Moreover, the two networks of this method are independently trained, and the adversarial game relationship between adversarial sample generation and repair is not considered. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] In view of the problem that current artificial intelligence models are vulnerable to adversarial samples and the lag of existing adversarial sample defense methods compared to attack methods, the present invention provides a method for improving the offensive and defensive performance of adversarial samples.
[0007] (2) Technical solution
[0008] To solve the above technical problems, the present invention provides a method for improving the offensive and defensive performance of adversarial samples based on the GRAN architecture, including the following steps:
[0009] First step: Design the structure of the GRAN network;
[0010] Second step: Based on the designed GRAN network, improve the ability to generate and repair adversarial samples.
[0011] Preferably, in the first step, the GRAN network is designed to consist of the following five parts: target model F, input sample set X, model output set Y, generator G, and repairer R;
[0012] (1) Target model F:
[0013] The target model F is represented as a mapping F: X → Y and is implemented by a deep learning model;
[0014] (2) Input sample set X:
[0015] The original sample set X is composed of all possible inputs of the target model F, and each element is a quantized input sample;
[0016] (3) Model output set Y:
[0017] The model output set Y is composed of all possible results of the target model F;
[0018] (4) Generator G:
[0019] For a target model F, the generator G can be represented as a mapping G: X → X, and its function is to generate an adversarial sample x for any original input sample x ∈ X A ;
[0020] G can be implemented by a functional function, that is, by an adversarial sample generation algorithm or by a generative neural network;
[0021] When G is implemented using an adversarial sample generation algorithm, the target model F and the repairer R need to be used as inputs when generating adversarial samples, that is, it is represented as x A = G(x, t; F, R), where t ∈ Y represents the target of the adversarial sample attack;
[0022] When G is implemented using a generative neural network, the objective function for training the neural network G can be expressed as L G = L GS (x, x A ) + λ G ·L GA (t, y R ), where x A = G(x), representing the adversarial example generated by G; y R = F(R(x A )) represents the model output result obtained by inputting the repaired sample R(x A ) into F after repairing the adversarial example x A ) using the repairer R; L GS , L GA respectively represent a loss function; λ G is the weight factor;
[0023] In the objective function L G , this part of L GS is used to promote the similarity between the adversarial example x A and the original input sample x, and this part of L GA is used to make the adversarial example x A still have an attack effect after being repaired by the repairer R;
[0024] (5) Repairer R:
[0025] For a target model F, the repairer R can be expressed as a mapping R: X → X, and its function is to generate a repaired sample x A from the adversarial example x R ;
[0026] R is implemented using a generative neural network;
[0027] The objective function for training the neural network R can be expressed as L R + L RS (x, x R ) + λ R ·L RA (y, y R ), where y ∈ Y represents the output result of the target model F under the original input sample x, that is, y = F(x); x R = R(x A ), representing the repaired sample generated by R; y R = F(x R ), representing the model output result of the target model F when the input is x R ; L RS , L RA respectively represent a loss function; λ R is the weight factor;
[0028] Objective function L R In it, L RS This part is used to promote the repair sample x R to be close to the original input sample x, L RA This part is used to promote the adversarial sample repair ability improvement of R.
[0029] Preferably, the second step is a method for improving the attack and repair capabilities of adversarial samples based on the GRAN architecture by training the GRAN network.
[0030] Preferably, the second step includes the following steps:
[0031] (1) Given the target model F, the training sample set attack target t, initialize the generator G and the repairer R;
[0032] (2) Select a batch of training data x from X Train and input it into the generator G to obtain the corresponding adversarial sample x A ;
[0033] (2a) If G is implemented by an adversarial sample generation algorithm, regard the target model F and the repairer R as a new target model F(R(·)) to obtain the corresponding adversarial sample x A = G(x, t; F, R), and at this time G does not participate in the parameter update of step (5);
[0034] (2b) If G is implemented by a generative neural network, the adversarial sample x A = G(x);
[0035] (3) Input the adversarial sample x A into the repairer R to obtain the repaired sample x R = R(x A );
[0036] (4) Input the original training data x and the adversarial sample x A into the target model to obtain the output result y = F(x), y A = F(x A );
[0037] (5) Calculate the objective function L G = L GS (x, x A ) + λ G ·L GA (t, y R ), L R = L RS (x, x R ) + λ R ·LRA (y, y R ), perform backpropagation and update the parameters of G and R using an optimization algorithm;
[0038] (6) Repeat steps (2) to (5) until the objective functions L G , L R both converge;
[0039] After the network training is completed, the restorer R can be used to restore the adversarial samples to restored samples that can be correctly processed by the target model. Connect R to the target model F to improve the robustness of the target model against adversarial samples; if G is implemented by a neural network, then generate adversarial samples with a certain attack ability through G; if G is implemented by an adversarial sample generation algorithm, then by connecting the restorer R in front of the target model F and regarding the two networks as a new target model, obtain adversarial samples with stronger attack ability.
[0040] Preferably, the deep learning model includes a convolutional network model for image data, a network model for speech and text data, and a network model for deep reinforcement learning.
[0041] Preferably, the quantized input samples include the pixel matrix of image samples, the Mel frequency cepstral coefficient features of speech samples, the word vectors of text data, and the quantization features of network traffic data.
[0042] Preferably, all possible results of the target model F include the class labels in the classification problem model, the value range of the predicted values in the regression model, and the action space of the reinforcement learning model.
[0043] Preferably, the adversarial sample generation algorithms include FGSM, DeepFool, JSMA, C&W.
[0044] Preferably, the generative neural network is an autoencoder.
[0045] Preferably, the optimization algorithm refers to a first-order gradient optimization algorithm.
[0046] (III) Beneficial Effects
[0047] Based on the idea of attack-defense confrontation game, the present invention proposes a generative repair adversarial network architecture GRAN, and based on this architecture, proposes an adversarial sample generation and repair method. This architecture studies adversarial sample generation and repair under the same architecture, and proposes a method for improving the attack and defense performance of adversarial samples. Continuously perform adversarial sample attacks and defenses on the target model through the generator and the restorer to improve the ability of adversarial sample attacks and defenses. Finally, make the target model more robust against adversarial sample attacks, that is, it can resist stronger adversarial sample attacks. Description of the Drawings
[0048] Figure 1 is the method flow chart of the present invention;
[0049] Figure 2 is the work flow chart of GRAN when the generator G is implemented by a neural network;
[0050] Figure 3 is the work flow chart of GRAN when the generator G is implemented by a certain adversarial sample generation algorithm. Specific Embodiments
[0051] To make the objectives, contents and advantages of the present invention clearer, the following further describes in detail the specific embodiments of the present invention with reference to the accompanying drawings and embodiments.
[0052] Aiming at the problem that current artificial intelligence models are vulnerable to adversarial samples attacks and solving the lag problem of existing adversarial sample defense methods compared with attack methods, the present invention proposes a Generative Repair Adversarial Network (GRAN) based on the idea of attack and defense confrontation game, and proposes an adversarial sample generation and repair method based on this architecture. This architecture studies adversarial sample generation and repair under the same architecture, and proposes a method for improving the attack and defense performance of adversarial samples. Through the generator and the repairer, continuous adversarial sample attacks and defenses are carried out on the target model to improve the capabilities of adversarial sample attacks and defenses. Finally, the target model is more robust when facing adversarial sample attacks, that is, it can resist stronger adversarial sample attacks.
[0053] The present invention proposes the following technical solution: a method for improving the attack and defense performance of adversarial samples based on the GRAN architecture. This solution includes two parts: 1) GRAN network structure: based on the adversarial game principle of the Generative Adversarial Network (GAN), the discriminator D is replaced with the repairer R, and the corresponding objective function is modified to achieve a zero-sum game between the generator G and the repairer R; 2) The adversarial sample generation and repair method based on GRAN: based on the GRAN architecture, the generator G is used to generate adversarial samples, and the repairer R repairs the adversarial samples to achieve a zero-sum game process between adversarial sample generation and repair, mutually improving their respective performances. Finally, the repairer R can repair adversarial samples with stronger attack effects, thereby improving the robustness of the target model when facing adversarial sample attacks.
[0054] The following will introduce each part of the solution in detail.
[0055] (I) Design of the GRAN Network Structure
[0056] This part mainly presents the GRAN network architecture. This network architecture mainly consists of the following five parts: the target model F, the input sample set X, the model output set Y, the generator G, and the repairer R. Each component will be introduced separately below.
[0057] (1) Target model F:
[0058] The target model F can be represented as a mapping F: X → Y, which is implemented by a certain deep learning model, including but not limited to convolutional network models such as ResNet, FasterRCNN, Yolo-v3 commonly used for image data, network models such as LSTM, Transformer, BERT commonly used for speech and text data, and network models such as DQN, A3C, PPO used for deep reinforcement learning.
[0059] (2) Input sample set X:
[0060] The original sample set X consists of all possible inputs of the target model F, where each element is generally a quantized input sample, such as the pixel matrix of an image sample, the Mel-scale Frequency Cepstral Coefficients (MFCC) features of a speech sample, the word vectors of text data, the quantized features of network traffic data, etc.
[0061] (3) Model output set Y:
[0062] The model output set Y consists of all possible results of the target model F, such as the various category labels in a classification problem model, the value range of predicted values in a regression model, the action space of a reinforcement learning model, etc.
[0063] (4) Generator G:
[0064] (4a) For a certain target model F, the generator G can be represented as a mapping G: X → X, and its function is to generate an adversarial sample x for any original input sample x ∈ X A ;
[0065] (4b) G can be implemented by a certain functional function, such as white-box adversarial sample generation algorithms like FGSM, DeepFool, JSMA, C&W, or can also be implemented by a neural network, such as a generative network like an autoencoder;
[0066] (4c) When G is implemented using a certain adversarial sample generation algorithm, the target model F and the repairer R need to be used as inputs when generating adversarial samples, that is, it can be expressed as x A = G(x, t; F, R), where t ∈ Y represents the target of the adversarial sample attack;
[0067] (4d) When G is implemented using a generative neural network, the objective function for training the neural network G can be expressed as L G = L GS (x, x A ) + λ G ·L GA (t, y R ), where x A = G(x), representing the adversarial sample generated by G; y R = F(R(x A )) represents the model output result obtained by inputting the repaired sample R(x A ) into F after repairing the adversarial sample x A ) using the repairer R; L GS , L GA respectively represent certain loss functions, such as cross-entropy loss, mean squared error loss, etc.; λ G is the weight factor;
[0068] (4e) In the objective function L G , the L GS part is used to promote the similarity between the adversarial sample x A and the original input sample x, and the L GA part is used to make the adversarial sample x A still have an attack effect after being repaired by the repairer R.
[0069] (5) Repairer R:
[0070] (5a) For a certain target model F, the repairer R can be expressed as a mapping R: X → X, and its function is to generate a repaired sample x A from the adversarial sample x R ;
[0071] (5b) R is implemented by a neural network, such as a generative network like an autoencoder;
[0072] (5c) The objective function for training the neural network R can be expressed as L R = L RS (x, x R ) + λ R ·L RA (y, y R ), where y ∈ Y represents the output result of the target model F under the original input sample x, that is, y = F(x); x R = R(x A ), representing the repaired sample generated by R; y R = F(x R ), representing the model output result of the target model F when the input is x R ; L RS , L RArespectively represent a certain loss function, such as cross - entropy loss, mean - square error loss, etc.; λ R is the weight factor.
[0073] (5d) The objective function L R In L RS part is used to promote the repaired sample x R to be close to the original input sample x, and L RA part is used to promote the improvement of the adversarial sample repair ability of R.
[0074] (2) Method for improving the generation and repair ability of adversarial samples based on GRAN
[0075] This part mainly realizes the method for improving the attack and repair ability of adversarial samples based on the GRAN architecture by training the GRAN network. This method mainly includes the following steps:
[0076] (1) Given the target model F, the training sample set attack target t, initialize the generator G and the repairer R;
[0077] (2) Select a batch of training data x from X Train and input it into the generator G to obtain the corresponding adversarial sample x A ;
[0078] (2a) If G is implemented by a certain adversarial sample generation algorithm, regard the original target model F and the repairer R as a new target model F(R(·)) to obtain the corresponding adversarial sample x A = G(x, t; F, R), and at this time G does not participate in the parameter update in step (5);
[0079] (2b) If G is implemented by a generative neural network, the adversarial sample x A = G(x);
[0080] (3) Input the adversarial sample x A into the repairer R to obtain the repaired sample x R = R(x A );
[0081] (4) Input the original training data x, the adversarial sample x A into the target model to obtain the output results y = F(x), y A = F(x A );
[0082] (5) Calculate the objective function L G = L GS (x, x A ) + λ G ·L GA (t, yR ),L R = L RS (x, x R ) + λ R ·L RA (y, y R ), perform backpropagation, and update the parameters of G and R using a certain optimization algorithm, where the optimization algorithm refers to a first-order gradient optimization algorithm, including Stochastic Gradient Descent (SGD), Root Mean Square Prop (RMSProp), Adaptive Moment Estimation (Adam), etc.;
[0083] (6) Repeat steps (2) to (5) until the objective functions L G , L R both converge.
[0084] (7) After the network training is completed, the restorer R can be used to restore the adversarial sample to a restored sample that the target model can correctly process. Connecting R to the target model F can improve the robustness of the target model against adversarial samples; if G is implemented by a neural network, then adversarial samples with strong attack capabilities can be quickly generated through G; if G is implemented by a certain adversarial sample generation algorithm, then by connecting the restorer R in front of the target model F and treating the two networks as a new target model, adversarial samples with stronger attack effects can be obtained.
[0085] The technical solution of the present invention proposes the GRAN architecture and designs a method for improving the attack capabilities of adversarial samples and the defense capabilities of the target model based on GRAN. The core of this technical solution is to propose the GRAN architecture for improving the offensive and defensive performance of adversarial samples. The GRAN architecture has the following advantages:
[0086] First, when G is represented by a neural network, the relationship between the generator G and the repairer R in the GRAN architecture is similar to the relationship between the generator and the discriminator in the GAN (Generative Adversarial Network) architecture: 1) From the perspective of the objective function, the update of the generator network parameters in both architectures depends on the output of the repairer or the discriminator, and the update of the repairer or discriminator network parameters depends on the output of the generator; 2) The purpose of the generator in GRAN is to generate adversarial samples that still have an attack effect after being repaired by the repairer, and the purpose of the repairer is to repair the adversarial samples into repaired samples that can be correctly processed by the target model. The adversarial game relationship between the two is similar to that of GAN. Based on the idea of adversarial game, the GRAN architecture improves the adversarial sample generation ability and defense ability of both the generator and the repairer simultaneously. When the training of the two networks converges, the generator can generate more aggressive adversarial samples for testing the robustness of the target model or the defense ability of other defense methods; the repairer can repair the adversarial samples into repaired samples that can be correctly processed by the target model to improve the robustness of the target model when facing more aggressive adversarial samples.
[0087] When G is implemented using a certain adversarial sample generation algorithm, the GRAN architecture only needs to train the repairer R, and the generator G regards the original objective function F and the repairer R as a new target model to generate corresponding adversarial samples. As the repair ability of R improves, the attack ability of the adversarial samples generated by G will also increase. When the training of R converges, on the one hand, it can repair the adversarial samples to improve the defense ability of the target model; on the other hand, it can use R to obtain more aggressive traditional adversarial samples based on the adversarial sample generation algorithm to test the robustness of the target model or the defense ability of other defense methods.
[0088] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principles of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A method for improving the attack and defense performance of adversarial samples based on the GRAN architecture, characterized in that It includes the following steps: The first step is to design the structure of the GRAN network; The second step is to improve the generation and repair capabilities of adversarial samples based on the designed GRAN network; In the first step, the GRAN network is designed to consist of the following five parts: the target model F, the input sample set X, the model output set Y, the generator G, and the repairer R; (1) Target model F: The target model F is represented as a mapping F: X → Y and is implemented by a deep learning model; (2) Input sample set X: The original sample set X is composed of all possible inputs of the target model F, where each element is a quantized input sample; (3) Model output set Y: The model output set Y is composed of all possible results of the target model F; (4) Generator G: For a target model F, the generator G can be represented as a mapping G: X → X, whose function is to generate an adversarial sample x for any original input sample x ∈ X A ; G can be implemented by a functional function, that is, by an adversarial sample generation algorithm or by a generative neural network; When G is implemented using an adversarial example generation algorithm, the target model F and the fixer R need to be used as inputs when generating adversarial examples, which is expressed as x A = G(x, t; F, R), where t ∈ Y represents the target of the adversarial example attack; When G is implemented using a generative neural network, the objective function for training the neural network G can be expressed as L G = L GS (x, x A ) + λ G ·L GA (t, y R ), where x A = G(x) represents the adversarial sample generated by G; y R = F(R(x A )) represents the output result of the model obtained by inputting the repaired sample R(x A ) into F after repairing the adversarial sample x A ) using the repairer R; L GS , L GA represent a loss function respectively; λ G is the weight factor; Objective function L G In L GS This part is used to promote the adversarial sample x A to be close to the original input sample x, and L GA this part is used to make the adversarial sample x A still have an attack effect after being repaired by the repairer R; (5) Repairer R: For a target model F, the fixer R can be represented as a mapping R: X → X, whose function is to generate a fixed sample x from the adversarial sample x A ; R ; R is implemented by a generative neural network; The objective function for training neural network R can be expressed as L R = L RS (x, x R ) + λ R ·L RA (y, y R ), where y ∈ Y represents the output result of the target model F under the original input sample x, that is, y = F(x); x R = R(x A ), representing the repaired sample generated by R; y R = F(x R ), representing the model output result of the target model F when the input is x R ; L RS , L RA respectively represent a loss function; λ R is the weight factor; Objective function L R In L RS This part is used to promote the repaired sample x R to be close to the original input sample x. L RA This part is used to promote the adversarial sample repair ability of R The deep learning model includes a convolutional network model for image data, a network model for speech and text data, and a network model for deep reinforcement learning; The quantized input samples include the pixel matrix of image samples, the Mel frequency cepstral coefficient features of speech samples, the word vectors of text data, and the quantized features of network traffic data.
2. The method according to claim 1, wherein The second step is a method for improving the attack and repair capabilities of adversarial samples based on the GRAN architecture by training the GRAN network.
3. The method according to claim 1, wherein The second step includes the following steps: (1) Given a target model F and a training sample set attack target t, initialize the generator G and the repairer R; (2) Select a batch of training data x from X Train and input it into the generator G to obtain the corresponding adversarial sample x A ; (2a) If G is implemented by an adversarial sample generation algorithm, the target model F and the repairer R are regarded as a new target model F(R(·)) to obtain the corresponding adversarial sample x A = G(x, t; F, R), and at this time, G does not participate in the parameter update of step (5); (2b) If G is implemented by a generative neural network, then the adversarial sample x A = G(x); (3) Input the adversarial example \(x\) A into the repairer \(R\) to obtain the repaired example \(x\) R \(= R(x\) A ); (4) Input the original training data x and the adversarial example x A into the target model to obtain the output results y = F(x), y A = F(x A ); (5) Calculate the objective function L G = L GS (x, x A ) + λ G ·L GA (t, y R ), L R = L RS (x, x R ) + λ R ·L RA (y, y R ), perform backpropagation, and use an optimization algorithm to update the parameters of G and R; (6) Repeat steps (2) to (5) until the objective functions L G , L R both converge; After the network training is completed, the repairer R can be used to repair the adversarial samples into repair samples that can be correctly processed by the target model. Connect R to the target model F to improve the robustness of the target model against adversarial samples; if G is implemented by a neural network, then generate adversarial samples with a certain attack ability through G; if G is implemented by an adversarial sample generation algorithm, then connect the repairer R in front of the target model F and regard the two networks as a new target model to obtain adversarial samples with stronger attack ability.
4. The method according to claim 1, characterized in that All possible results of the target model F include the category labels in the classification problem model, the value range of the predicted values in the regression model, and the action space of the reinforcement learning model.
5. The method according to claim 1, wherein The adversarial sample generation algorithms include FGSM, DeepFool, JSMA, and C&W.
6. The method according to claim 1, characterized in that, The generative neural network is an autoencoder.
7. The method according to claim 3, wherein The optimization algorithm refers to the first-order gradient optimization algorithm.
Citation Information
Patent Citations
Image restoration method based on deep learning
CN112669224A
Defect detection method based on semi-supervised learning
CN113808035A