A Method for Improving the Generalization Ability of Adversarial Training Based on Fisher Information

Optimized weighted coefficients to generate optimal generalized adversarial images through Fisher information and particle swarm algorithm, solving the problem of insufficient generalization ability of existing adversarial training methods and achieving effective defense against various and unknown adversarial attack methods.

CN115984667BActive Publication Date: 2025-07-22TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310008684.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-04
Publication Date
2025-07-22
Estimated Expiration
2043-01-04

AI Technical Summary

Technical Problem

The existing adversarial training methods are insufficient in generalization when facing different and unknown adversarial attack methods and cannot effectively defend against multiple types of adversarial images.

Method used

The generalization ability of the adversarial image is calculated through Fisher information, and the weighting coefficient is optimized using the particle swarm algorithm to generate the optimal generalization adversarial image for adversarial training, improving the generalization ability of the model.

Benefits of technology

The trained model not only maintains high classification accuracy under known adversarial attack methods, but also has strong generalization to the images generated by unknown adversarial attack methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984667B_ABST
    Figure CN115984667B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for improving the generalization ability of adversarial training based on Fisher information, comprising the following steps: S1, generating different adversarial images from clean images through different adversarial attack methods; S2, linearly superimposing different adversarial images through Fisher information to obtain a generalized adversarial image; S3, optimizing the weighted coefficients of the generalized adversarial image through a particle swarm algorithm to obtain an optimal generalized adversarial image; S4, using the optimal generalized adversarial image for adversarial training to obtain a classification model with stronger generalization ability. The present invention can not only maintain the classification accuracy of traditional adversarial attack methods on clean images, but also has strong generalization ability for adversarial images generated by various and unknown adversarial attack methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of adversarial security technology, and particularly to a method for improving the generalization ability of adversarial training based on Fisher information. Background Art

[0002] Although deep neural networks (DNNs) have been widely applied in fields such as images, videos, and texts with excellent advantages, DNN-based models perform poorly on adversarial images. An attacker can mislead the output category of a classifier by superimposing imperceptible adversarial perturbations on clean samples. Once adversarial images are used to attack DNN systems with high security requirements, adverse consequences will occur. For example, adversarial images can cause an autonomous driving vehicle to recognize a "stop" sign as a "speed limit"; maliciously modified videos can cause a video classification system to misclassify a "robbery" behavior as "rope skipping". Therefore, it is crucial to improve the robustness and generalization of DNN-based models.

[0003] To be able to defend against adversarial images, a large number of adversarial defense methods have emerged, including denoising methods using encoders and decoders, compression methods using image compression to eliminate the influence of adversarial perturbations, smoothing methods using random smoothing for robust authentication of regions around samples, etc. However, the earlier proposed adversarial training is still the best among all defense methods, although adversarial training requires a large amount of training cost. Adversarial training mainly improves the robustness of the model to adversarial images by introducing adversarial images into the training process, enabling the classifier to learn the features of adversarial images. However, existing adversarial training methods all adopt a single adversarial image generation method, and the trained model is only effective for adversarial images generated by the adversarial attack method used during training, and does not have generalization for multiple or even unknown adversarial images.

[0004] Prior art one proposed a deep learning adversarial training method based on data augmentation. This method performs multiple data augmentations on clean data samples, generates adversarial attack samples, and trains the model together. Finally, the trained classification model can improve the classification accuracy of the classification model for adversarial images and alleviate the overfitting phenomenon of traditional adversarial training methods. However, the model trained by this solution is only effective for adversarial images generated by a single adversarial attack method and does not have generalization for different types of adversarial images.

[0005] Prior Art II proposes a method for defending adversarial examples based on the combination of image preprocessing and adversarial training. This method first performs DCT transform (Discrete Cosine Transform) on clean images and adversarial images, and designs a quantization table. Then, it compresses the clean images with different compression ratios and adds noise for training to obtain adversarial images. Finally, it trains multiple classifiers at different compression ratios and obtains the classification results through voting. This scheme improves the classification accuracy, but requires a large amount of computational effort and has poor generalization ability for unknown adversarial images.

[0006] Prior Art III proposes an adversarial example defense method based on saliency adversarial training. This method uses the Project Gradient Descent (PGD) method to generate adversarial images, and divides the saliency map of the adversarial images into several small blocks. It compresses the images by calculating the average saliency value of each small block, thereby training the model. This method requires image compression first when inputting images, which improves the robustness of the model against adversarial images. However, saliency compression requires prior information, and this defense method does not have generalization ability for different and unknown types of adversarial images.

[0007] It should be noted that the information disclosed in the above background art section is only used for understanding the background of the present application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0008] The main objective of the present invention is to solve the problem of weak generalization ability of existing adversarial training techniques under different and unknown adversarial attack methods.

[0009] To this end, the present invention proposes a method for improving the generalization ability of adversarial training based on Fisher information, including the following steps: S1, generating different adversarial images from clean images through different adversarial attack methods; S2, linearly superimposing different adversarial images through Fisher information to obtain a generalized adversarial image; S3, optimizing the weighted coefficients of the generalized adversarial image through the particle swarm algorithm to obtain the optimal generalized adversarial image; S4, using the optimal generalized adversarial image for adversarial training to obtain a classification model with stronger generalization ability.

[0010] In some embodiments of the present invention, in step S1, the different adversarial attack methods are selected from two or more of the Fast Gradient Sign Method (FGSM), the Project Gradient Descent method PGD, C&W, JSMA, and DeepFool; the types of the different adversarial attack methods are selected from different types of adversarial attack methods or a larger number of the same type of adversarial attack methods.

[0011] In some embodiments of the present invention, in step S2, the generalization ability of the Fisher information calculation adversarial image is specifically as follows: Given an image classifier f, f(y|x) represents the distribution of the output label y of the classifier when the input clean image is x; at this time, the KL divergence of the output distributions after the clean image x and the image x+η after adding the perturbation η pass through the classifier can be expressed as:

[0012]

[0013] where G x represents the Fisher information matrix.

[0014] In some embodiments of the present invention, the objective function for restricting the perturbation η in the KL divergence is:

[0015]

[0016] where ε represents the perturbation restriction range.

[0017] In some embodiments of the present invention, the objective function for restricting the perturbation η in the KL divergence by using the Lagrange multiplier method is optimized to obtain the optimization variable G x η = λη, where λ is the Lagrange multiplier.

[0018] In some embodiments of the present invention, the optimization variable further optimizes the objective function of converting the perturbation η in the KL divergence into the maximum eigenvalue of the Fisher information matrix G x as:

[0019]

[0020] where where s represents the softmax output of the model f, and G s is a positive definite matrix; p i (s) represents the confidence score that the classifier outputs the i-th category after passing through the softmax, and M represents the number of classification categories.

[0021] In some embodiments of the present invention, in step S2, the linear superposition is to linearly superpose the adversarial images x1, x2,..., x N generated by using N different adversarial methods through the particle positions to obtain the generalized adversarial image where α i is the weighting coefficient, satisfying

[0022] In some embodiments of the present invention, in step S2, the generalized adversarial image after the linear superposition is expressed as:

[0023]

[0024] Among them β1, β2, ..., β N-1 represent the N-1 dimensional particle positions corresponding to the clean image.

[0025] In some embodiments of the present invention, in step S3, for optimizing the fitness function of different particles of the generalized adversarial image by the particle swarm algorithm, the fitness function to be optimized is expressed as:

[0026]

[0027] where B is the batch size and M is the number of classification categories, represents the weighted sum of adversarial images generated by the j-th image in a batch through different adversarial attack methods; the optimization variable α of the fitness function to be optimized, the weighted sum of the adversarial images of the fitness function to be optimized; i (i = 1, 2, ..., N) corresponds to β in the particle swarm algorithm i (i = 1, 2, ..., N-1).

[0028] In some embodiments of the present invention, in step S4, the performing of the adversarial training includes the following steps:

[0029] S4-1. Calculate the fitness function to be optimized for different particles of the optimal generalized adversarial image, and save the current optimal position of each particle and the global optimal position among all particles; S4-2. Perform multiple search iterations until the fitness function to be optimized converges or reaches the maximum number of iterations; S4-3. Calculate a batch of optimal generalized adversarial images through the global optimal position in the particles; S4-4. Calculate the gradient of the entire batch with respect to the classification model, and perform gradient descent to optimize the classification model parameters;

[0030] S4-5. Perform adversarial training on different batches; perform multi-generation training on all training set data;

[0031] S4-6. When the maximum number of training generations is reached, output a more generalized classification model.

[0032] The present invention has the following beneficial effects:

[0033] A method for improving the generalization ability of adversarial training based on Fisher information proposed by the present invention. By proposing a generalization ability index of Fisher information, adversarial images generated by different common adversarial attack methods are weighted and superimposed, and the weighting coefficients are optimized by a particle swarm algorithm to obtain the adversarial image with the optimal generalization ability. Finally, adversarial training is carried out with the optimal generalization adversarial image. The present invention can not only maintain the classification accuracy of traditional adversarial attack methods on clean images, but also has strong generalization ability for adversarial images generated by various and unknown adversarial attack methods. Brief Description of the Drawings

[0034] Figure 1 It is the flowchart of the working process in the embodiment of the present invention;

[0035] Figure 2(a) is the initial clean image in the Cifar10 dataset in the embodiment of the present invention;

[0036] Figure 2(b) is the adversarial image generated by using the fast gradient sign method in the embodiment of the present invention;

[0037] Figure 2(c) is the adversarial image generated by using the projected gradient descent method in the embodiment of the present invention;

[0038] Figure 2(d) is the adversarial image generated by using the C&W adversarial attack method in the embodiment of the present invention;

[0039] Figure 2(e) is the optimal generalization adversarial image obtained by optimizing Figure 2(a) in the embodiment of the present invention;

[0040] Figure 3 It is the simulation result of adversarial training on the Cifar10 dataset in the embodiment of the present invention;

[0041] Figure 4 It is the flowchart of Embodiment 1 of the present invention. Detailed Embodiment

[0042] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that only the parts related to the present invention are shown in the drawings for the convenience of description, rather than all the structures.

[0043] In the field of adversarial defense, compared with methods such as image denoising, image compression, and random smoothing, adversarial training is considered to be one of the most effective methods for defending against adversarial examples. Although existing research can defend against adversarial examples through adversarial training and improve the robustness of the model, however, the method of adversarial training is only effective for the adversarial example generation methods used during training and has insufficient generalization ability for various and unknown adversarial example generation methods. Therefore, to address the problem of poor generalization of existing adversarial training methods to various and unknown adversarial examples, this solution proposes a method for enhancing the generalization ability of adversarial training based on Fisher Information, uses Fisher Information to calculate the generalization level of the model for different types of adversarial examples, and generates more representative generalized adversarial examples for adversarial training, thereby improving the model's generalization ability to different and unknown types of adversarial examples.

[0044] The following embodiments of the present invention disclose a method for enhancing the generalization ability of adversarial training based on Fisher Information, belonging to the field of adversarial defense, and solving the problem of weak generalization ability of existing adversarial training technologies under different and unknown adversarial attack methods. The following embodiments of the present invention first convert the maximization of the KL divergence of the output distributions of clean images and perturbed images after passing through the classifier into the minimization of the trace of the Fisher Information matrix under softmax by means of Fisher Information, and use it to evaluate the generalization ability of adversarial images. Then, the adversarial images generated by different adversarial attack types are weighted and superimposed, and the particle swarm algorithm is used to optimize the weighting coefficients to obtain the adversarial image with the optimal generalization ability. Finally, the optimal generalized adversarial image is used for adversarial training. The trained classification model has generalization ability for adversarial examples generated by various and unknown adversarial attack methods.

[0045] The following embodiments of the present invention use a generalization ability measurement index based on Fisher Information to evaluate the generalization ability of adversarial images, and optimize the optimal generalized adversarial image for adversarial training, aiming to propose a method for enhancing the generalization ability of adversarial training based on Fisher Information. The model trained in the following embodiments of the present invention has generalization ability for adversarial images generated by various and unknown adversarial attack methods.

[0046] The following embodiments of the present invention provide a method for improving the generalization ability of adversarial training based on Fisher information. The embodiments of the present invention are mainly divided into three parts. The first is that this solution proposes to use Fisher information to measure the generalization ability of adversarial images. The second is that this solution uses a variety of adversarial attack methods to generate adversarial images, and optimizes the weighting coefficients between adversarial images through Fisher information to obtain the adversarial image with the optimal generalization ability. The third is that this solution uses the generated optimal adversarial images to perform adversarial training on the classification model, and the trained model can improve the generalization ability for different types of adversarial samples.

[0047] The overall flowchart of the embodiments of the present invention is as Figure 1 shown. For clean images, first use a variety of different attack methods to generate multiple adversarial images, then use Fisher information to calculate the generalization ability of the adversarial images, and use the particle swarm optimization algorithm to optimize the weighting coefficients of the adversarial images. Finally, the optimized optimal generalization adversarial images are used for adversarial training to obtain an adversarial defense model with stronger generalization performance for different and unknown types of adversarial samples.

[0048] I. Generalization ability measurement index

[0049] The generalization abilities of adversarial samples generated by different adversarial attack methods are different. During the adversarial training process, if adversarial samples with stronger generalization ability are used for training, it will help the model learn more extensive adversarial sample information, thereby improving the generalization ability of the trained model for different types of adversarial samples. Therefore, it is necessary to propose a measurement index for measuring the generalization ability of adversarial samples. The embodiments of the present invention propose to use Fisher information to calculate the generalization ability of adversarial samples. Specifically, given an image classifier f, f(y|x) represents the distribution of the classifier outputting label y when the input clean image is x; at this time, the KL divergence (Kullback-Leibler Divergence) of the output distributions of the clean image x and the image x + η after adding perturbation η after passing through the classifier can be expressed as:

[0050]

[0051] where G x represents the Fisher information matrix.

[0052] Since adversarial samples are expected to make the classifier output labels completely different from those of clean samples, for attackers, the larger the KL divergence, the better. On the other hand, the adversarial perturbation needs to be as unobservable to the human eye as possible, so the perturbation η needs to be restricted. At this time, the objective function can be expressed as:

[0053]

[0054] where ε represents the perturbation limit range.

[0055] Using the Lagrange multiplier method to optimize the objective function, G can be obtained x η = λη, where λ is the Lagrange multiplier. From the definition of the spectral norm, the optimization variable can be transformed from η to the largest eigenvalue of the Fisher information matrix G x of.

[0056] Define where s represents the softmax output of the model f, and G s is a positive definite matrix. Since when the largest eigenvalue of G s is very large, the perturbation η is more likely to misclassify the classifier. Therefore, to reduce the computational complexity, minimizing the largest eigenvalue of G s can be transformed into minimizing the largest eigenvalue of G s Moreover, since minimizing the largest eigenvalue of a positive definite matrix can be transformed into minimizing the trace of the matrix, the objective function is further transformed into:

[0057]

[0058] where p i (s) represents the confidence score that the classifier outputs the i-th class after softmax, and M represents the number of classification categories.

[0059] In summary, maximizing the KL divergence between the output distributions of clean samples and adversarial samples can be transformed into minimizing In the extreme case, if the adversarial sample has the optimal attack ability, the prediction result of the classifier for the adversarial sample should be close to random guessing. In other words, when the confidence scores of each class in the classifier output are equal, is minimized, which is consistent with the conclusion that the adversarial sample can prevent the classifier from correctly classifying. Since the stronger the generalization ability of the adversarial sample, the more representative it is for different types of adversarial attacks, therefore, can be used to measure the generalization ability of the adversarial sample.

[0060] The generalization ability measurement index adopted in the embodiments of the present invention can be used to evaluate the generalization performance of adversarial samples generated by different adversarial attack methods.

[0061] II. Generalized sample generation

[0062] Since the situation where the confidence scores of each class in the classifier output are equal is too extreme, by simply minimizing Obtaining adversarial examples will lead to a decrease in the generalization of the examples to traditional adversarial attack methods. To obtain the optimal generalized adversarial image that better conforms to the adversarial attack paradigm, the embodiments of the present invention use the weighted sum of adversarial images generated by three traditional adversarial attack methods to obtain the generalized adversarial image. Specifically, the fast gradient sign method, the projected gradient descent method, and the C&W method are used to generate adversarial images x1, x2, and x3, and they are linearly superimposed to obtain the generalized adversarial image where α1, α2, and α3 are weighting coefficients that satisfy α1 + α2 + α3 = 1. Then, the particle swarm algorithm is used to optimize the weighting coefficients to obtain the optimal generalized adversarial image. The fitness function to be optimized can be expressed as:

[0063]

[0064] The three adversarial attack methods used in the embodiments of the present invention can be replaced with other advanced adversarial attack methods, and the number of attack methods can also be increased, which all contribute to improving the optimized generalized adversarial image and have no impact on other steps in the solution.

[0065] The iteration termination condition of the particle swarm optimization in the embodiments of the present invention can be flexibly adjusted to reduce the time of adversarial training.

[0066] III. Adversarial Training

[0067] After the generalized adversarial examples are generated, they will be used for adversarial training to improve the robustness of the classifier to adversarial examples and the generalization to different types of adversarial examples. Specifically, during the training process, to reduce the computational amount, given a batch of images, the embodiments of the present invention select the same set of weighting coefficients for a batch of images. At this time, the fitness function to be optimized is modified as:

[0068]

[0069] where B is the batch size, represents the weighted sum of the adversarial images generated by the three adversarial attack methods for the j-th image in a batch.

[0070] The embodiments of the present invention use the above fitness function to optimize the optimal weighting coefficients for each batch of images and perform adversarial training on the classification model. At this time, the trained model will have generalization to different types of adversarial attack methods. Since the embodiments of the present invention use Fisher information to optimize the optimal generalized adversarial image, the features learned by the model training can ensure its generalization to unknown adversarial attack methods.

[0071] The technology in the embodiments of the present invention is applicable to the field of countermeasure defense. The proposed generalization ability measurement index helps to find the adversarial samples with the optimal generalization ability. In addition, the model trained by the proposed adversarial training method not only has good defense capabilities against various known adversarial attack methods, but also has generalization ability for unknown adversarial samples.

[0072] As Figure 4 shown, the execution process of Embodiment 1 is as follows:

[0073] S1. Generate different adversarial images from clean images through different adversarial attack methods;

[0074] S2. Linearly superimpose different adversarial images through Fisher information to obtain a generalized adversarial image;

[0075] S3. Optimize the weighting coefficients of the generalized adversarial image through the particle swarm algorithm to obtain the optimal generalized adversarial image;

[0076] S4. Use the optimal generalized adversarial image for adversarial training to obtain a classification model with stronger generalization ability.

[0077] The execution process of Embodiment 2 is as follows:

[0078] (1) Input a clean image;

[0079] (2) Generate adversarial images using three adversarial attack methods and perform weighted superposition;

[0080] (3) Optimize the weighting coefficients according to the fitness function to obtain the optimal generalized adversarial image;

[0081] (4) Use the optimal generalized adversarial image for adversarial training.

[0082] Embodiment 3:

[0083] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the protection scope of the present invention.

[0084] The application principle of the present invention will be described in detail below with reference to the accompanying drawings.

[0085] In this embodiment, the dataset is divided into a training set and a test set in a ratio of 7:3; for each batch of clean images in the training set, three adversarial attack methods, namely FGSM, PGD, and C&W, are used to attack the clean images and obtain the corresponding adversarial images. The step size of FGSM is 0.05, the step size of PGD is 0.007, the perturbation threshold is 0.031, the number of iterations is 10, and the number of searches for C&W is 10. Figures 2(a), 2(b), 2(c), 2(d), and 2(e) show a clean image in the Cifar10 dataset and its corresponding adversarial images. Figure 2(a) is the initial clean image, and Figures 2(b), 2(c), and 2(d) are the adversarial images generated by the three adversarial attacks of FGSM, PGD, and C&W respectively. Figure 2(e) shows the optimal generalized adversarial image obtained by optimizing the clean image in Figure 2(a).

[0086] Initialize all parameters of the particle swarm optimization algorithm. The number of particles is 10, the dimension of the particle position is twice the batch size, the maximum number of iterations is 100, the inertia weight is 0.8, the learning factor is 2, and the positions and velocities of all particles are randomly initialized between 0 and 1.

[0087] Use the particle positions to linearly superimpose the three adversarial images x1, x2, and x3 generated by the adversarial attacks to generate a generalized adversarial image x gen . The generalized adversarial image after the linear superposition can be expressed as:

[0088]

[0089] where β1 and β2 respectively represent the particle positions of the two dimensions corresponding to the clean image.

[0090] Calculate the fitness function of a batch of images. The fitness function to be optimized can be expressed as:

[0091]

[0092] where B is the batch size and M is the number of classification categories, represents the weighted sum of the adversarial images generated by the three adversarial attack methods for the j-th image in a batch.

[0093] The optimization variable α i (i = 1, 2, 3) corresponds to β1 and β2 in the particle swarm algorithm.

[0094] Calculate the fitness functions of different particles, and save the current best position of each particle and the global best position among all particles.

[0095] Perform multiple search iterations until the fitness function converges or the maximum number of iterations is reached.

[0096] Calculate the optimal generalization adversarial images for a batch through the global optimal particle positions. Figure 2(e) shows the optimal generalization adversarial images obtained by optimizing the clean image in Figure 2(a).

[0097] Calculate the gradient of the entire batch with respect to the classification model and perform gradient descent to optimize the classification model parameters; perform adversarial training for different batches; perform multi-generation training on all training set data.

[0098] When the maximum number of training generations is reached, output a more robust classification model. Figure 3 Shows the classification accuracy of the test set data during 50 generations of adversarial training simulation on the Cifar10 dataset, where the abscissa represents the number of iterations and the ordinate represents the accuracy. The solid line represents the classification accuracy on clean images, and the dashed line represents the classification accuracy on adversarial images generated by randomly selecting an unknown adversarial attack method. The method for improving the generalization ability of adversarial training based on Fisher information can not only maintain the classification accuracy of traditional adversarial attack methods on clean images, but also has strong generalization ability for adversarial samples generated by unknown adversarial attack methods.

[0099] In view of the problem that the existing adversarial training methods are only effective for the adversarial samples used during training and have insufficient generalization ability for various and unknown adversarial sample generation methods, the embodiment of the present invention proposes a method for improving the generalization ability of adversarial training based on Fisher information, which uses the Fisher information matrix to optimize the generalization ability of adversarial samples, thereby improving the generalization ability of the model after adversarial training for various and unknown adversarial sample generation methods. Specifically, the embodiment of the present invention transforms the adversarial image optimization problem into the trace of the Fisher information matrix under softmax through Fisher information and uses it to measure the generalization performance of adversarial images. Then, the embodiment of the present invention performs weighted superposition on the adversarial images generated by various different adversarial attack methods and uses the particle swarm algorithm to optimize the weighting coefficients to obtain the adversarial images with the optimal generalization ability. Finally, use the optimal generalization adversarial images for adversarial training to obtain a classification model with stronger generalization ability.

[0100] The embodiment of the present invention helps to find the adversarial samples with the optimal generalization ability, and the trained classification model has generalization ability for adversarial samples generated by various and unknown adversarial attacks.

[0101] The embodiment of the present invention also has the following characteristics:

[0102] 1. Generalization ability index based on Fisher information: To improve the generalization ability of adversarial samples used in adversarial training, the embodiments of the present invention propose a generalization ability index based on Fisher information, which transforms maximizing the KL divergence of the output distributions of clean images and images after adding perturbations through a classifier into minimizing the trace of the Fisher information matrix under softmax, thereby obtaining adversarial samples with stronger generalization ability. Currently, none of the adversarial defense methods consider generalized adversarial samples.

[0103] 2. Optimizing the optimal generalized adversarial samples: Currently, the methods of adversarial training all use a single attack method to generate adversarial samples, and the classification models trained in this way do not have generalization ability against other types of attacks. To enable the model obtained by adversarial training to defend against various types of adversarial attacks, the embodiments of the present invention perform weighted superposition on the adversarial samples generated by three common adversarial attack methods, and optimize the weighting coefficients through a particle swarm algorithm to obtain the adversarial samples with the optimal generalization ability.

[0104] 3. Using the optimal generalized adversarial samples for adversarial training: To shorten the training cost, the embodiments of the present invention use the same set of weighting coefficients for the same batch of samples, and use the mean value of the generalization index as the fitness function. Then, use the optimized optimal generalized adversarial samples for adversarial training. The trained classification model can not only defend against different types of adversarial attack samples, but also has generalization against unknown adversarial attack methods.

[0105] The above content is a further detailed description of the present invention in combination with specific / preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, they can also make several substitutions or modifications to these described embodiments, and these substitution or modification methods should all be regarded as belonging to the protection scope of the present invention. In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "preferred embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. Without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples. Although the embodiments of the present invention and their advantages have been described in detail, it should be understood that various changes, substitutions, and alterations can be made herein without departing from the scope of protection of the patent application.

Claims

1. A method for improving the generalization ability of adversarial training based on Fisher information, characterized in that It includes the following steps: S1. Generate different adversarial images from clean images through different adversarial attack methods; S2. Linearly superimpose different adversarial images through Fisher information to obtain a generalized adversarial image; S3. Optimize the weighting coefficients of the generalized adversarial image through the particle swarm algorithm to obtain the optimal generalized adversarial image; S4. Use the optimal generalized adversarial image for adversarial training to obtain a classification model with stronger generalization ability; In step S2, the specific calculation of the generalization ability of the adversarial image by the Fisher information is as follows: Given an image classifier f, f(y|x) represents the distribution of the output label y of the classifier when the input clean image is x; at this time, the KL divergence of the output distributions of the clean image x and the image x + η after adding the perturbation η after passing through the classifier can be expressed as: wherein G x denotes the Fisher information matrix; The objective function for restricting the perturbation η in the KL divergence is: where ε represents the perturbation restriction range; Optimize the objective function that restricts the perturbation η in the KL divergence using the Lagrange multiplier method to obtain the optimized variable G x η = λη, where λ is the Lagrange multiplier; The optimization variable transforms the perturbation η in the KL divergence into the Fisher information matrix G x The objective function of the maximum eigenvalue of is further optimized as follows: where where s represents the softmax output of model f, and G s is a positive definite matrix; p i (s) represents the confidence score that the classifier outputs the i-th category after softmax, and M represents the number of classification categories; In step S2, the linear superposition is performed by linearly superposing adversarial images x1, x2,..., x generated by using N different adversarial methods for pairs of particle positions to obtain a generalized adversarial image N , where α is a weighting coefficient satisfying i ​ In step S2, the linearly superimposed generalized adversarial image is expressed as: Among them β1, β2, ..., β N-1 represent the N-1 dimensional particle positions corresponding to the clean image.

2. The method for improving the generalization ability of adversarial training based on Fisher information according to claim 1, wherein In step S1, the different adversarial attack methods are selected from two or more of the fast gradient sign method FGSM, projected gradient descent method PGD, C&W, JSMA, and DeepFool; the types of the different adversarial attack methods are selected from different types of adversarial attack methods or more of the same type of adversarial attack methods.

3. The method for improving the generalization ability of adversarial training based on Fisher information according to claim 1, wherein In step S3, the fitness function of different particles for optimizing the generalized adversarial image by the particle swarm algorithm, and the fitness function to be optimized is expressed as: where B is the batch size and M is the number of classification categories, represents the weighted sum of adversarial images generated by different adversarial attack methods for the j-th image in a batch; The optimization variable α of the fitness function to be optimized i (i = 1, 2,..., N) corresponding to β in the particle swarm optimization algorithm i (i = 1, 2,..., N - 1).

4. The method for improving the generalization ability of adversarial training based on Fisher information according to claim 1, wherein, In step S4, the adversarial training includes the following steps: S4-1. Calculate the fitness function to be optimized for different particles of the optimal generalized adversarial image, and save the current optimal position of each particle and the global optimal position among all particles; S4-2. Conduct multiple search iterations until the fitness function to be optimized converges or reaches the maximum number of iterations; S4-3. Calculate a batch of optimal generalized adversarial images through the global optimal position in the particles; S4-4. Calculate the gradient of the entire batch with respect to the classification model, and perform gradient descent to optimize the classification model parameters; S4-5. Conduct adversarial training for different batches; conduct multi-generation training for all training set data; S4-6. When the maximum number of training generations is reached, output a classification model with stronger generalization ability.

Citation Information

Patent Citations

  • Incremental learning method based on Fisher information matrix

    CN113469262A

  • Learning method of neural network model for language generation and apparatus for performing the learning method

    US20210089904A1