A Two-Stage Adversarial Defense Method and System for Image Classification

By training multiple deep neural networks with different structures to build the optimal model pool, and performing random geometric transformation and random selection of classifiers on the adversarial images, the comprehensiveness and efficiency of adversarial attacks in the existing technology are solved, and effective defense against adversarial samples is achieved.

CN114881104BActive Publication Date: 2025-07-18CHENGDU SHUSHENG LANGLANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210297830.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-25
Publication Date
2025-07-18
Estimated Expiration
2042-03-25

AI Technical Summary

Technical Problem

The prior art is difficult to fully resist adversarial attacks of different types and different perturbation intensity, and the existing multi-model integrated defense methods increase time cost and are susceptible to attack migration.

Method used

By training multiple deep neural networks with different structures, the optimal model pool is built, and random geometric transformation and random selection classifiers are performed on the adversarial images to classify, combining data and model-level defense measures.

Benefits of technology

Effective defense against multiple attack types and different perturbation intensity is achieved, and it hardly affects the model's classification performance on clean samples, reducing the probability of attack success.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114881104B_ABST
    Figure CN114881104B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of adversarial sample defense, and particularly relates to a two-stage adversarial defense method and system for image classification. The method includes: training multiple deep neural networks with different structures using the same training set to obtain multiple trained classifiers as alternative models; constructing an optimal model pool from the multiple alternative models; performing random geometric transformation on the adversarial images to destroy the specific spatial structure of the perturbations in the adversarial images; randomly selecting a classifier from the optimal model pool to classify the images after the random geometric transformation. The adversarial defense method proposed by the present invention constructs defenses from both the data and model levels, can more comprehensively resist adversarial images of various types and different perturbation intensities, and hardly affects the classification performance of the model for clean samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of adversarial sample defense, and particularly relates to a two-stage adversarial defense method and system for image classification. Background Art

[0002] Deep learning is an algorithm based on representing learning of data. Since its inception, it has shown excellent results in various fields such as computer vision, natural language processing, and speech recognition. As the core technology in the field of deep learning, deep neural networks have currently been widely applied to various task systems, such as image classification, object detection, face recognition, machine translation, etc. Although deep neural networks have excellent performance in various tasks, researchers have found that DNNs are vulnerable and are extremely susceptible to attacks by adversarial samples. An adversarial sample refers to a new sample obtained by adding specific adversarial perturbations to the original clean data sample, where the perturbations are carefully crafted and difficult to detect by the human eye, but it will induce the deep neural network model to output incorrect results with a high confidence when making predictions. In the image classification task, an adversarial sample is an image with perturbation noise maliciously synthesized by an attacker. The adversarial image is indistinguishable from its corresponding natural image, and it is difficult for humans to perceive the adversarial perturbation through visual senses. Therefore, the adversarial sample does not affect human understanding of the image semantics, but when the adversarial image is fed into an intelligent system built based on a deep neural network, it can change the classification result of the model. In addition, adversarial samples do not only exist in the digital world. In the real physical world, attackers can also use adversarial samples to deceive intelligent systems. For example, an attacker only needs to add some specific "stickers" to a road sign to make it adversarial, which can make the target recognition system misidentify the road sign, thereby leading to serious traffic accidents; some researchers use light projection means to project adversarial patterns onto people's faces to deceive or evade face recognition systems. The existence of adversarial samples has made people start to pay attention to the security and credibility of intelligent systems built based on deep neural networks. And how to make deep neural networks effectively resist adversarial attacks and improve the adversarial robustness of the model has become a very important topic.

[0003] Most existing defenses have a common drawback: they have good performance against some types of adversarial attacks with small perturbation intensities, which makes the defense lack comprehensiveness, that is, they cannot resist adversarial attacks of different types and different perturbation intensities. In addition, existing research uses the method of multi-model integrated decision-making to resist adversarial samples, which not only significantly increases the time cost required for prediction, but also when multiple integrated models are similar, adversarial samples can rely on their migration ability to attack the integrated model, thus failing to achieve the purpose of defense. Therefore, how to effectively integrate models to make the models diverse is also an issue to be solved. Summary of the Invention

[0004] In order to improve the effect of model defense against attacks, the present invention proposes a two-stage defense method and system for image classification, the method comprising the following steps:

[0005] Train multiple deep neural networks with different structures using the same training set to obtain multiple trained classifiers as candidate models;

[0006] Build an optimal model pool using multiple candidate models;

[0007] Perform random geometric transformations on the adversarial image to destroy the specific spatial structure of the perturbation in the adversarial image;

[0008] A classifier is randomly selected from the pool of optimal models to classify the image after random geometric transformation.

[0009] Furthermore, the alternative model is a convolutional neural network including at least a convolutional layer, a pooling layer, a fully connected layer and a softmax layer.

[0010] Furthermore, the process of constructing the optimal model pool with a pool size of s includes:

[0011] Arbitrarily combine the trained candidate models into different model pools with a number s in the pool. As an optional implementation, the size of the optimal model pool is {2, 3, …, m-1}, where m is the total number of candidate models;

[0012] Calculate the diversity-accuracy index DAA for each model pool pool ;

[0013] All DAA pool The model pool corresponding to the maximum value of the indicator is taken as the optimal model pool when the model pool size is s.

[0014] Furthermore, the diversity-accuracy index DAA of the model pool pool It is expressed as:

[0015]

[0016] in, It means pairing the classifiers in the model pool with each other, with a total of right; Represents the diversity-accuracy index between the nth pair of models.

[0017] Furthermore, a pair of models F i and Model F j Diversity-Accuracy Index (DAA) i,j It is expressed as:

[0018] DAA i,j =PDi,j +Acc i,j ;

[0019] Among them, DAA i,j represents the diversity-accuracy index of models F i and model F j , PD i,j represents the diversity measure between paired models F i and F j , and Acc i,j represents the accuracy measure between the two models.

[0020] Furthermore, the diversity measure PD i,j and the accuracy measure Acc i,j between paired models are calculated as follows:

[0021]

[0022]

[0023] Among them, x k represents the k-th sample, k = 1, 2,..., N; d(F i (x k ), F j (x k )) represents the difference between models F i and model F j on sample x k ; Acc i , Acc j respectively represent the accuracies of models F i and model F j on N samples.

[0024] Furthermore, the difference d(F i (x), F j (x)) shown by paired models on the sample is expressed as:

[0025]

[0026] Among them, F i (x), F j (x) respectively represent the softmax outputs of models F i and F j for the input data x, F i (x) = [p i,1 (x), p i,2 (x),..., p i,C (x)], F j (x) = [p j,1 (x), pj,2 (x),...,p j,C (x)], p i,c (x), p j,c (x) are respectively the probabilities of the model F i and the model F j to classify the sample x into the c-th class, where C is the total number of classes.

[0027] The present invention proposes a two-stage adversarial defense system for image classification, which is used to implement a two-stage adversarial defense method for image classification. The system includes a data acquisition module, a classifier training module, and an optimal model pool selection module, where:

[0028] The data acquisition module acquires historical data for training the classifier and classifies the data to be predicted using a classifier randomly selected from the optimal model pool;

[0029] The classifier training module includes multiple deep neural networks with different structures. Multiple classifiers are obtained by training multiple deep neural networks with different structures using the same training set;

[0030] The optimal model pool selection module selects s classifiers from the trained classifiers for combination according to different model pool sizes s, calculates the diversity-accuracy index of each combination, and takes the combination with the largest diversity-accuracy index as the optimal model pool corresponding to the current pool size s. Then, the model pools corresponding to different pool sizes are used to predict the same batch of clean data samples, the classification accuracies of each model pool are obtained, and the model pool with the highest accuracy is used as the finally constructed optimal model pool.

[0031] Advantages of the present invention:

[0032] 1. The present invention is more comprehensive in terms of defense capabilities, can resist adversarial samples of multiple attack types and different perturbation intensities, and hardly affects the classification performance of the model for clean samples.

[0033] 2. The present invention determines an effective defense combination, which can construct defense measures from both the data level and the model level simultaneously to achieve the purpose of dual defense. Brief Description of the Drawings

[0034] Figure 1 is the structural diagram of the two-stage adversarial defense in the present invention;

[0035] Figure 2 is the relationship diagram between the DAA pool index of the model pool with a size of 5 and the pool classification accuracy;

[0036] Figure 3 is the DAA poolRelationship diagram between indicators and pool classification accuracy;

[0037] Figure 4 It is the effect diagram of the pool classification accuracy of the optimal model pool of different scales;

[0038] Figure 5 It is the defense effect diagram of the two-stage adversarial defense against different attacks in the present invention;

[0039] Figure 6 It is the comparison diagram of the defense effects of different algorithms. Specific implementation manners

[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0041] The present invention proposes a two-stage adversarial defense method for image classification, which specifically includes the following steps:

[0042] Train multiple deep neural networks with different structures using the same training set to obtain multiple trained classifiers as candidate models;

[0043] Construct an optimal model pool using multiple candidate models;

[0044] Perform random geometric transformation on the adversarial image to destroy the specific spatial structure of the perturbation in the adversarial image;

[0045] Randomly select a classifier from the optimal model pool to classify the image after random geometric transformation.

[0046] In this embodiment, a two-stage adversarial defense method for image classification is proposed, as Figure 1As shown in the figure, the method of the present invention includes two stages, namely: the first stage of defense is the RRF transformation, which refers to the combined transformation of image scaling and image flipping of the input image; the second stage of defense is random prediction, in which a classifier is randomly selected from the constructed optimal model pool to complete the classification prediction of the transformed image. The image transformation in the first stage is an ordinary image enhancement for clean samples. However, for adversarial samples, since the perturbations in the samples are sensitive to positions, the performance of the same perturbation varies greatly at different positions. Therefore, through this transformation, the specific spatial structure of the adversarial perturbation can be destroyed, thereby weakening the aggressiveness of the adversarial samples. The random prediction in the second stage constructs a defense from the model level by utilizing the weak transferability of adversarial samples among different models. The random idea makes it impossible for the attacker to determine the classifier used for prediction each time, thereby reducing the probability of targeted attacks. Compared with the existing classification methods, whether the input image is the original image or the adversarial image, the method of the present invention has good classification performance.

[0047] In this embodiment, multiple classifiers are trained on deep neural networks with different network structures using the same training set as alternative models. The deep neural networks with different network structures at least include a convolutional layer, a pooling layer, a fully connected layer, and a softmax layer. Those skilled in the art can set the specific number of convolutional layers, pooling layers, etc. according to actual needs, which will not be elaborated in the present invention; only the most basic network structure of the deep neural network is defined in the present invention. Those skilled in the art can add other structures to this most basic network structure. For example, structures such as the attention mechanism can be added to the deep neural network, and the present invention does not limit this.

[0048] The adversarial image is input into the data preprocessing layer for random geometric transformation to destroy the specific spatial structure of the perturbations in the image. As an alternative implementation, the random geometric transformation refers to the combined transformation of image scaling and image flipping of the target image, and the change size during the image scaling operation is not fixed.

[0049] An optimal model pool is constructed using the trained classifiers, and then a classifier is randomly selected from the optimal model pool to perform classification prediction on the adversarial image after random geometric transformation to obtain the classification result. The process of constructing the optimal model pool includes:

[0050] Step 1: Construct optimal model pools corresponding to different pool sizes s (s = 2, 3,..., m - 1).

[0051] Step 2: Select s classifiers from all the classifiers as a model pool.

[0052] Step 3: Calculate the diversity-accuracy index DAA of each model pool pool, when calculating this metric, pair up the classifiers in a model pool, and a total of pairing results are obtained. Then, the diversity-accuracy metric DAA i,j of each pair of models is expressed as:

[0053] DAA i,j = PD i,j + Acc i,j ;

[0054] Among them, DAA i,j represents the diversity-accuracy metric of model F i and model F j . PD i,j represents the diversity measure between the paired models F i and F j , and Acc i,j represents the accuracy measure between the two models.

[0055] The calculation process of the diversity measure PD i,j and the accuracy measure Acc i,j between paired models includes:

[0056]

[0057]

[0058]

[0059] Among them, x k represents the k-th sample, k = 1, 2,..., N; d(F i (x k ), F j (x k )) represents the difference between model F i and F j on the sample x k . Acc i , Acc j respectively represent the accuracies of the two models on N samples. F i (x), F j (x) respectively represent the softmax outputs of model F i and F j for the input data x. F i (x) = [p i , (1x), p i (x,), 2..., p iC (x)], F j (x) = [p j,1 (x), p j,2 (x),..., pj,C (x)],p ,c p(x) is the probability that model F classifies sample x into the c-th class, and C is the total number of classes.

[0060] Calculate the DAA of all pairwise models i,j The mean value of the metrics to obtain the diversity-accuracy metric DAA of the model pool pool , denoted as:

[0061]

[0062] Step 4, take all DAA pool The model pool corresponding to the maximum value among the metrics as the optimal model pool when the model pool size is s, denoted as DAA opt = MAX(DAA pool ).

[0063] Step 5, classify and predict the same batch of clean data samples by the optimal model pools corresponding to different pool sizes s to obtain their respective pool classification accuracies.

[0064] Step 6, select the model pool with the highest pool classification accuracy as the optimal model pool.

[0065] This implementation also proposes a two-stage adversarial defense system for image classification, which is used to implement a two-stage adversarial defense method for image classification. The system includes a data acquisition module, a classifier training module, and an optimal model pool selection module, where:

[0066] The data acquisition module acquires historical data for training the classifier and acquires the data to be predicted for classification by a classifier randomly selected from the optimal model pool;

[0067] The classifier training module includes multiple deep neural networks with different structures, and uses the same training set to train multiple deep neural networks with different structures to obtain multiple classifiers;

[0068] The optimal model pool selection module selects s classifiers from the trained classifiers for combination according to different model pool sizes s, calculates the diversity-accuracy metric for each combination, and takes the combination with the largest diversity-accuracy metric as the optimal model pool corresponding to the current pool size s. Then, the model pools corresponding to different pool sizes are used to predict the same batch of clean data samples to obtain the classification accuracies of each model pool, and the model pool with the highest accuracy is used as the finally constructed optimal model pool.

[0069] In the present invention, a two-stage adversarial defense system for image classification can be a hardware system with modules connected, or a computer program stored in a medium and implemented by a processor to realize a two-stage adversarial defense method for image classification. For example, a computer device for two-stage adversarial defense of image classification includes a memory and a processor. The memory is used to store a computer program, and the processor is used to run the computer program stored in the memory to realize a two-stage adversarial defense method for image classification.

[0070] This embodiment presents an application of the method of the present invention on the ImageNet dataset. Experiments were conducted using the ImageNet dataset, and the images in this dataset cover most of the image categories that can be seen in daily life. 5000 clean images that can be correctly classified by all classifiers were selected from its validation set as source images, and a total of 10 groups of adversarial samples with different perturbation intensities were generated using FGSM, BIM, PGD, and CW attack algorithms to attack the target model, with 5000 adversarial samples in each group.

[0071] The comparison metric used in the present invention is the classification accuracy Acc of the model. The methods of comparative experiments adopted in the present invention include:

[0072] WF: Implement defense using WebP compression and image flipping transformation;

[0073] PD: Implement defense using the pixel flipping method;

[0074] BDR: Implement defense using pixel compression;

[0075] ComDefend: Implement defense using an end-to-end image compression method.

[0076] In this embodiment, an experiment on constructing a model pool was first conducted, with the aim of obtaining an optimal model pool for subsequent defense classification. In this embodiment, the diversity-accuracy index DAA of the model pool formed by different model combinations was calculated on the ImageNet dataset pool and the corresponding classification accuracy of each model pool was obtained through experiments. In this example, 12 common convolutional neural networks with different network structures were trained as candidate models for constructing the model pool, namely InceptionV3, ResNet50, VGG16, SqueezeNet, GoogleNet, ResNet101, DenseNet169, ShuffleNetV2, MNasNet1, MobileNetV2, VGG16bn, WideResNet50, which are numbered M1 to M12 for convenient description. For a model pool of a certain scale (scale of 5, 7), 5 groups of model pools were selected to display DAApool The relationship with the pool classification accuracy, and the results are as Figures 2-3 shown. In this embodiment, there are a total of 12 classifiers, M1 to M12. Figure 2 For the DAA with the optimal model pool size of 5 pool The relationship diagram between the index and the pool classification accuracy. Figure 3 For the DAA with the optimal model pool size of 7 pool The relationship diagram between the index and the pool classification accuracy. In both figures, the abscissa is the value of the DAA pool index, and the ordinate is the model accuracy. Among them, Figure 2 The first empirical pool given from left to right consists of classifier M4, classifier M5, classifier M8, classifier M9, and classifier M10. The second empirical pool consists of classifier M4, classifier M5, classifier M7, classifier M8, and classifier M12. The third empirical pool consists of classifier M2, classifier M6, classifier M7, classifier M9, and classifier M10. The fourth empirical pool consists of classifier M1, classifier M2, classifier M6, classifier M9, and classifier M11. The fifth empirical pool consists of classifier M1, classifier M5, classifier M6, classifier M7, and classifier M12; Figure 3 The first empirical pool given from left to right consists of classifier M1, classifier M2, classifier M5, classifier M6, classifier M7, classifier M11, and classifier M12. The second empirical pool consists of classifier M1, classifier M2, classifier M3, classifier M6, classifier M7, classifier M9, and classifier M11. The third empirical pool consists of classifier M3, classifier M5, classifier M6, classifier M7, classifier M8, classifier M9, and classifier M10. The fourth empirical pool consists of classifier M3, classifier M4, classifier M5, classifier M9, classifier M10, classifier M11, and classifier M12. The fifth empirical pool consists of classifier M2, classifier M3, classifier M4, classifier M6, classifier M8, classifier M9, and classifier M10. It can be found that as the model pool diversity-accuracy index DAA pool increases, the classification performance shown by the model pool is better, which indicates that the combination of diversity and accuracy of the models in the pool is better. For the relationship between the optimal model pool corresponding to different scales and the pool classification accuracy, the results are as Figure 4 shown. It can be found that the classification performances shown by the optimal model pools of different scales are very little different. However, when the scale is too large or too small, the accuracy is relatively low. Therefore, when selecting the scale, it is best to randomly select the scale between 1 / 3 and 3 / 4 of the total number of classifiers. In this embodiment, in order to save memory costs, the model pool with a size of 5 corresponding to a high pool classification accuracy and a small size is selected as the optimal model pool for the subsequent experiments.

[0077] This embodiment verifies the defense effect of the two-stage adversarial defense (TPAD) against various attacks, and the results are as follows Figure 5 shown. The abscissa represents the types of various attacks with different types and different perturbation intensities, specifically including FGSM0.008, FGSM0.02, FGSM0.04, FGD0.04, FGD0.06, FGD0.08, BIM0.02, BIM0.05, BIM0.08, and C&W. Each attack type includes four prediction methods, which are, from left to right, no defense, RRF transformation, random prediction column, and the TPAD of the present invention. The ordinate represents the accuracy corresponding to each defense method. It can be found that the proposed method shows good defense effects against adversarial attacks with different types and different perturbation intensities. In addition, by comparing the results of implementing defense from a single perspective (RRF transformation and random prediction column) in the figure, it can be seen that the effect of constructing defense from both the data and model perspectives is better.

[0078] This embodiment is compared with existing defense methods, and the results are as follows Figure 6 shown Figure 6 In the figure, the abscissa represents the types of various attacks with different types and different perturbation intensities, specifically including FGSM0.008, FGSM0.02, FGSM0.04, FGD0.04, FGD0.06, FGD0.08, BIM0.02, BIM0.05, BIM0.08, and C&W. Each attack type includes five prediction methods, which are, from left to right, no defense, ComDefend, PD, BDR, WF, and the TPAD of the present invention. The ordinate represents the accuracy corresponding to each defense method. Compared with other defenses, the TPAD method proposed by the present invention shows the optimal defense ability against all attacks, and as the perturbation intensity of various attacks increases, the defense effect shown by the TPAD method becomes more significant. This shows that the TPAD method is more prominent in defense performance and can better resist high-intensity adversarial attacks than other methods. In addition, in this embodiment, the influence of each defense on the classification performance of clean samples is also tested. From the clean data in Figure 6 , it can be seen that this method reduces the classification accuracy of the model for clean samples by less than 2%, and the degree is minor. In summary, it shows that the TPAD defense method proposed by the present invention can not only resist various types and intensities of adversarial attacks, but also hardly affect the classification performance of the model for clean samples.

[0079] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirits of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A two-stage adversarial defense method for image classification, characterized in that, Specifically, it includes the following steps: Train multiple deep neural networks with different structures using the same training set to obtain multiple trained classifiers as alternative models; Construct an optimal model pool using multiple alternative models, specifically including: Arbitrarily combine the trained alternative models into different model pools pool with the number s in the pool; Calculate the diversity-accuracy metric DAA of each model pool pool , which is expressed as: Among them, means pairing classifiers in the model pool pairwise, with a total of pairs; represents the diversity-accuracy index between the nth pair of models; Take all DAA pool The model pool corresponding to the maximum value among all the DAA metrics is used as the optimal model pool when the model pool size is s; Perform random geometric transformations on the adversarial images to destroy the spatial structure of the perturbations in the adversarial images; Randomly select a classifier from the optimal model pool and classify the images after random geometric transformation.

2. The two-stage adversarial defense method for image classification according to claim 1, wherein The alternative model is a convolutional neural network that at least includes a convolutional layer, a pooling layer, a fully connected layer, and a softmax layer.

3. A two-stage adversarial defense method for image classification according to claim 1, characterized in that, A pair of models F i and model F j The diversity-accuracy index DAA i,j is expressed as: DAA i,j = PD i,j + Acc i,j ; Among them, DAA i,j represents the diversity-accuracy metric of model F i and model F j , and PD i,j represents the diversity measure between paired models F i and F j , and Acc i,j represents the accuracy measure between the two models.

4. A two-stage adversarial defense method for image classification according to claim 3, characterized in that, Diversity measure PD between paired models i,j and accuracy measure Acc i,j The calculation of includes: where, x k represents the k-th sample, k = 1, 2,..., N; d(F i (x k ), F j (x k )) represents the difference between model F i and model F j on sample x k ; Acc i , Acc j respectively represent the accuracies of model F i and model F j on N samples.

5. A two-stage adversarial defense method for image classification according to claim 4, characterized in that The difference d(F i (x), F j (x)) shown by the paired model is expressed as: Among them, F i (x k ) and F j (x k ) respectively represent the softmax outputs of model F i and model F j for the input data x k . It is expressed as F i (x k ) = [p i,1 (x k ), p i,2 (x k ),..., p i,C (x k )], F j (x k ) = [p j,1 (x k ), p j,2 (x k ),..., p j,C (x k )]. p i,c (x k ) and p j,c (x k ) are the probabilities that model F i and model F j classify the sample x k into the c-th category, and C is the total number of categories.

6. A two-stage adversarial defense system for image classification, characterized in that, This system is used to implement any one of the two-stage adversarial defense methods for image classification described in claims 1 to 5. This system includes a data acquisition module, a classifier training module, and an optimal model pool selection module, where: The data acquisition module acquires historical data for training the classifier and acquires the data to be predicted for classification by a classifier randomly selected from the optimal model pool; The classifier training module includes multiple deep neural networks with different structures in this module. Use the same training set to train multiple deep neural networks with different structures to obtain multiple classifiers; The optimal model pool selection module selects s classifiers from the trained classifiers for combination according to different model pool scales s, calculates the diversity-accuracy index of each combination, and uses the combination with the largest diversity-accuracy index as the optimal model pool corresponding to the current pool scale s; Use the model pools corresponding to different pool scales to predict the same batch of clean data samples, obtain the classification accuracy of each model pool, and use the model pool with the highest accuracy as the finally constructed optimal model pool.

7. A two-stage adversarial defense system for image classification according to claim 6, characterized in that, When selecting the optimal model pool, the model pool with the largest diversity-accuracy metric is taken as the optimal model pool, i.e., DAA opt = MAX(DAA pool ), and the diversity-accuracy metric DAA of the model pool pool is expressed as: DAA i,j = PD i,j + Acc i,j ; Among them, DAA opt is the diversity-accuracy index of the optimal model pool; DAA i,j represents the diversity-accuracy index of a pair of models F i and model F j ; represents the nth pair of models in a certain model pool, that is, the diversity-accuracy index of model F i and model F j ; PD i,j represents the diversity measure between the paired models F i and model F j , Acc i,j represents the accuracy measure between model F i and model F j ; represents the number of pairings of models when the size of the optimal model pool is s.

8. A two-stage adversarial defense system for image classification according to claim 7, characterized in that, Paired model F i and F j The diversity measure PD i,j is expressed as: Model F i With Model F j The accuracy metric Acc i,j Is expressed as: Among them, x k represents the k-th sample, where k = 1, 2,..., N; d(F i (x k ), F j (x k )) represents the difference between model F i and model F j on sample x k ; Acc i , Acc j respectively represent the accuracies of model F i and model F j on N samples; F i (x k ), F j (x k ) respectively represent the softmax outputs of model F i , model F j for the input data x k , expressed as F i (x k ) = [p i,1 (x k ), p i,2 (x k ),..., p i,C (x k )], F j (x k ) = [p j,1 (x k ), p j,2 (x k ),..., p j,C (x k )], where p i,c (x k ), p j,c (x k ) are the probabilities that model F i and model F j classify sample x k into the c-th class, and C is the total number of classes.