A method for enhancing the robustness of image classification

By generating adversarial sample detection networks and retraining the enhanced model, the problem of inaccurate classification results in the deep learning image classification model in the face of perturbation is solved, and the robustness and defense capabilities of the model are significantly improved.

CN112926661BActive Publication Date: 2025-05-30UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110222508.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-26
Publication Date
2025-05-30
Estimated Expiration
2041-02-26

AI Technical Summary

Technical Problem

Existing deep learning image classification models are prone to attack when faced with perturbations that are difficult to distinguish between naked eyes, resulting in inaccurate classification results and affecting the security of the model in practical applications.

Method used

A robust enhancement method for image classification model is proposed. By generating an adversarial sample detection network, each sample is evaluated, and an enhanced image classifier is trained in combination with the original model, using this model to classify images, and defending against most traditional white box adversarial samples.

Benefits of technology

This method can significantly improve the model's resistance to most commonly used white box anti-samples, with a wide range of defenses, and does not modify the parameters and structure of the original model, so that the model can be used in combination with other enhancement methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112926661B_ABST
    Figure CN112926661B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of computer software, specifically a method for enhancing the robustness of image classification, which can enhance the anti-interference and robustness of the classification model. The purpose is to be able to defend against attacks from most traditional white-box adversarial samples. It mainly includes an adversarial sample detection network generation module: by adding a neural network layer on the basis of the original classifier to construct an adversarial sample detection network, which mainly identifies adversarial samples; a judgment threshold generation module: using the commonly used adversarial sample method to find a suitable judgment threshold for the adversarial sample detection network. Enhanced model generation module: based on the classification of the image classifier of the original model, combined with the classification results of the aforementioned detection network, further training is performed to obtain an enhanced image classifier, and finally the enhanced image classifier is used to classify the image, thereby improving the robustness of the classifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer software, and specifically relates to a method for enhancing the robustness evaluation of an image classification model, which can enhance the anti-interference ability and robustness of the classification model. Background Art

[0002] In recent years, deep learning has been widely applied and achieved good results in image classification, face recognition, and language processing. Especially in image recognition, it can even match the performance of humans, and the existing technology has achieved an identification rate of over 99%. However, most researchers are more concerned about the performance of the model (such as accuracy) and ignore the vulnerability and robustness of the model. Foreign researchers such as Szegedy et al. found in experiments that by adding perturbations that are difficult to distinguish by the naked eye to images, the final model cannot obtain the correct classification result. Subsequently, Szegedy et al. proposed using the constrained L-BFGS algorithm to calculate perturbations, Goodfellow et al. proposed calculating perturbations based on the fast gradient sign algorithm, and Papernot et al. used an alternative neural network to fit the unknown neural network and then generated perturbations according to the alternative neural network. These algorithms can all generate good perturbations, resulting in the model making classification errors or classifying into the classification desired by the attacker. These problems have led people to start paying attention to whether personal security can be guaranteed when applying deep learning to real-world scenarios. Summary of the Invention

[0003] Aiming at the robustness problem of the above-mentioned existing model, the present invention proposes a method for enhancing the robustness of an image classification model. An adversarial sample detection network is generated based on the original model. Each sample is first evaluated, and finally, an enhanced image classifier is retrained in combination with the original model. The model is used for image classification. This method can defend against most traditional white-box adversarial samples without modifying the parameters of the original model, and can also improve the robustness of the model by combining other defense methods.

[0004] The present invention adopts the following technical solutions to solve the above technical problems:

[0005] A method for enhancing the robustness of image classification, including:

[0006] An adversarial sample detection network generation module: Obtain the structure of the original image classifier, obtain the hidden layer before the last fully connected layer, and add several fully connected layers on this basis to form an image detection network (the number of added layers can be set manually). The detection network can directly perform preliminary detection on adversarial samples. The last layer of the image detection network will be mapped to the size of the original image. The detection network is trained by optimizing the L2 distance between the original image and the output image of the image detection network. The purpose of training is to train a detection network that can preliminarily detect adversarial samples;

[0007] Judgment threshold generation module: Use common adversarial sample methods (PGD, C&W, BIM, etc.) to generate a certain number of adversarial samples, and combine them with a normal image dataset to obtain a judgment threshold n;

[0008] Enhanced model generation module: Based on the original image classification result, further train a combined enhanced model with the aforementioned image detection network, and use the enhanced image classification result as the final image classification result.

[0009] In the above technical solution, the image classifier mainly refers to an identification classifier based on a neural network algorithm, including any one of a fully connected neural network DNN, a convolutional neural network CNN, a staggered neural network ResNet, Xception, VGG19, and InceptionV3.

[0010] In the above technical solution, in the image detection network generation module, the steps of constructing the network are as follows:

[0011] S3.1: Obtain the hidden layer before the last fully connected layer of the CNN network and perform a flattening operation to change the multi-dimensional structure of the hidden layer into a one-dimensional structure. If it is a DNN network, no flattening operation is performed because the parameters of the convolutional network model, which is the mainstream of neural networks, will be represented in a computer similar to an array like [n, n], which is not convenient for subsequent combination, so it needs to be flattened into [n^2, 1];

[0012] S3.2: Add a fully connected layer layer1, where the number of neurons is twice the number of pixels of the original image. Build an autoencoder on the basis of the original model to identify adversarial samples without specific classification.

[0013] S3.3: Then add another fully connected layer layer2, where the number of neurons is the number of pixels of the original image.

[0014] In the above technical solution, in S4.1, during the process of training the network, the L2 distance between the reconstructed image output by the detection network and the original image is used as the objective function for iterative training;

[0015] S4.2: Perform a combination operation on the output of the fully connected layer layer1. Combine them in pairs according to the order of neurons, that is, add the corresponding outputs of 2 neurons. When fully connecting with the layer2 layer, the input data is the output after the final combination.

[0016] S4.3: The final output of the detection network is a reconstructed image of the original image size.

[0017] In the above technical solution, the training of the judgment threshold generation module has the following characteristics:

[0018] The following steps are adopted when generating the judgment threshold:

[0019] S5.1: Combine the normal image dataset and adversarial sample images with a quantity not less than 15%-25% of the normal image dataset to obtain the image dataset R;

[0020] S5.2: Input the image dataset R into the adversarial sample detection network and obtain the reconstructed image set R' output by the detection network;

[0021] S5.3: Calculate the L2 distance between all the output reconstructed images and the corresponding input image data, denote Lmax as the maximum value and Lmin as the minimum value;

[0022] S5.4: Set the parameter n to Lmin;

[0023] S5.5; Statistically calculate the probability tf that all correct image samples are determined as adversarial sample images under the condition of the current parameter n, and the probability fp that all determined adversarial sample images are actually adversarial sample images, denote p i = fp - tf, p i represents the effectiveness of the threshold value during the i-th search;

[0024] S5.6: When the value of n is not greater than Lmax, update and execute step S5.5, otherwise execute S5.7, where K is the number of iterations;

[0025] S5.7: Find the value n' of n that makes p i the largest, and set n = n'.

[0026] In the above technical solution, the input construction steps of the enhanced image classifier in the enhanced image classifier generation module are as follows:

[0027] S6.1. For each image data input x i , i represents the i-th input data. First, calculate the L2 distance l from the output of the image detection network. When l > n, set the corresponding adv i to 1, otherwise set it to 0. adv i represents the classification result of the detection network for it. The classification result refers to the specific classification label, specifically the category corresponding to this data;

[0028] S6.2. For each x i Take the output result logit i of its corresponding logits layer of the original model without passing through the softmax layer and adv iConstruct the training data tran_data for the subsequent enhanced model. This input data is the original output result with the judgment result of the detection network added; construct the training data for the subsequent enhanced model. The enhanced model is a model retrained through the classification result of the original model and the classification result of a detection network.

[0029] In the above technical solution, the construction of the training data labels of the enhanced image classifier network is as follows: Add one dimension to the label of each image data to mark whether it is an adversarial sample. For normal data, the value of this dimension is 0, and the values of other dimensions remain unchanged. For adversarial samples, its value is 1, and the values of other dimensions are set to 0; The training of the model requires input data and the corresponding labels of the data. For example, for a classifier that identifies cats and dogs, the labels of its data are [1, 0] or [0, 1]. What we need to do now is to modify this label to add one dimension representing the label of the adversarial sample. So the classifier becomes to identify cats and dogs, and the corresponding labels for adversarial samples should be [1, 0, 0], [0, 1, 0], [0, 0, 1].

[0030] In the above technical solution, 30% of the training data of the enhanced image classifier are adversarial samples (generated by PGD, C&W, FGSM methods). Using the stochastic gradient descent method, minimize the cross-entropy between the model output and the training data labels. The network structure uses a common convolutional network. This is mainly to ensure that the number of adversarial samples is not too small, which can be set according to the situation. If the success rate of identifying adversarial samples is low, more adversarial samples are needed.

[0031] In the above technical solution, the input data of the enhanced image classifier consists of the output of the logits layer of the original model and the detection result of the adversarial sample detection network. The output classification result directly includes the probability that the input picture is an adversarial sample.

[0032] In the above technical solution, the image classification result of the final enhanced image classifier is processed in the following way: First, obtain the image classification result of the enhanced image classifier. If the probability of belonging to an adversarial sample is the largest, then the input sample is determined to be an adversarial sample. Otherwise, re-normalize the prediction probabilities of the remaining categories and output the final image classification result.

[0033] Compared with the prior art, the beneficial effects of the present invention are shown in:

[0034] First, the present invention is applicable to applications such as face recognition and driverless that commonly use deep neural networks such as DNN or CNN, and can greatly improve the resistance of the model to most existing common white-box adversarial samples, with a wide range of defenses.

[0035] Second, without modifying the parameters and structure of the original model, the model can still be combined with other enhancement methods, such as adversarial training, distillation defense, etc.

[0036] Third, the prediction process is unique and different from the commonly used direct defense methods based on the original model, making it not easily broken. Description of the Drawings

[0037] Figure 1 is the overall architecture diagram of the present invention. Detailed Embodiments

[0038] The present invention will be further described below in conjunction with the drawings and specific embodiments.

[0039] A method for enhancing the robustness of image classification includes:

[0040] Adversarial sample detection network generation module: Obtain the structure of the original image classifier, obtain the hidden layer before the last fully connected layer, and add several fully connected layers on this basis to form an image detection network (the number of added layers can be set manually). The last layer of the image detection network will be mapped to the size of the original image, and the detection network is trained by optimizing the L2 distance between the original image and the output image of the image detection network;

[0041] Judgment threshold generation module: Use common adversarial sample methods (PGD, C&W, BIM, etc.) to generate a certain number of adversarial samples, and combine them with a normal image data set to obtain a judgment threshold n;

[0042] Enhanced model generation module: On the basis of the original image classification result, further train with the aforementioned image detection network to obtain an enhanced model, and the enhanced image classification result is used as the final image classification result.

[0043] In the above technical solution, the image classifier mainly refers to an identification classifier based on a neural network algorithm, including any one of a fully connected neural network DNN, a convolutional neural network CNN, a staggered neural network ResNet, Xception, VGG19, and lnceptionV3.

[0044] In the above technical solution, in the image detection network generation module, the steps for constructing the network are as follows:

[0045] S3.1: Obtain the hidden layer before the last fully connected layer of the CNN network and perform a flattening operation to change the multi-dimensional structure of the hidden layer into a one-dimensional structure. If it is a DNN network, no flattening operation is performed because the parameters of the convolutional network model, which is the mainstream of neural networks, will be represented in the form of an array similar to [n, n] in the computer, which is not convenient for subsequent combination, so it needs to be flattened into [n^2, 1];

[0046] S3.2: Add a fully-connected layer layer1, where the number of neurons is twice the number of pixels of the original image. Based on the original model, construct an autoencoder to identify adversarial samples without specific classification.

[0047] S3.3: Then add another fully-connected layer layer2, where the number of neurons is the number of pixels of the original image.

[0048] In the above technical solution, in S4.1, during the process of training the network, the objective function is to minimize the L2 distance between the reconstructed image output by the detection network and the original image, and iterative training is performed.

[0049] S4.2: Perform a combination operation on the output of the fully-connected layer layer1. Combine them in pairs according to the order of neurons, that is, add the corresponding outputs of 2 neurons. When fully connecting with the layer2 layer, the data passed in is the output after the final combination.

[0050] S4.3: The final output of the detection network is a reconstructed image of the size of the original image.

[0051] In the above technical solution, the training judgment threshold generation module has the following characteristics:

[0052] The following steps are adopted when generating the judgment threshold:

[0053] S5.1: Use the normal image dataset and combine it with adversarial sample images whose quantity is not less than 15%-25% of the normal image dataset to obtain the image dataset R.

[0054] S5.2: Input the image dataset R into the adversarial sample detection network and obtain the reconstructed image set R' of the output of the detection network.

[0055] S5.3: Calculate the L2 distance between all the output reconstructed images and the corresponding input image data. Denote Lmax as the maximum value and Lmin as the minimum value.

[0056] S5.4: Set the parameter n to Lmin.

[0057] S5.5; Statistically calculate the probability tf that all correct image samples are determined as adversarial sample images and the probability fp that all determined adversarial sample images are actually adversarial sample images under the condition of the current parameter n. Denote p i = fp - tf, p i represents the effectiveness of the threshold during the i-th search.

[0058] S5.6: When the value of n is not greater than Lmax, then update and execute step S5.5. Otherwise, execute S5.7, where K is the number of iterations.

[0059] S5.7: Find the value of n, denoted as n′, that maximizes p i and set n = n′.

[0060] In the above technical solution, the input construction steps of the enhanced image classifier in the enhanced image classifier generation module are as follows:

[0061] S6.1. For each input image data x i , where i represents the i-th input data, first calculate the L2 distance l from the output of the image detection network. When l > n, set the corresponding adv i to 1, otherwise set it to 0. adv i represents the classification result of the detection network for it;

[0062] S6.2. For each x i Construct the output result logit of the logits layer of its corresponding original model without passing through the softmax layer i and adv i into the training data tran_data of the subsequent enhanced model. This input data is the original output result with the judgment result of the detection network added; construct the training data of the subsequent enhanced model. The enhanced model is a model that is retrained through the classification result of the original model and a classification result of the detection network.

[0063] In the above technical solution, the training data labels of the enhanced image classifier network are constructed as follows: Add one dimension to the label of each image data to mark whether it is an adversarial sample. For normal data, the value of this dimension is 0, and the values of other dimensions remain unchanged. For adversarial samples, the value is 1, and the values of other dimensions are set to 0; The training of the model requires input data and the corresponding labels of the data. For example, for a classifier that identifies cats and dogs, the labels of its data are [1, 0] or [0, 1]. What we need to do now is to modify this label to add one dimension representing the label of the adversarial sample. So the classifier becomes one that identifies cats and dogs, and the corresponding labels of the adversarial samples should be [1, 0, 0], [0, 1, 0], [0, 0, 1].

[0064] In the above technical solution, 30% of the training data of the enhanced image classifier are adversarial samples (generated using the PGD, C&W, FGSM methods). Using the stochastic gradient descent method, minimize the cross-entropy between the model output and the training data labels. The network structure uses a common convolutional network. This is mainly to ensure that the number of adversarial samples is not too small and can be set according to the situation. If the success rate of identifying adversarial samples is low, more adversarial samples are needed.

[0065] In the above technical solution, the input data of the enhanced image classifier consists of the output of the logits layer of the original model and the detection results of the adversarial sample detection network, and the classification result directly includes the probability that the input picture is an adversarial sample.

[0066] In the above technical solution, the image classification result of the final enhanced image classifier is processed in the following way: First, obtain the image classification result of the enhanced image classifier. If the probability of belonging to an adversarial sample is the largest, then the input sample is determined to be an adversarial sample. Otherwise, the prediction probabilities of the remaining categories are renormalized and the final image classification result is output.

[0067] Embodiment

[0068] These experimental steps are based on the Windows 10 platform, the language used is python3.6, with dependencies such as tensorflow, theano, keras, etc., and the compilation software is pycharm. Table 1 shows the tools used in this embodiment.

[0069]

[0070]

[0071] Table 2 shows the main interface APIs used in this embodiment.

[0072] Serial Number API Description 1 CreatNet This interface creates a detection network based on the original model 2 GetN Obtain the optimal detection threshold 3 GetResults The determination result of this function's detection network 4 CreateData Create enhanced model training data based on the determination result 5 TrainModel Train the enhanced model

[0073] The specific implementation steps are executed according to the above modules:

[0074] I. Generation of the detection network: We select 60,000 samples from the Minist library, 80% as the training set, the remaining 10% as the training threshold dataset, and another 10% as the final test set. According to the described construction method, new network layers are added on the basis of the original model, and the network is trained by optimizing the L2 distance between the input and output.

[0075] II. Determination of the decision threshold: According to the method described above, first use the common white-box method to generate a total of 1,200 adversarial samples. Here, 3 methods, namely C&W, FGSM, and I-FGSM, are used. Together with the aforementioned 6,000 training threshold datasets as the input to the trained detection network, the order of the datasets is randomly shuffled, and the best n is iteratively generated according to the steps of finding the best n described above.

[0076] III. Training of the enhanced model: According to the method described above, use the previous detection network and decision threshold to construct data labels and train a new enhanced model.

[0077] IV. Test results: First, several common white-box attack methods such as C&W, BIM, JSMA, and One Pixel Attack are used to generate 1,000 adversarial samples. Together with the aforementioned 6,000 test data sets, they are used as inputs and fed into the above model output module. In the model output module, the distances between all samples and the reconstructed images are calculated. Based on the found n and combined with the original model, a new data set and data labels are constructed, and then the aforementioned enhanced model is trained. The test data is sequentially fed into the original model to obtain the output of the logits layer, and combined with the results of the detection network, it is fed into the enhanced model. Finally, the accuracy rate is statistically calculated based on the results.

[0078] The above are only representative embodiments among the numerous specific application scopes of the present invention, and do not constitute any limitation to the protection scope of the present invention. Any technical solutions formed by transformation or equivalent replacement fall within the scope of the rights protection of the present invention.

Claims

1. A method for enhancing the robustness of image classification, characterized in that, it includes: Adversarial sample detection network generation module: Obtain the structure of the original image classifier, obtain the hidden layer before the last fully connected layer, and add several fully connected layers on this basis to form an image detection network. The last layer of the image detection network will be mapped to the size of the original image, and the detection network is trained by optimizing the L2 distance between the original image and the output image of the image detection network; Judgment threshold generation module: Use common adversarial sample methods to generate a certain number of adversarial samples, and combine them with a normal image dataset to obtain a judgment threshold n; Enhanced model generation module: On the basis of the original image classification result, further train with the aforementioned image detection network to obtain an enhanced model, and the enhanced image classification result is used as the final image classification result; In the detection network generation module: S4.

1. During the training of the network, minimize the L2 distance between the reconstructed image output by the detection network and the original image as the objective function, and perform iterative training; S4.

2. Perform a combination operation on the output of the fully connected layer layer1. Combine them in pairs according to the order of neurons, that is, add the corresponding outputs of 2 neurons. When fully connecting with layer2, the input data is the output after the final combination; S4.

3. The final output of the detection network is a reconstructed image of the size of the original image; In the above solution, the training of the judgment threshold generation module has the following characteristics: The following steps are adopted when generating the judgment threshold: S5.1: Combine a normal image dataset with adversarial sample images whose quantity is not less than 15% - 25% of that of the normal image dataset to obtain an image dataset ; S5.2: Input the image dataset into the adversarial sample detection network and obtain the reconstructed image set of the output of the detection network ; S5.3: Calculate the L2 distance between all output reconstructed images and the corresponding input image data, denote Lmax as the maximum value and Lmin as the minimum value; S5.4: Set the parameter n to Lmin; S5.5: Statistically calculate the probability tf that an image sample judged as an adversarial sample image among all correct image samples under the condition of the current parameter n, and the probability fp that an image actually being an adversarial sample image among all images judged as adversarial sample images, and record , representing the effectiveness of the threshold value during the i-th search; S5.6: Update when the value of n is not greater than Lmax , and execute step S5.5; otherwise, execute S5.7, where K is the number of iterations; S5.7: Find the value of n that maximizes and set to ; The input construction steps of the enhanced image classifier in the enhanced image classifier generation module are as follows: S6.

1. For each input of image data , i representing the i th input data, first calculate the L2 distance from the output of the image detection network . When , set the corresponding to 1, otherwise set it to 0, representing the classification result of the detection network for it; S6.

2. For each output result of the logits layer of its corresponding original model without passing through the softmax layer and construct it into the training data of the subsequent enhanced model , and this input data is the original output result added with the judgment result of the detection network.

2. According to the method for enhancing the robustness of image classification described in claim 1, characterized in that: The image classifier mainly refers to a recognition classifier based on a neural network algorithm, including any one of a fully connected neural network DNN, a convolutional neural network CNN, a staggered neural network ResNet, Xception, VGG19, and InceptionV3.

3. According to the method for enhancing the robustness of image classification described in claim 1, characterized in that, In the image detection network generation module, the steps for constructing the network are as follows: S3.1: Obtain the hidden layer before the last fully connected layer of the CNN network and perform a flattening operation to change the multi-dimensional structure of the hidden layer into a one-dimensional structure. If it is a DNN network, no flattening operation is performed; S3.2: Add a fully connected layer layer1, where the number of neurons is twice the number of pixels of the original image; S3.3: Then add a fully connected layer layer2, where the number of neurons is the number of pixels of the original image.

4. According to the method for enhancing the robustness of image classification described in claim 1, characterized in that: The construction of the training data labels for the enhanced image classifier network is as follows: For each image data label, add one dimension to mark whether it is an adversarial sample. For normal data, the value of this dimension is 0, and the values of other dimensions remain unchanged. For adversarial samples, the value is 1, and the values of other dimensions are set to 0.

5. A method for enhancing the robustness of image classification according to claim 1, characterized in that: 30% of the training data for the enhanced image classifier is adversarial samples. Using the stochastic gradient descent method, minimize the cross-entropy between the model output and the training data labels. The network structure uses a common convolutional network.

6. A method for enhancing the robustness of image classification according to claim 1, characterized in that: The input data of the enhanced image classifier consists of the output of the logits layer of the original model and the detection results of the adversarial sample detection network. The output classification result directly includes the probability that the input picture is an adversarial sample.

7. A method for enhancing the robustness of image classification according to claim 1, characterized in that: The image classification result of the finally enhanced image classifier is processed in the following way: First, obtain the image classification result of the enhanced image classifier. If the probability of belonging to an adversarial sample is the largest, then the input sample is determined to be an adversarial sample. Otherwise, the predicted probabilities of the remaining classes are renormalized and the final image classification result is output.

Citation Information

Patent Citations

  • Method for enhancing robustness of neural network model

    CN110443367A

  • Adversarial sample defense system and method for artificial intelligence classification

    CN110569916A