Classification model defense method based on dynamic adversarial training and attention area denoising

Through dynamic adversarial training and area-focused denoising technology, the problems of overfitting and robustness in existing defense methods are solved, and the model's high robustness and accuracy under multiple adversarial attacks are achieved.

CN120125907APending Publication Date: 2025-06-10ANHUI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510280162.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing adversarial sample defense methods have problems such as overfitting, large computational overhead and poor handling of local perturbations, resulting in insufficient robustness in the face of multiple adversarial attacks.

Method used

A classification model defense method based on dynamic adversarial training and denoising of attention areas is adopted, and the learning rate is dynamically adjusted and a variety of adversarial attack methods are used for training, and a robust region of attention is determined in combination with class activation graph technology, and denoising is performed to weaken the adversarial perturbation.

Benefits of technology

It significantly improves the robustness and classification accuracy of the model under multiple adversarial attacks, avoids overfitting a single attack method, and effectively retains the key information of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125907A_ABST
    Figure CN120125907A_ABST
Patent Text Reader

Abstract

The invention discloses a classification model defense method based on dynamic adversarial training and attention area denoising, and aims to improve the robustness of a classification model under adversarial attack. In the training stage, FGSM, PGD and Camp are dynamically selected according to the number of training rounds; w and other different confrontation sample generation methods are adopted, so that specific attacks of model over-fitting are avoided, and the generalization ability to various attacks is enhanced. Meanwhile, a CAM (Class Activation Mapping) is used for positioning an input image attention area, wavelet denoising processing is adopted, weighted fusion with an original image is carried out, countermeasure disturbance is accurately weakened, and key information is reserved. Experiments show that on a CIFAR-10 data set, the method significantly improves the model performance in FGSM, PGD and Camp; compared with a traditional method, the robustness accuracy is improved by 10%-20%, and the method can be widely applied to the fields of automatic driving, security and protection monitoring and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, relates to image security and adversarial sample defense technologies, and particularly relates to a classification model defense method based on dynamic adversarial training and region-of-interest denoising. Background Art

[0002] With the development of deep learning technology, neural network models have made remarkable progress in fields such as computer vision, autonomous driving, and security monitoring. However, these deep models are vulnerable to adversarial attacks, that is, adding tiny perturbations to the input data can cause the model to output incorrect classification results. Adversarial attacks not only threaten the reliability of the model but also may cause serious security hazards, especially in critical application scenarios such as autonomous driving and medical diagnosis.

[0003] Currently, a variety of adversarial sample defense methods have been proposed, including adversarial training, input transformation, model regularization, etc. Among them, adversarial training is one of the most effective methods, that is, adding adversarial samples during the model training process to improve the robustness of the model. However, traditional adversarial training methods still have the following deficiencies:

[0004] 1. Limited robustness: Conventional adversarial training mainly relies on fixed attack methods, such as FGSM (Fast Gradient Sign Method) or PGD (Projected Gradient Descent). The model is prone to overfitting to specific types of adversarial attacks and still lacks robustness against unknown or stronger attack methods.

[0005] 2. High computational cost: Adversarial training requires continuously generating adversarial samples and updating the model, with a relatively high computational cost. Especially when the deep network structure is complex, the training cost increases significantly.

[0006] 3. Ineffective suppression of local perturbations of adversarial samples: Research shows that adversarial perturbations often concentrate on the key feature regions of the target object, while existing defense methods usually process the entire image rather than targeting local key regions, resulting in the model failing to effectively resist attacks. Summary of the Invention

[0007] Object of the Invention: The object of the present invention is to solve the deficiencies existing in the prior art and provide a classification model defense method based on dynamic adversarial training and region-of-interest denoising to solve problems such as overfitting and loss of key information existing in existing defense methods, significantly improve the robustness and classification accuracy of the classification model under various adversarial attacks, and ensure the reliability and security of the model in practical applications.

[0008] Technical solution: A classification model defense method based on dynamic adversarial training and denoising of the attention area of the present invention includes the following steps:

[0009] Step 1, in the adversarial training stage of the classification model, prepare a classification model (such as ResNet-18), a training data set, and various adversarial attack methods, perform dynamic adversarial training on the original basic model (such as ResNet-18), optimize the model parameters through backpropagation, dynamically adjust the learning rate, and use the cross-entropy loss function to calculate the loss;

[0010] Step 2, in the robust attention area denoising stage, first generate adversarial samples of the test data, and use the class activation map CAM to determine the robust attention area of the adversarial samples, and then perform denoising processing on the obtained robust attention area to weaken the influence of adversarial perturbations. The specific steps are as follows:

[0011] Step 2.1, use various adversarial attack methods to generate adversarial samples of the test data; for example, only FGSM can be used in the early stage, FGSM and PGD can be randomly selected in the middle stage, and FGSM, PGD, and C&W can be randomly selected in the later stage;

[0012] Step 2.2, different from the prior art that only relies on a single highest-probability category to generate the activation map, here the information of the top k categories is comprehensively considered and weighted to generate the final CAM to determine the robust attention area. The class activation map is calculated by the following formula:

[0013]

[0014] Among them, is the comprehensive weight coefficient obtained after optimizing the weights of the fully connected layers corresponding to the top k categories, F k (x) is the k-th channel of the feature map; the value of k depends on the characteristics of the data set, and a reasonable value needs to be selected through experimental testing to ensure that the key area information related to the true classification of the adversarial samples can be effectively captured;

[0015] Step 2.3, perform denoising processing on the robust attention area of the adversarial sample to weaken the influence of adversarial perturbations. Here, local average denoising or wavelet denoising methods are used;

[0016] Step 2.4, perform weighted fusion of the denoised image and the original image, and the formula is:

[0017] X D =α·X d +β·X c

[0018] Among them, α and β are weight coefficients, usually set to 0.7 and 0.3, X d is the denoised image, X cis a clean image;

[0019] Step 3: Classification stage. Input the adversarial samples after robust region-of-interest denoising into the model that has completed adversarial training, and then classification can be performed to obtain the final result and detect the classification accuracy rate of the model.

[0020] Further, the method for constructing the dataset in Step 1 is as follows: Select a dataset and divide it into a training set and a test set. Perform preprocessing operations on the images, such as randomly flipping them horizontally (with a 50% probability), normalizing (normalizing the R, G, and B channels), and Cutout (randomly occluding parts of the images).

[0021] Use ResNet-18 as the original classification model and appropriately modify the model considering the image size of the dataset.

[0022] The steps of dynamic adversarial training are as follows:

[0023] Step a): The training set for adversarial training includes clean samples X c and adversarial samples X adv . For the adversarial samples X adv , the method for generating adversarial samples is dynamically selected according to the number of training epochs. For example, only FGSM can be used in the early stage, FGSM and PGD can be randomly selected in the middle stage, and FGSM, PGD, and C&W can be randomly selected in the later stage.

[0024] Step b): Use the torch.cat() function to perform tensor concatenation operations. Concatenate the clean samples and adversarial samples along the batch dimension to form a training input mixed sample X mix , whose shape is (2B, C, H, W), and the batch size becomes twice the original. Since the adversarial samples are generated based on the clean samples and their true classes are the same as the corresponding clean samples, the target labels themselves are still concatenated along the batch dimension to obtain the mixed target label Y mix ;

[0025] Step c): After obtaining the concatenated mixed sample X mix and the mixed target label Y mix , start training, and at the same time use the cross-entropy loss function L train to calculate the loss. The formula is as follows:

[0026]

[0027] Further, the steps for optimizer selection and learning rate scheduling in the adversarial training stage of the classification model in Step 1 include: Use the Adam optimizer to update the model parameters, with an initial learning rate of 0.001 and a weight decay of 5e-4; the learning rate is halved every 5 epochs, and if the validation loss decreases, the model parameters are saved.

[0028] Further, the steps in the classification stage of step 3 include:

[0029] Load the model weights after adversarial training, and input the denoised adversarial samples in the robust attention area into the model after adversarial training for classification; calculate the classification accuracy of the model for the denoised adversarial samples, and evaluate the robustness of the model.

[0030] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0031] (1) Improve model robustness: Through dynamic adversarial training, the model can learn various types of adversarial features, avoiding overfitting to a single attack method, so as to maintain a high classification accuracy when facing various complex adversarial attacks. Experimental results show that under various attacks such as FGSM, PGD, and C&W, the robustness accuracy of the model trained by the method of the present invention is increased by 10%-20% compared with the traditional adversarial training method.

[0032] (2) Retain key image information: The robust attention area denoising method breaks through the traditional limitations and abandons the single-category weight mode by means of the class activation mapping (CAM) technology. By comprehensively considering the information of the top k categories, the weights of the fully connected layer are optimized to generate more targeted weight coefficients, which can more accurately locate the areas in the image that play a key role in the classification decisions of multiple related categories, so as to determine the attention area of the adversarial samples, and then perform targeted denoising on the robust attention area. This method effectively avoids the loss of key information caused by undifferentiated denoising of the entire image and greatly retains the core details of the image.

[0033] (3) Strong adaptability: The attention area denoising and adversarial training methods of the present invention have good adaptability and can be flexibly adjusted according to different data sets, model structures, and attack types, and are applicable to the security protection of various classification models. Description of the Drawings

[0034] Figure 1 is the overall flowchart of the present invention;

[0035] Figure 2 is the selection logic diagram of the attack method in the dynamic adversarial training of the present invention;

[0036] Figure 3 is the core processing flowchart based on wavelet transform of the present invention;

[0037] Figure 4 is the schematic diagram of threshold processing of the present invention;

[0038] Figure 5 is the schematic diagram of the details of wavelet denoising of the present invention;

[0039] Figure 6This is the flowchart of the adversarial training of the present invention;

[0040] Figure 7 This is the schematic diagram of the technical implementation principle of the present invention;

[0041] Figure 8 This is the schematic diagram of the adversarial sample obtained by using the FGSM method in the embodiment;

[0042] Figure 9 This is the heat map of the adversarial sample obtained in the embodiment;

[0043] Figure 10 This is the image after denoising the robust attention area in the embodiment. Detailed implementation manners

[0044] The technical solution of the present invention will be described in detail below, but the protection scope of the present invention is not limited to the described embodiments.

[0045] As Figure 1 shown, the classification model defense method based on dynamic adversarial training and attention area denoising of the present invention includes the following steps:

[0046] Step 1, in the adversarial training stage of the classification model, prepare a classification model (such as ResNet-18), a training data set, and various adversarial attack methods, perform dynamic adversarial training on the original basic model (such as ResNet-18), optimize the model parameters through backpropagation, dynamically adjust the learning rate, and calculate the loss using the cross-entropy loss function;

[0047] Step 2, in the robust attention area denoising stage, first generate adversarial samples of the test data, and use the class activation map CAM to determine the robust attention area of the adversarial samples, and then perform denoising processing on the obtained robust attention area to weaken the influence of adversarial perturbations. The specific steps are as follows:

[0048] Step 2.1, use various adversarial attack methods to generate adversarial samples of the test data; for example, only use FGSM in the early stage, randomly select FGSM and PGD in the middle stage, and randomly select FGSM, PGD, and C&W in the later stage;

[0049] Step 2.2, comprehensively consider the information of the top k categories and generate the final CAM by weighted summation to determine the robust attention area. The class activation map is calculated by the following formula:

[0050]

[0051] where is the comprehensive weight coefficient obtained by optimizing the weights of the fully connected layers corresponding to the top k categories, and f k(x) is the k-th channel of the feature map; the value of k depends on the characteristics of the dataset and a reasonable value needs to be selected through experimental testing to ensure that key regional information related to the true classification of adversarial samples can be effectively captured;

[0052] Step 2.3, Denoise the robust attention region of the adversarial sample to weaken the influence of the adversarial perturbation. Here, local average denoising or wavelet denoising methods are used;

[0053] Step 2.4, Perform weighted fusion of the denoised image and the original image. The formula is:

[0054] X D = α·X d + β·X c

[0055] where α and β are weight coefficients, usually set to 0.7 and 0.3, X d is the denoised image, and X c is the clean image;

[0056] Step 3, In the classification stage, input the adversarial sample after denoising the robust attention region into the model that has completed adversarial training, and then classification can be performed to obtain the final result and detect the classification accuracy rate of the model.

[0057] The robust attention region denoising method of this embodiment is different from the traditional method that solely relies on the highest probability category to locate the attention region. Instead, it uses the class activation mapping (CAM) technology and conducts targeted optimization. Considering the characteristic that adversarial samples may keep the top k most likely categories unchanged in the image scenario with adversarial perturbations, by integrating the information of the top k categories, through in-depth analysis of the feature map of the model and the weights of the fully connected layer, calculate the comprehensive contribution of each pixel point to the final classification results of multiple relevant categories, so as to more comprehensively and accurately find the attention regions that play a key role in classification. The robust localization method can effectively avoid the problem of misjudging the highest probability category caused by adversarial perturbations and then wrongly determining the attention region. The steps of the above-mentioned robust attention region denoising method are as follows:

[0058] S001, Obtain the output of the last convolutional layer of the classification model as the feature map and record the weights of the fully connected layer. For a specific dataset, in combination with its category characteristics, reasonably select the top k categories through experimental testing.

[0059] S002, For each channel in the feature map, perform weighted summation with the weights of the fully connected layer corresponding to the top k categories respectively to obtain k two-dimensional class activation mapping graphs. Then perform exponential weighted averaging on these k mapping graphs to highlight the mapping influence of the category with higher correlation to the correct classification, thereby obtaining a comprehensive CAM graph.

[0060] S003. Normalize the comprehensively obtained CAM map so that its value range is between [0, 1]. Different from the traditional fixed threshold method, this method dynamically determines the threshold according to the statistical characteristics of the CAM map (combining the noise, perturbations, etc. of the dataset images, and determining the threshold based on adaptive methods such as the Bayesian shrinkage algorithm), and marks the regions in the CAM map greater than the threshold as the regions of interest.

[0061] In this embodiment, the wavelet denoising method is used to denoise the robust region of interest, and the image of the robust region of interest can be decomposed into sub-bands of different frequencies. Adversarial perturbations usually concentrate in the high-frequency part of the image. By performing wavelet transform on the region of interest, separating the high-frequency part and the low-frequency part, then performing threshold processing on the wavelet coefficients of the high-frequency part to remove the adversarial perturbations therein, and finally performing inverse wavelet transform to restore the image. The specific steps are as follows:

[0062] S001. Select a suitable wavelet basis, such as the Haar wavelet basis, and perform two-dimensional discrete wavelet transform on the region of interest to decompose the image into a low-frequency sub-band and three high-frequency sub-bands (horizontal, vertical, and diagonal directions).

[0063] S002. Use the Bayesian shrinkage algorithm to determine the threshold to process the wavelet coefficients. The Bayesian shrinkage algorithm can adaptively determine the threshold according to the statistical characteristics of the wavelet coefficients, effectively removing noise while retaining the important features of the image.

[0064] S003. Perform inverse wavelet transform on the processed wavelet coefficients to obtain the denoised image of the region of interest.

[0065] When this embodiment performs weighted fusion of the denoised image and the original image, it is necessary to perform weighted fusion according to the preset weights, which can not only retain the overall information of the original image but also make full use of the advantage of the denoised image of the region of interest in removing adversarial perturbations, and avoid losing key information due to excessive denoising. Specific steps: Set the weight α of the denoised image and the weight β of the original image, where the value range of α is 0.6 - 0.8, the value range of β is 0.2 - 0.4, and α + β = 1. Perform fusion on the pixel values at the corresponding positions of the denoised image of the region of interest and the original image according to the weighted formula X D = α·X d + β·X c to obtain the final denoised image, where X D is the fused image, X d is the denoised image of the region of interest, and X c is the original image.

[0066] During the dynamic adversarial training process of this embodiment, different adversarial sample generation methods are dynamically selected according to the number of training rounds. In the initial stage of training, the model's learning of data features is not yet sufficient. At this time, the FGSM attack, which is simple to calculate and fast, is selected to generate adversarial samples, which can quickly guide the model to learn adversarial features. As the training progresses, the model gradually acquires a certain ability to resist attacks. The PGD attack is introduced to generate more challenging adversarial samples, further enhancing the model's robustness. In the later stage of training, the C&W attack is used to generate high-intensity adversarial samples, enabling the model to cope with relatively more complex attack scenarios. This dynamic adjustment strategy avoids the overfitting of the model to a single attack method and improves the model's generalization ability to different types of attacks.

[0067] Division of training rounds: The entire training process is divided into three stages. The first stage is when the number of training rounds is less than 30% of the total number of training rounds. The second stage is when the number of training rounds is greater than or equal to 30% and less than 60% of the total number of training rounds. The third stage is when the number of training rounds is greater than or equal to 60% of the total number of training rounds.

[0068] Selection of attack methods: In the first stage, the FGSM attack is used to generate adversarial samples. The perturbation intensity ε of FGSM can be adjusted according to the characteristics of the dataset and the model, and generally ranges from 0.1 to 0.3. In the second stage, the FGSM or PGD attack is used to generate adversarial samples. The perturbation intensity ε of PGD is similar to that of FGSM, and the step size α ranges from 0.01 to 0.05, and the number of iteration steps steps ranges from 20 to 50. In the third stage, the C&W attack is introduced, and any one of the methods can be randomly selected to generate adversarial samples. The parameter c of C&W ranges from 0.1 to 1, kappa is set to 0, and the number of iteration steps steps ranges from 500 to 1500.

[0069] Generation and training of adversarial samples: In each training batch, the corresponding attack method is selected according to the current training stage to generate adversarial samples. The generated adversarial samples are concatenated with the original clean samples as the training input, and the cross-entropy loss function is used to calculate the loss, and the model parameters are updated through an optimizer (such as the Adam optimizer).

[0070] To verify the feasibility and effectiveness of the present invention, the technical solution of the present invention is applied in this embodiment, including the following steps:

[0071] 1. Model training stage:

[0072] (1) Data preparation: Use the CIFAR-10 dataset as the training and test data.

[0073] The CIFAR-10 dataset contains 60,000 32x32 color images in 10 categories, with 50,000 for training and 10,000 for testing. After reading the data, the following preprocessing operations are performed

[0074] · Random horizontal flipping: Flip the image with a 50% probability.

[0075] · Normalization: Normalize the R, G, and B channels of the image separately

[0076] · Cutout: Randomly occlude a part of the image to enhance the robustness of the model.

[0077] These operations enhance the diversity of the training data and improve the generalization ability of the model. The Dataloader processes the dataset in batches, with each batch size of 64.

[0078] (2) Model construction: Use ResNet-18 as the base model. Considering the image size of CIFAR-10 is 32x32, the following modifications are made to the model: Change the convolutional kernel size of the first convolutional layer from 7x7 to 3x3, the stride from 2 to 1, and the padding to 1. Modify the output dimension of the last fully connected layer to 10 categories.

[0079] (3) Dynamic adversarial training stage:

[0080] a) The training set for adversarial training includes clean samples X c and adversarial samples X adv , where the adversarial sample generation method is dynamically selected according to the number of training epochs:

[0081] In the early stage (epoch < 60): Only use FGSM for attack, as Figure 8 shown.

[0082] In the middle stage (60 ≤ epoch < 120): Randomly select FGSM and PGD to generate adversarial samples.

[0083] In the late stage (epoch ≥ 120): Randomly select FGSM, PGD, and C&W to generate adversarial samples.

[0084] FGSM is a classic white-box adversarial attack method. First, calculate the gradient of the model loss function with respect to the input x Use the sign function to determine the perturbation direction to maximize the loss function, and add a perturbation to the original input X C to generate an adversarial sample. The formula for the FGSM attack is as follows:

[0085]

[0086] where, X advis the generated adversarial sample; X c is the original input sample; ε is the perturbation size; J(x, y) represents the function loss, represents the gradient of the loss function with respect to x; sign represents the sign function; y represents the true label of the clean sample.

[0087] PGD is an iterative optimization adversarial attack method that gradually adjusts the input through multiple iterations to deviate it from the correct classification. First, using the original sample X C as the starting point, calculate the gradient of the loss function with respect to the input, and along the direction of the gradient, make a small modification to the input to make it closer to the decision boundary of the model, obtaining the updated adversarial sample Perform a "projection" operation on the updated sample to ensure that it remains within a predefined perturbation range, which is controlled by a parameter representing the maximum allowed modification amount. Repeat the above steps until the maximum number of iterations is reached or the termination condition is satisfied. The attack formula of PGD is as follows:

[0088]

[0089] where is the adversarial sample after the (t + 1)-th iteration, that is, the new value of the input after the attack; is the value of the current adversarial sample at the t-th iteration; α controls the size of the attack perturbation each time and determines the amplitude of the update of the adversarial sample each time. If the step size is too small, the attack effect may not be obvious; if the step size is too large, it will affect the attack quality; is the gradient direction of the loss function; Clip x,ε is to prevent the change of the adversarial sample from exceeding the maximum perturbation range and affecting the perceptibility.

[0090] The C&W attack is a powerful white-box adversarial attack method that generates imperceptible adversarial samples by minimizing the L2 norm of the perturbation. First, define the objective function f(x + δ) to measure the classification error degree of the adversarial sample; by optimizing min δ ||δ|| 2 + c·f(x + δ), while ensuring the classification error, minimize the L2 norm of the perturbation. Adjust the parameters c and κ to control the strength and confidence of the attack. The C&W attack objective function is as follows:

[0091] min δ ||δ|| 2 + c·f(x + δ)

[0092] f(x + δ) = max(max{Z(x + δ) i : i ≠ t} - Z(x + δ) t , -κ)

[0093] c·f(x + δ) represents the target network and adjusts the confidence of classification error by adjusting the size of κ. The adversarial examples generated by this method have a certain degree of transferability, but it takes a lot of time.

[0094] b) Concatenate the original image and the adversarial example to form the training input X mix :

[0095] X mix = torch.cat((X c , X adv ), dim = 0)

[0096] The target labels also need to be concatenated to ensure that each input corresponds to the same label. The concatenated target label Y mix is expressed as:

[0097] Y mix = torch.cat((y, y), dim = 0)

[0098] c) Loss calculation

[0099] During training, use the cross-entropy loss function to calculate:

[0100]

[0101] where L train is the training loss, N is the number of samples, CE is the cross-entropy loss function, is the prediction of the model for the mixed samples, and Y i is the true label

[0102] (4) Optimizer and learning rate scheduling:

[0103] a) Use the Adam optimizer to update the model parameters. The initial learning rate is 0.001, and the weight decay is 5e-4. Halve the learning rate every 5 epochs. The learning rate scheduling formula is:

[0104]

[0105] where, lr t is the learning rate for the t-th training, γ is the learning rate decay factor, and stepsize is the step length of learning rate decay.

[0106] b) During the training process, if the validation loss decreases, save the model parameters to the specified path

[0107] 2. Region of interest denoising stage

[0108] (1) Adversarial example generation:

[0109] Generate adversarial samples of test data using multiple adversarial attack methods (such as FGSM, PGD, C&W) to simulate diverse attacks.

[0110] (2) Class Activation Map (CAM) generation:

[0111] Adversarial samples often exhibit the characteristic of relative stability in the top-k most likely classes. Therefore, by integrating the information of the top-k classes, a weighted average class activation map is obtained, which can more accurately locate the attention regions crucial for classification and effectively avoid the problems of misjudgment of the highest probability class and misidentification of the attention region caused by adversarial perturbations in traditional methods.

[0112] a) Extract the output of the last convolutional layer of the classification model as the feature map and record the weights of the fully connected layer. For specific datasets with different characteristics, multiple groups of experimental tests are carried out according to their class distribution characteristics. Taking the CIFAR-10 dataset as an example, through repeated experimental verification, selecting k = 3 can efficiently capture the key region information closely related to the true classification of adversarial samples, and this value is determined based on in-depth research on the characteristics of this dataset and adversarial samples.

[0113] b) For each channel f k (x) of the feature map, perform weighted summation with the comprehensive weight coefficients optimized according to the dataset characteristics corresponding to the top-k classes respectively, and the formula is as follows:

[0114]

[0115] Generate k two-dimensional class activation mapping diagrams through the operation, and then perform exponential weighted average on these mapping diagrams to generate a comprehensive CAM diagram to comprehensively reflect the key information related to multiple relevant classes of the image.

[0116] c) Normalize the comprehensively obtained CAM diagram to the interval [0, 1], as Figure 9 shown. Different from the traditional fixed threshold method, the present invention dynamically determines the threshold by using an adaptive method such as the Bayesian shrinkage algorithm according to the statistical characteristics of the CAM diagram, combined with the actual situations such as the image noise level and adversarial perturbation intensity of the dataset, and marks the regions in the CAM diagram greater than the dynamic threshold as the attention regions.

[0117] (3) Denoising of attention regions:

[0118] Wavelet shrinkage is a process based on the sparse characteristics of real-world signals by discrete wavelet transform, leveraging the advantages of wavelet denoising to perform denoising processing on the attention regions of adversarial samples and weaken the impact of adversarial perturbations.

[0119] a) Use the Bayesian shrinkage algorithm as an adaptive method for wavelet soft thresholding to determine the threshold by comparing the relationship between the noise variance and the signal plus noise variance.​

[0120] b) X c is a clean image, X adv is an adversarial example, then X adv = X c + ρ, where ρ is the added perturbation. Let the wavelet sub - band variance of the adversarial example be Then there is:

[0121]

[0122] where, W m is the sub - band wavelet and M is the total number of wavelet coefficients in the sub - band.

[0123] c) Soft threshold t bs is calculated as follows:

[0124]

[0125] If the noise variance is less than the signal - plus - noise variance then it is considered that the coefficients in this wavelet sub - band are mainly composed of signals, and the threshold is selected as Conversely, if the noise variance is large, it is considered that the coefficients in this wavelet sub - band are mainly composed of noise, and the threshold is selected as the maximum amplitude of the coefficients in this wavelet sub - band.

[0126] (4) Denoising result fusion:

[0127] Fuse the denoised image (as shown in Figure 10 ) with the original image with weights. The formula is:

[0128] X D = α·X d + β·X c

[0129] where α and β are weight coefficients, usually set to 0.7 and 0.3, X d is the denoised image, and X c is the clean image.

[0130] 3. Classification stage

[0131] Load the model weights after adversarial training, and input the denoised adversarial example in the region of interest into the model after adversarial training for classification. The classification result is obtained through the following formula

[0132]

[0133] Calculate the classification accuracy of the model for the denoised adversarial example to evaluate the robustness of the model. The following shows the test correct rate of using the present invention and the jpeg method for defense:

[0134] Table 1 Test Results of Region Denoising and Adversarial Training

[0135] FGSM PGD C&W No defense 63% 62% 61% jepg 74% 72% 66% This method 83% 78% 79%

[0136] It can be found that this method can greatly improve the classification accuracy of adversarial samples.

[0137] The present invention enhances the adversarial robustness of the deep model by combining dynamic adversarial training and region-of-interest denoising. The core idea of this method is as follows: through dynamic adversarial training, the adaptability of the model to different types of adversarial attacks is improved, overfitting to specific attack methods is avoided, and at the same time, the key regions of interest in the image are located using the Class Activation Map (CAM), and denoising processing is performed on these regions to weaken the influence of adversarial perturbations, thereby improving the classification ability of the model for adversarial samples.

Claims

1. A classification model defense method based on dynamic adversarial training and focus area denoising, characterized in that: The following steps are included: Step 1: In the adversarial training stage of the classification model, first prepare the classification model, training data set and multiple adversarial attack methods, perform dynamic adversarial training on the original classification model, optimize the model parameters through back propagation, dynamically adjust the learning rate, and use the cross entropy loss function to calculate the loss; Step 2: In the robust attention region denoising stage, the adversarial sample of the test data is first generated, and the class activation map CAM is used to determine the robust attention region of the adversarial sample. Then, the obtained robust attention region is denoised to weaken the impact of the adversarial perturbation. The specific steps are as follows: Step 2.1, use multiple adversarial attack methods to generate adversarial samples of test data; Step 2.2: Comprehensively consider the information of the first k categories and weight them to generate the final CAM class activation map to determine the robust attention area. The calculation formula of the class activation map is as follows: in, is the comprehensive weight coefficient obtained after optimizing the weights of the fully connected layer corresponding to the first k categories, f k (x) is the kth channel of the feature map; Step 2.3, denoising the robust attention area of ​​the adversarial sample; Step 2.4: Weighted fusion of the denoised image and the original image. The formula is: X D =α·X d +β·X c Among them, α and β are weight coefficients, usually set to 0.7 and 0.3, d is the denoised image, X c For clean images; Step 3: In the classification stage, the adversarial samples that have been denoised by the robust attention area are input into the model that has completed adversarial training, and the final result can be obtained by classification to detect the classification accuracy of the model.

2. The classification model defense method based on dynamic adversarial training and focus area denoising according to claim 1 is characterized in that: The method for constructing the data set in step 1 is as follows: selecting a data set and dividing it into a training set and a test set, and randomly horizontally flipping, standardizing, and randomly occluding the images; Use ResNet-18 as the original classification model and modify the model appropriately considering the image size of the dataset; The steps of dynamic adversarial training are as follows: Step a), the training set of adversarial training includes clean samples X c With adversarial sample X adv ; Step b) Use the torch.cat() function to perform tensor concatenation operations to concatenate clean samples and adversarial samples along the batch dimension to form a training input mixed sample X mix , the true category of the adversarial sample is the same as the corresponding clean sample, so the target label itself is concatenated along the batch dimension to obtain the mixed target label Y mix ; Step c), obtain the spliced ​​mixed sample X mix With mixed target label Y mix After that, start training and use the cross entropy loss function L train To calculate the loss, the formula is as follows: N is the total number of samples.

3. The classification model defense method based on dynamic adversarial training and focus area denoising according to claim 1 or 2, characterized in that: Step 1: The steps of optimizer selection and learning rate scheduling in the adversarial training phase of the classification model include: using the Adam optimizer to update the model parameters, with an initial learning rate of 0.001 and a weight decay of 5e-4; the learning rate is halved every 5 epochs, and if the verification loss is reduced, the model parameters are saved.

4. The classification model defense method based on dynamic adversarial training and focus area denoising according to claim 1 is characterized in that: Step 3: The steps in the classification phase include: Load the adversarial trained model weights, input the adversarial samples after robust attention region denoising into the adversarial trained model for classification; calculate the classification accuracy of the model for the denoised adversarial samples to evaluate the robustness of the model.

Citation Information

Cited By

  • Double-stage decoupling robust adversarial training method

    CN121257648A