Adversarial training defense method based on adversarial label guidance

Through an adversarial training method based on adversarial label-guided, the problem of fragility of adversarial attacks in the image processing field is solved, and the high accuracy classification of adversarial samples and the robustness of the model is improved.

CN119940468APending Publication Date: 2025-05-06YUNNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510028191.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Deep learning models are fragile in the field of image processing, and the existing adversarial training methods are insufficient, and the distribution and characteristics of samples after adversarial attacks are not fully explored.

Method used

Adversarial training method based on adversarial tags is adopted to improve the robustness of the model by generating adversarial samples, extracting adversarial attack characteristics, generating adversarial tags and counting their tag category probability distribution, adding regularization terms to the loss function, and imposing constraints on the current adversarial tags.

Benefits of technology

It significantly improves the classification accuracy of adversarial samples, enhances the robustness and generalization capabilities of the model, and can more effectively defend against adversarial attacks, especially when the model capacity is limited.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940468A_ABST
    Figure CN119940468A_ABST
Patent Text Reader

Abstract

The invention introduces an adversarial training defense method based on adversarial label guidance, and aims to enhance the ability of a deep learning model to resist adversarial attacks. By generating an adversarial sample, extracting adversarial features and generating an adversarial label, the method introduces a regularization item into a loss function to constrain the adversarial label, so that adversarial sample prediction is closer to a real label. CIFAR-10 and CIFAR-100 data sets are used in experiments, and compared with an existing method, the method provided by the invention not only improves the classification accuracy of adversarial samples, but also realizes good balance between natural accuracy and robustness, and especially shows excellent defensive performance when facing various attacks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of adversarial defense in artificial intelligence and machine learning, and in particular relates to an adversarial training defense method guided by adversarial labels. Background Art

[0002] With the widespread application of deep neural networks (DNNs), they have achieved remarkable success in image recognition, natural language processing and other fields. However, studies have found that DNN adversarial examples are very vulnerable. Adversarial examples are generated by applying tiny and human-imperceptible perturbations to the input data, which can cause the model to output incorrect predictions. This vulnerability has aroused widespread concern about the security and robustness of DNNs. To deal with these attacks, researchers have proposed a variety of defense strategies, including adversarial training, noise denoising, defensive distillation, gradient regularization / shielding, and detection-only methods. Among them, adversarial training is considered to be one of the most effective defense strategies.

[0003] Due to the effectiveness of adversarial training, many related studies have started from different perspectives to further enhance the adversarial robustness of DNN by improving the performance of adversarial training. For example, TRadeoff-inspired Adversarial DEfense via Surrogate-loss minimization (TRADES) method, Geometry-Aware Instance-Reweighted Adversarial Training (GAIRAT) method, Friendly Adversarial Training (FAT), Misclassification-Aware Adversarial Training (MART), free adversarial training (free-AT), TRADES achieves a trade-off between the robustness and accuracy of the model by minimizing a loss function consisting of two parts; FAT alleviates the cross-mixing problem and focuses on generating "friendly" adversarial samples to improve the natural generalization ability of the model; free adversarial training (free-AT) improves the generalization ability of the model by simultaneously updating the parameters of the model and the perturbations of the image; MART proposes a weighted loss function from the perspective of misclassified samples; the GAIRAT method considers the problem from a geometric perspective, believing that different data points have different importance and contributions to robustness due to their different positions in the data space, and proposes to adjust the weights of different data points through geometric perception; and the SAT regularization method uses the Taylor expansion of small Gaussian noise to better achieve a balance between robustness and accuracy.

[0004] Although previous studies have pointed out that it is necessary to pay attention to misclassified data in model training, the distribution of samples after adversarial attacks has not been studied in depth. The decision boundary is only optimized by learning the misclassified samples near the decision boundary. Therefore, past research still has shortcomings, and adversarial training also faces some challenges and limitations. The present invention further studies the characteristics of adversarial attacks and improves adversarial training accordingly. Summary of the invention

[0005] The purpose of the present invention is to provide an adversarial training defense method based on adversarial label guidance to solve the vulnerability problem of deep learning models facing adversarial attacks in the field of image processing.

[0006] The technical solution adopted by the present invention is an adversarial training defense method based on adversarial label guidance, comprising the following steps:

[0007] Step S1, image data preparation: generating adversarial samples;

[0008] Step S2: extract the features of the adversarial attack, generate adversarial labels, and count the label category probability distribution of the adversarial labels;

[0009] Step S3: conduct adversarial training based on the adversarial label guidance, add a regularization term to the loss function, and impose constraints on the current adversarial label.

[0010] Furthermore, the specific process of step S1 is as follows:

[0011] Step S11: Get the original input sample x from the public dataset cifar10 i and its corresponding true label y i , as the target sample;

[0012] Step S12: Select the loss function L1 and calculate the loss on the original sample x, specifically:

[0013]

[0014] Where L1(·) represents the loss function, x represents the input sample, y represents the true label, C represents the total number of label categories, and y i represents the actual label probability of the i-th label category, Represents the probability of the i-th label category predicted by the model for the input sample x, and i represents the current sequence number of the traversed label category;

[0015] Step S13: Based on the loss function obtained in step S12, the gradient of the loss function with respect to the current input sample is calculated. The specific formula is as follows:

[0016]

[0017] Among them, g represents the gradient of the loss function to the current input sample, represents the gradient of the input sample x, and L1(x,y) is the loss function obtained in step S12;

[0018] Step S14: Based on the gradient calculated in step S13, generate a preliminary update of the adversarial sample. The specific update formula is:

[0019] x'=x+a·sign(g)

[0020] Where x' represents the adversarial sample, x represents the input sample, a represents the step size, which can control the magnitude of each update; sign(g) represents the sign function of the gradient;

[0021] Step S15: Apply the projection operation to ensure that the generated adversarial sample is within the allowed perturbation range. The specific operation is:

[0022] x'=clip(x',x-ε,x+ε)

[0023] Where x' represents the adversarial sample, ε represents the maximum perturbation amplitude allowed, and clip(·) represents the clipping function;

[0024] Step S16: Repeat steps S12 to S15 for multiple iterations until the predetermined number of iterations is reached or the value of the loss function no longer changes significantly, generating the final adversarial sample.

[0025] Furthermore, the specific process of feature extraction and adversarial label generation in step S2 is as follows:

[0026] Step S21a, calculating the input data in sequence through the convolution layer, the fully connected layer and the activation layer to extract features;

[0027]

[0028] Among them, F i l represents the i-th feature map of the l-th layer, is the convolution kernel weight of layer l, represents the j-th feature map of the l-1th layer, represents the bias term, C in Indicates the input channel, is the activated feature map, ReLU(·) represents the activation function; max(0,F i l ) means that when the input sample is greater than 0, the value is directly output, and when the input sample is less than 0, 0 is output;

[0029] Step S21b: Generate a feature map and flatten the feature map into a one-dimensional vector

[0030]

[0031] in, represents the flattened one-dimensional vector, A L represents the output feature map of the last convolutional layer, and flatten(·) represents the flattening operation;

[0032] Step S21c: vector Perform linear transformation to generate the final feature vector;

[0033]

[0034] in, represents the flattened one-dimensional vector, represents the output feature vector, W represents the weight matrix, represents the bias vector;

[0035] Step S21d: At the output layer, the model converts the feature vector into a prediction result to obtain the probability distribution of the predicted label category;

[0036]

[0037] Among them, P i represents the predicted probability of the i-th label category, C is the total number of label categories, z i represents the output value of the i-th label category, i represents the current sequence number of the traversed label category, and softmax(·) represents a function that converts any real number vector into a probability distribution. Represents the exponential function of the i-th label category in the vector z, that is, The number of values ​​of the i-th label category in ;

[0038] Step S21e, recording the most likely result in the adversarial prediction output as an adversarial label;

[0039] y adv = argmax(p i )

[0040] Among them, y adv represents the adversarial label, P i represents the predicted probability of the i-th label category, and argmax(·) represents the function of selecting the label category with the largest probability as the adversarial label.

[0041] Furthermore, in step 2, based on the CIFAR-10 dataset, the probability distribution of the label categories of the adversarial labels is statistically analyzed, and the results of natural training and standard adversarial training are statistically analyzed using ResNet-18, and the number of samples such as 1st, 2nd, 3rd, ..., 10th are counted; among them, 1st represents the number of samples with high accuracy and robustness, and the remaining 2nd, 3rd, ..., 10th represent the ranking of the number of successfully attacked adversarial samples, and also represent the number of samples with high accuracy but low robustness.

[0042] Furthermore, the specific process of step 3 is as follows:

[0043] S31. Input the original sample and the generated adversarial sample into the model and perform forward propagation to obtain the true label and adversarial sample label. The formula is as follows:

[0044] y=f(x,θ)

[0045]

[0046] Among them, f(·) represents the forward propagation function of the model, y represents the true label, represents the adversarial sample label, θ represents the model parameter, x represents the original sample, represents adversarial examples;

[0047] S32. Design a loss function and calculate the loss. Calculate the difference between the adversarial sample prediction and the true label. The formula is:

[0048]

[0049] Among them, L CE (·) is the cross entropy loss function, which means calculating the adversarial sample prediction The loss value between the actual label y; y i represents the actual label probability of the i-th label category, represents the predicted probability of label category i, and C represents the total number of label categories;

[0050] Then the cross entropy function is modified to obtain the inverse cross entropy loss function, the formula is:

[0051]

[0052] Among them, L ICE (·) is the inverse cross entropy loss function, which means calculating the adversarial sample prediction With adversarial label y adv The loss value between represents the adversarial label of label category i, represents the predicted probability of label category i, C represents the total number of label categories, and λ is a constant used to prevent the denominator from being too small;

[0053] When L CE (·) and L ICE (·) When used together, on the premise that natural samples are correctly classified, the regularization term is designed as follows:

[0054] L ALAT =L CE (p(x adv ,θ),y)+L ICE (p(x adv ,θ),y adv )

[0055] Among them, p(·) represents the predicted value of the model, x adv represents the input variable, θ represents the parameters of the model, y represents the true label, and y adv represents the adversarial label, L ALAT represents the regularization term, L CE (·) represents the cross entropy loss function, L ICE (·) represents the inverse cross entropy loss function;

[0056] The total loss function involved in the adversarial training technique based on adversarial labels is as follows:

[0057] L total =L CE (p(x,θ),y)+β·L ALAT

[0058] Among them, L total Represents the total loss function, L CE(·) represents the cross entropy loss function, P(·) represents the predicted value of the model, x represents the input sample, θ represents the parameter of the model, y represents the true label, β represents the adjustable scaling parameter that balances the final loss of the two parts, and L ALAT represents the regularization term;

[0059] S33, back propagation and parameter update, the gradient of the composite loss function to the model parameter θ is calculated through back propagation, the formula is as follows:

[0060]

[0061] in, represents the gradient of parameter θ, L total Represents the total loss function, L CE(·) represents the cross entropy loss function, P(·) represents the predicted value of the model, x represents the input sample, θ represents the parameter of the model, y represents the true label, β represents the adjustable scaling parameter that balances the final loss of the two parts, and LALAT represents the regularization term;

[0062] Use the gradients to update the model parameters:

[0063]

[0064] Among them, θ t+1 represents the updated model parameters, θ t represents the model parameters at the tth iteration, η represents the learning rate, ▽ θ represents the gradient of the parameter θ, L total Represents the total loss function.

[0065] The beneficial effects of the present invention are:

[0066] 1. This paper proposes a new adversarial training method, namely, Adversarial Label Guided Adversarial Training (ALAT). Compared with the prior art, the Adversarial Label Guided Adversarial Training (ALAT) method significantly improves the classification accuracy of adversarial samples under the condition of limited model capacity.

[0067] 2. Add an additional regularization term in the adversarial optimization process to incorporate the adversarial label. This technical feature guides the prediction of adversarial samples closer to the true label, increases the possibility of natural prediction labels, and effectively moves away from the current adversarial label.

[0068] 3. The present invention not only provides a more effective adversarial training approach for the model, but also shows significant robustness improvement in experiments on CIFAR-10 and CIFAR-100 datasets, and has obvious effectiveness and superiority in practical applications.

[0069] 4. Based on the classification behavior of the model for natural samples and adversarial samples, this paper proposes four possible scenarios of the model: being able to correctly predict both natural samples and adversarial samples, being able to correctly predict only natural samples, being unable to correctly predict both natural samples and adversarial samples, and being able to correctly predict only adversarial samples. By distinguishing the performance of the model on natural samples and adversarial samples, we can analyze the defects of the model in more detail, better understand the impact of adversarial attacks, and formulate more targeted improvement strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0071] Figure 1This is an illustration of generating adversarial samples.

[0072] Figure 2 There are four situations to counter the attack.

[0073] Figure 3 This is an illustration of ALAT.

[0074] Figure 4 is the definition graph of the adversarial label.

[0075] Figure 5 This is the distribution of adversarial labels in CIFAR10. DETAILED DESCRIPTION

[0076] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0077] Example 1

[0078] like Figure 1-5 As shown, the present invention provides an adversarial training defense method based on adversarial label guidance, and the specific steps include:

[0079] Step S1, image data preparation, that is, generating adversarial samples, using the projected gradient descent (PGD) method to generate adversarial samples and improve the robustness of the model. PGD is an iterative optimization technique that aims to maximize the prediction error of the model by applying small perturbations to the input samples. The cifar10 benchmark dataset is used, which contains 60,000 color images divided into 10 categories, with 6,000 images in each category.

[0080] Step S11: Get the original input sample x from the public dataset cifar10 i and its corresponding true label y i , as the target sample.

[0081] Step S12: Select the loss function L1 (the present invention selects the cross entropy loss function) to calculate the loss on the original sample x, specifically:

[0082]

[0083] Where L1(·) represents the loss function, x represents the input sample, y represents the true label, C represents the total number of label categories, and y i represents the actual label probability of the i-th label category, It represents the probability of the i-th label category predicted by the model for the input sample x, and i represents the current sequence number of the traversed label category.

[0084] Step S13: Based on the loss function obtained in step S12, the gradient of the loss function to the current input sample is calculated to determine how to adjust the input sample to increase the prediction error of the model. The specific formula is as follows:

[0085]

[0086] Among them, g represents the gradient of the loss function to the current input sample, represents the gradient of the input sample x, and L1(x, y) is the loss function obtained in step S12.

[0087] Step S14: Based on the gradient calculated in step S13, generate a preliminary update of the adversarial sample. The specific update formula is:

[0088] x'=x+a·sign(g)

[0089] Among them, x' represents the adversarial sample, x represents the input sample, a represents the step size, which can control the amplitude of each update; sign(g) represents the sign function of the gradient, which can return the positive and negative signs of the gradient.

[0090] Step S15: Apply the projection operation to ensure that the generated adversarial sample is within the allowed perturbation range. The specific operation is:

[0091] x'=clip(x',x-ε,x+ε)

[0092] Here, x' represents the adversarial sample, ε represents the maximum perturbation amplitude allowed, and clip(·) represents the clipping function, which can limit the input value within the specified limit to ensure that the adversarial sample does not exceed the perturbation range of the original sample.

[0093] Step S16: Repeat steps S12 to S15 for multiple iterations until the predetermined number of iterations is reached or the value of the loss function no longer changes significantly, generating the final adversarial sample.

[0094] Step S2: extract the features of the adversarial attack, generate adversarial labels, and calculate the probability distribution of the label categories of the adversarial labels.

[0095] When the model inputs natural samples and adversarial samples, the present invention can divide the output results into four cases, such as Figure 2 As shown. and adversarial prediction labels The results are classified as follows:

[0096] 1. Ideal situation and The model can correctly predict natural samples and adversarial samples. The model can effectively recognize and classify images of various different label categories and maintain good performance in the face of adversarial samples. This is our ideal high-performance robust model.

[0097] 2. Natural samples are accurate but adversarial samples are poor. and When , the model can successfully predict natural samples, but performs poorly on adversarial samples. This situation is common in natural training models that have good performance but lack robustness.

[0098] 3. Both natural samples and adversarial samples are poor. and When the model is not able to process natural samples effectively, the first issue is not to improve the robustness of the model, but to improve the model's processing performance on natural samples.

[0099] 4. The situation that does not exist is and The goal of white-box adversarial attack is to generate adversarial examples based on loss maximization. Therefore, the situation where the robust prediction loss is greater than the natural prediction loss will not happen.

[0100] Among the above four situations, represents the natural prediction label, represents the predicted probability of label category i, y i Indicates the actual label probability of the i-th label category. In the following, the feature analysis is mainly carried out on the data of the second case.

[0101] Step S21: Feature extraction and generation of adversarial labels. In order to observe the characteristics of adversarial attacks, input data with predicted probabilities are input into the model for forward propagation. The specific process is as follows:

[0102] Step S21a: The input data is calculated in sequence through the convolution layer, the fully connected layer and the activation layer to perform feature extraction.

[0103]

[0104] Among them, F i l represents the i-th feature map of the l-th layer, is the convolution kernel weight of layer l, represents the j-th feature map of the l-1th layer, represents the bias term, C in Indicates the input channel, is the activated feature map, ReLU(·) represents the activation function; It means that when the input sample is greater than 0, the value is directly output, and when the input sample is less than 0, 0 is output.

[0105] Step S21b: Generate a feature map and flatten the feature map into a one-dimensional vector

[0106]

[0107] in, represents the flattened one-dimensional vector, A L represents the output feature map of the last convolutional layer, and flatten(·) represents the flattening operation.

[0108] Step S21c: vector Perform linear transformation to generate the final feature vector.

[0109]

[0110] in, represents the flattened one-dimensional vector, represents the output feature vector, W represents the weight matrix, Represents the bias vector.

[0111] Step S21d: At the output layer, the model converts the feature vector into a prediction result to obtain the probability distribution of the predicted label category.

[0112]

[0113] Among them, P i represents the predicted probability of the i-th label category, C is the total number of label categories, z i represents the output value of the i-th label category, i represents the current sequence number of the traversed label category, and softmax(·) represents a function that converts any real number vector into a probability distribution. Represents the exponential function of the i-th label category in the vector z, that is, The number of values ​​of the i-th label category in .

[0114] Step S21e: record the most likely result in the adversarial prediction output as the adversarial label, such as Figure 4 shown.

[0115] y adv = argmax(p i )

[0116] Among them, y adv represents the adversarial label, P i represents the predicted probability of the i-th label category, and argmax(·) represents the function of selecting the label category with the largest probability as the adversarial label.

[0117] S22: Count the probability distribution of the label categories of the adversarial labels to obtain the distribution characteristics of the adversarial labels.

[0118] Under the premise that natural samples can be correctly predicted, the probability distribution of the adversarial labels is statistically analyzed based on the CIFAR-10 dataset. At the same time, in order to ensure the universality of the phenomenon, the output results of the two adversarial training methods, TRADES and MART, are statistically observed. The results are as follows: Figure 5 The details are as follows:

[0119] The results of natural training and standard adversarial training are counted using ResNet-18, and the number of samples such as 1st, 2nd, 3rd, ..., 10th are counted. Among them, 1st represents the number of samples with high accuracy and robustness (the first ideal case), and the remaining 2nd, 3rd, ..., 10th represent the ranking of the number of adversarial samples successfully attacked, and also represent the number of samples with high accuracy but low robustness (the second case). In order to enhance the robustness of the model to adversarial samples, the goal of the present invention is to maximize the number of samples of the first category (i.e., the model can correctly handle adversarial samples), and by converting the samples of the second category (the model's incorrect handling of adversarial samples) into the first category, the adversarial defense capability of the model is effectively improved. This conversion can not only improve the robustness of the model, but also enhance the generalization ability of the model in the face of adversarial attacks.

[0120] S3. Adversarial training is conducted based on adversarial label guidance. A regularization term is added to the loss function to make the prediction of adversarial samples closer to the true label, thereby improving robustness accuracy. Constraints are imposed on the current adversarial label to make the prediction of adversarial samples far away from the current adversarial label.

[0121] S31. Input the original sample and the generated adversarial sample into the model and perform forward propagation to obtain the true label and adversarial sample label. The formula is as follows:

[0122] y=f(x,θ)

[0123]

[0124] Where f(·) represents the forward propagation function of the model, usually a neural network, and y represents the true label. represents the adversarial sample label, θ represents the model parameter, x represents the original sample, Represents adversarial examples.

[0125] S32. Design loss function and calculate loss: Use the cross entropy loss function to calculate the difference between the adversarial sample prediction and the true label to make it closer to the distribution of the true label. The formula is:

[0126]

[0127] Among them, L CE (·) is the cross entropy loss function, which means calculating the adversarial sample prediction The loss value between the actual label y; y i represents the actual label probability of the i-th label category, represents the predicted probability of label category i, and C represents the total number of label categories.

[0128] In order to calculate the loss value between the true probability distribution and the predicted probability distribution and make the true probability of the predicted label close to 0 through back propagation, the cross entropy function is modified to obtain the inverse cross entropy loss function, the formula is:

[0129]

[0130] Among them, L ICE (·) is the inverse cross entropy loss function, which means calculating the adversarial sample prediction With adversarial label y adv The loss value between represents the adversarial label of label category i, Represents the predicted probability of label category i, C represents the total number of label categories, and λ is a constant used to prevent the denominator from being too small.

[0131] Cross entropy loss function L CE (·)make Close to 1, but because the inverse of the cross entropy loss function input is taken, the modified inverse cross entropy loss function L ICE (·) can make Far away from 1, that is, close to 0, so that the prediction result is as far away from the current adversarial label as possible

[0132] Since the inverse cross entropy loss function is is close to 0, so when gradient descent is performed multiple times, the gradient converges to 0. CE (·) and L ICE (·) When used together, on the premise that natural samples are correctly classified, the regularization term is designed as follows:

[0133] L ALAT =L CE (p(x adv ,θ),y)+L ICE (p(x adv ,θ),y adv )

[0134] Among them, p(·) represents the predicted value of the model, x adv represents the input variable, θ represents the parameters of the model, y represents the true label, and y adv represents the adversarial label, L ALAT represents the regularization term, L CE (·) represents the cross entropy loss function, L ICE (·) represents the inverse cross entropy loss function.

[0135] For data where natural samples are predicted correctly but adversarial samples are predicted incorrectly, ALAT guides the adversarial prediction label to be closer to the true label, increasing the possibility of the natural prediction label. At the same time, it imposes a penalty on it to keep it away from the current adversarial label, helping to guide the model to conduct more effective adversarial training and guide the adversarial sample prediction errors to correct predictions as much as possible. The total loss function involved in the adversarial training technology based on adversarial labels is as follows:

[0136] L total =L CE (p(x,θ),y)+β·L ALAT

[0137] Among them, L total Represents the total loss function, L CE(·) represents the cross entropy loss function, P(·) represents the predicted value of the model, x represents the input sample, θ represents the parameter of the model, y represents the true label, β represents the adjustable scaling parameter that balances the final loss of the two parts, and L ALAT Represents the regularization term, which is used to calculate the adversarial loss.

[0138] S33, back propagation and parameter update, the gradient of the composite loss function to the model parameter θ is calculated through back propagation, the formula is as follows:

[0139]

[0140] in, represents the gradient of parameter θ, L total Represents the total loss function, L CE(·) represents the cross entropy loss function, P(·) represents the predicted value of the model, x represents the input sample, θ represents the parameter of the model, y represents the true label, β represents the adjustable scaling parameter that balances the final loss of the two parts, and L ALAT Represents the regularization term, which is used to calculate the adversarial loss.

[0141] Use the gradients to update the model parameters:

[0142]

[0143] Among them, θ t+1 represents the updated model parameters, θt represents the model parameters at the tth iteration, η represents the learning rate, represents the gradient of the parameter θ, L total Represents the total loss function.

[0144] In the field of image processing, adversarial training is a commonly used technique that aims to enhance the ability of deep learning models to resist adversarial attacks. This paper deeply understands the nature of adversarial attacks and explores how to improve adversarial training methods to improve the defense capabilities of image classification models in the face of adversarial attacks. By optimizing the adversarial training strategy, the robustness of the model is improved so that it can still maintain high accuracy in an adversarial environment, and the generalization ability of the model can be enhanced to cope with a variety of attack scenarios.

[0145] Experimental verification

[0146] In this paper, an adversarial training defense method based on adversarial label guidance is designed and evaluated on CIFAR-10 and CIFAR-100 datasets.

[0147] Fast Gradient Sign Method (FGSM), Projected Gradient Descent (PGD), Carlini & Wagner Attack (C&W) are selected as baseline attacks, and PGD-10 is used to generate adversarial samples for training. In addition, ALAT is compared with the most comparable and valuable defense methods AT, TRADES, MART, GAIRAT, SAT (AWP-AT) to evaluate the method of the present invention.

[0148] For the CIFAR-10 dataset, three neural networks, ResNet-18, WideResNet-32, and WideResNet-34, were used for training. Set epoch = 130, perturbation ε = 8 / 255, perturbation step size 0.007, iteration number K = 10, learning rate 0.1, batch size 128, and use SGD optimizer (momentum 0.9 and weight decay 2e-4). For the CIFAR-100 dataset, ResNet-18 and WideResNet-32 were used for training, setting epoch = 130, perturbation ε = 8 / 255, perturbation step size 0.007, iteration number K = 10, learning rate 0.1, batch size 128, and use SGD optimizer (momentum 0.9 and weight decay 2e-4). The specific evaluation data is shown in the following table:

[0149] Table 1 Comparison of the robustness of various defense methods after ResNet-18 neural network training on the CIFAR-10 dataset

[0150] method Clean FGSM PGD-20 C&W AT 84.23% 60.77% 43.4% 43.78% TRADES 82.43% 62.73% 48.74% 47.1% MART 81.84% 63.00% 48.59% 45.72% GAIRAT 82.68% 62.11% 49.55% 37.96% SAT-AWP-AT 81.19% 62.19% 49.08% 46.97% The present invention (β=4) 82.32% 63.93% 53.68% 48.5%

[0151] Table 2 Comparison of the robustness of various defense methods after WRN-32 neural network training on the CIFAR-10 dataset

[0152] method Clean FGSM PGD-20 C&W AT 86.64% 63.69% 47.09% 47.62% TRADES 85.18% 63.46% 46.74% 47.10% MART 85.24% 63.69% 46.76% 46.76% GAIRAT 85.03% 64.55% 53.58% 43.61% The present invention (β=4) 85.92% 67.51% 55.19% 49.47%

[0153] Table 3 Comparison of the robustness of various defense methods after WRN-34 neural network training on the CIFAR-10 dataset

[0154] method Clean FGSM PGD-20 C&W AT 81.34% 56.78% 41.86% 42.34% TRADES 85.43% 63.48% 47.04% 47.64% MART 85.39% 63.89% 48.13% 47.29% GAIRAT 84.93% 63.14% 51.05% 43.91% The present invention (β=4) 84.56% 65.59% 53.63% 48.61%

[0155] The experimental results in Table 1-3 prove that the robust accuracy of the present invention is slightly lower than that of other methods only when defending against Clean attacks, and is higher than that of other methods when defending against the other three attacks.

[0156] Table 4 Comparison of the robustness of various defense methods after ResNet-18 neural network training on the CIFAR-100 dataset

[0157]

[0158]

[0159] Table 5 Robustness comparison of various defense methods after WRN-32 neural network training on CIFAR-100 dataset

[0160] method Clean FGSM PGD-20 C&W AT 59.36% 34.19% 23.33% 23.78% TRADES 57.13% 34.47% 25.09% 25.12% MART 58.54% 33.59% 23.37% 23.31% GAIRAT 59.58% 34.34% 24.19% 23.52% The present invention (β=11) 59.73% 36.45% 26.91% 24.31%

[0161] Tables 4-5 show the experimental results on the CIFAR-100 dataset. On this dataset, the method of the present invention still has good robustness. However, in terms of defending against C&W attacks, the present invention is slightly inferior to the TRADES method, but it still shows better performance than AT, MART, GAIRAT and SAT (AWP-AT), exceeding 2.05, 0.31, 2.41 and 1.30 percentage points on the ResNet-18 model, and achieving a robust accuracy of 22.70%. On the WideResNet-32 model, it exceeds AT, MART, and GAIRAT by 0.53, 1.0, and 0.79 percentage points, respectively, and achieves a robust accuracy of 24.31%. In addition, in terms of defending against FGSM and PGD attacks, the robust accuracy of the present invention is the best.

[0162] In addition, due to model capacity limitations or feature generalization and overfitting problems, adversarial training often sacrifices natural accuracy to a certain extent while improving robustness. However, through experiments on two datasets, the present invention achieves a good balance between accuracy and robustness. Compared with other methods, the present invention improves robustness while the natural accuracy rate only slightly decreases or remains comparable, or even shows better performance. For example, based on the CIFAR100 dataset, the natural accuracy rate of ALAT on the WRN-32 model exceeds that of other methods, reaching 59.73%.

[0163] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0164] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A method for adversarial training defense based on adversarial label guidance, characterized in that: The following steps are involved: Step S1, image data preparation: generating adversarial samples; Step S2: extract the features of the adversarial attack, generate adversarial labels, and count the label category probability distribution of the adversarial labels; Step S3: conduct adversarial training based on the adversarial label guidance, add a regularization term to the loss function, and impose constraints on the current adversarial label.

2. According to claim 1, the adversarial training defense method based on adversarial label guidance is characterized in that: The specific process of step S1 is as follows: Step S11: Get the original input sample x from the public dataset cifar10 i and its corresponding true label y i , as the target sample; Step S12: Select a loss function L1 and calculate the loss on the original sample x, specifically: Where L1(·) represents the loss function, x represents the input sample, y represents the true label, C represents the total number of label categories, and y i represents the actual label probability of the i-th label category, Represents the probability of the i-th label category predicted by the model for the input sample x, and i represents the current sequence number of the traversed label category; Step S13: Based on the loss function obtained in step S12, the gradient of the loss function with respect to the current input sample is calculated. The specific formula is as follows: Among them, g represents the gradient of the loss function to the current input sample, represents the gradient of the input sample x, and L1(x,y) is the loss function obtained in step S12; Step S14: Based on the gradient calculated in step S13, generate a preliminary update of the adversarial sample. The specific update formula is: x'=x+a·sign(g) Where x' represents the adversarial sample, x represents the input sample, a represents the step size, which can control the magnitude of each update; sign(g) represents the sign function of the gradient; Step S15: Apply the projection operation to ensure that the generated adversarial sample is within the allowed perturbation range. The specific operation is: x'=clip(x',x-ε,x+ε) Where x' represents the adversarial sample, ε represents the maximum perturbation amplitude allowed, and clip(·) represents the clipping function; Step S16: Repeat steps S12 to S15 for multiple iterations until the predetermined number of iterations is reached or the value of the loss function no longer changes significantly, generating the final adversarial sample.

3. The adversarial training defense method based on adversarial label guidance according to claim 1, characterized in that: The specific process of extracting the features of the adversarial attack and generating the adversarial labels in step S2 is as follows: Step S21a, calculating the input data in sequence through the convolution layer, the fully connected layer and the activation layer to extract features; ReLU(F i l )=max(0,F i l ) Among them, F i l represents the i-th feature map of the l-th layer, is the convolution kernel weight of layer l, represents the j-th feature map of the l-1th layer, represents the bias term, C in Indicates the input channel, is the activated feature map, ReLU(·) represents the activation function; max(0,F i l ) means that when the input sample is greater than 0, the value is directly output, and when the input sample is less than 0, 0 is output; Step S21b: Generate a feature map and flatten the feature map into a one-dimensional vector in, represents the flattened one-dimensional vector, A L represents the output feature map of the last convolutional layer, and flatten(·) represents the flattening operation; Step S21c: vector Perform linear transformation to generate the final feature vector; in, represents the flattened one-dimensional vector, represents the output feature vector, W represents the weight matrix, represents the bias vector; Step S21d: At the output layer, the model converts the feature vector into a prediction result to obtain the probability distribution of the predicted label category; Among them, P i represents the predicted probability of the i-th label category, C is the total number of label categories, z i represents the output value of the i-th label category, i represents the current sequence number of the traversed label category, and softmax(·) represents a function that converts any real number vector into a probability distribution. Represents the exponential function of the i-th label category in the vector z, that is, The number of values ​​of the i-th label category in ; Step S21e, recording the most likely result in the adversarial prediction output as an adversarial label; y adv =argmax(p i ) Among them, y adv represents the adversarial label, P i represents the predicted probability of the i-th label category, and argmax(·) represents the function of selecting the label category with the largest probability as the adversarial label.

4. The adversarial training defense method based on adversarial label guidance according to claim 1, characterized in that: In the step 2, based on the CIFAR-10 data set, the probability distribution of the label categories of the adversarial labels is statistically analyzed, and the results of natural training and standard adversarial training are statistically analyzed using ResNet-18 to calculate the number of samples such as 1st, 2nd, 3rd, ..., and 10th; wherein 1st represents the number of samples with high accuracy and robustness, and the remaining 2nd, 3rd, ..., and 10th represent the ranking of the number of adversarial samples that have been successfully attacked, and also represent the number of samples with high accuracy but low robustness.

5. The adversarial training defense method based on adversarial label guidance according to claim 1, characterized in that: The specific process of step 3 is as follows: S31. Input the original sample and the generated adversarial sample into the model and perform forward propagation to obtain the true label and adversarial sample label. The formula is as follows: y=f(x,θ) Where f(·) represents the forward propagation function of the model, y represents the true label, represents the adversarial sample label, θ represents the model parameters, x represents the original sample, represents adversarial examples; S32. Design a loss function and calculate the loss. Calculate the difference between the adversarial sample prediction and the true label. The formula is: Among them, L CE (·) is the cross entropy loss function, which means calculating the adversarial sample prediction The loss value between the actual label y; y i represents the actual label probability of the i-th label category, represents the predicted probability of label category i, and C represents the total number of label categories; Then the cross entropy function is modified to obtain the inverse cross entropy loss function, the formula is: Among them, L ICE (·) is the inverse cross entropy loss function, which means calculating the adversarial sample prediction With adversarial label y adv The loss value between represents the adversarial label of label category i, represents the predicted probability of label category i, C represents the total number of label categories, and λ is a constant used to prevent the denominator from being too small; When L CE (·) and L ICE (·) When used together, on the premise that natural samples are correctly classified, the regularization term is designed as follows: L ALAT =L CE (p(x adv ,θ),y)+L ICE (p(x adv ,θ),y adv ) Among them, p(·) represents the predicted value of the model, x adv represents the input variable, θ represents the parameters of the model, y represents the true label, and y adv represents the adversarial label, L ALAT represents the regularization term, L CE (·) represents the cross entropy loss function, L ICE (·) represents the inverse cross entropy loss function; The total loss function involved in the adversarial training technique based on adversarial labels is as follows: L total =L CE (p(x,θ),y)+β·L ALAT Among them, L total Represents the total loss function, L CE(·) represents the cross entropy loss function, P(·) represents the predicted value of the model, x represents the input sample, θ represents the parameter of the model, y represents the true label, β represents the adjustable scaling parameter that balances the final loss of the two parts, and L ALAT represents the regularization term; S33, back propagation and parameter update, the gradient of the composite loss function to the model parameter θ is calculated through back propagation, the formula is as follows: in, represents the gradient of parameter θ, L total Represents the total loss function, L CE(·) represents the cross entropy loss function, P(·) represents the predicted value of the model, x represents the input sample, θ represents the parameter of the model, y represents the true label, β represents the adjustable scaling parameter that balances the final loss of the two parts, and L ALAT represents the regularization term; Use the gradients to update the model parameters: Among them, θ t+1 represents the updated model parameters, θ t represents the model parameters at the tth iteration, η represents the learning rate, represents the gradient of the parameter θ, L total Represents the total loss function.