An Image Classification Method Based on a Fair and Robust Neural Network
The method addresses the class-specific robustness variability in neural networks by dynamically adjusting perturbation intensity and using parameter averaging, enhancing the robustness and fairness of image classification models, particularly in safety-critical domains like autonomous driving.
Patent Information
- Application Number
- CN202310159248.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-23
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-02-23
AI Technical Summary
Existing neural network-based machine learning models in image classification tasks suffer from significant variability in robustness across different classes, with some classes being more susceptible to adversarial attacks, posing a safety risk in critical applications like autonomous driving.
A method for fair and robust neural network training that adjusts the perturbation intensity and uses parameter averaging to enhance the robustness of the weakest class, ensuring high accuracy across all classes by dynamically adapting the perturbation radius based on class-specific accuracy and maintaining a minimum robustness threshold.
The proposed method significantly improves the robustness of the weakest class and enhances overall class fairness, leading to higher reliability and safety in critical applications.
Smart Images

Figure CN116091838B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine learning in artificial intelligence, and relates to the adversarial robustness and security of machine learning image processing. Specifically, it is an image classification method based on neural network adversarial training, which has the characteristics of high robustness and fairness between classes. Background Art
[0002] Machine learning methods based on neural networks have achieved great advantages in image classification tasks and are widely used in fields such as autonomous driving and intelligent diagnosis and treatment. By taking a number of samples with class labels as the training set and inputting them into the neural network for training, satisfactory classification results can be obtained in the test set. However, studies have found that neural network models trained by conventional machine learning generally have the problem of adversarial samples. Adversarial samples are a type of samples that add tiny perturbations to the original input samples, resulting in misclassification of the model. For a sample x belonging to class y and a classification prediction model f θ , if the model can correctly classify x without perturbation, that is, argmax k f θ (x) k = y, but misclassification occurs after adding a tiny perturbation δ, that is, argmax k f θ (x + δ) k ≠ y, then x + δ is called an adversarial sample, where f θ (x) k is the prediction probability of the model f θ for the sample x in the k-th class.
[0003] The discovery of adversarial samples reveals great hidden dangers in the security of artificial intelligence, enabling attackers to achieve adversarial attacks by adding perturbations to the input samples, thereby interfering with the performance of machine learning models. For example, by adding specific adversarial interference patterns to traffic road signs, attackers can cause the road sign classifier carried by autonomous vehicles to misjudge the road sign category during driving, thus posing a threat to traffic safety. In the field of intelligent military industry, attackers can also graffiti specific adversarial samples on equipment to bypass intelligent models performing reconnaissance tasks. Therefore, in some artificial intelligence applications in security-critical fields, it is necessary to design defense methods against adversarial samples to avoid being damaged by attackers adding adversarial samples to the model performance. These classification methods with adversarial defense technologies are collectively referred to as robust image classification methods.
[0004] Adversarial training technology is a current mainstream method for improving the robustness of image classification. Adversarial training improves the robustness of the model by adding adversarial samples to the machine learning training process, enabling it to improve the accuracy (i.e., robustness) under adversarial perturbations. In existing adversarial training methods, for a K-classification task, let D be the collected dataset, and (x, y) represent the samples and class labels sampled from the dataset respectively. Let θ represent the parameters of the fitted objective function f θ where f θ (x) ∈ [0, 1] k gives the predicted probability for each class, and L(θ; x, y) represents the loss function of the objective function f θ on the sample x and its corresponding label y. The goal of adversarial training is to use the samples with the highest loss function within the perturbation range for training. Generally, the maximum perturbation radius ∈ is preset in advance, and the perturbation range Δ = {δ: ||δ|| p ≤ ∈} is selected according to the l p norm. After obtaining a set of samples and their labels (x, y) by each sampling, find the perturbation δ that maximizes L(θ; x + δ, y) within the perturbation range Δ, and use x + δ to replace x for training. In summary, the general adversarial training method can be expressed as the following optimization problem: θ = argmin θ E (x,y)~D max δ∈Δ L(θ; x + δ, y). The model f θ obtained after adversarial training has better robustness than the model obtained by ordinary training and shows higher accuracy when predicting samples x + δ with adversarial interference.
[0005] However, recent research has shown that after adversarial training, the robustness of the model in classification tasks varies significantly across different classes. Specifically, the model has good robustness when processing samples of some classes, but poor robustness for some classes. This poses a further challenge to the security of artificial intelligence. For example, in the road sign classification model for autonomous driving, assuming that the model has good average robustness for hundreds of road signs, it may still have poor robustness in classifying some types of road signs (such as speed limit signs and no-entry signs). If an attacker deliberately adds adversarial interference to these road signs, it will pose a huge safety hazard to the moving vehicle. So far, how to improve the robust fairness of adversarial training through adversarial training, that is, to have better robustness for each class, remains an unsolved problem. Summary of the Invention
[0006] To overcome the significant differences in inter-class robustness brought about by existing adversarial training methods, the present invention proposes an image classification method based on a fair robust neural network, which mainly includes two parts: perturbation intensity calibration and parameter averaging calibration based on improving robust fairness. Thus, in the deployed neural network model, it has higher robustness to the worst-case classes, thereby improving the credibility and security of the classification task.
[0007] The technical solution provided by the present invention is as follows:
[0008] An image classification method based on a fair robust neural network, comprising the following steps:
[0009] A. Collect the dataset D for the classification task. Suppose there are K classes in the classification task, then the same number of samples need to be collected from each class y ∈ {1, 2,..., K} to form several sample-label pairs (x, y), which are included in the dataset D. Then divide the training set D into the training set D train and the validation set D valid . Next, initialize a neural network f θ (such as the common ResNet-18 model), where θ is the network parameter. The preset perturbation radius is ∈, generally selected as 8 / 255. Then randomly initialize a set of neural network parameters whose network structure is the same as that of f θ for maintaining the model parameter averaging. Set the robust fairness threshold γ, generally selectable as 0.2. Suppose the model is trained for N rounds in total.
[0010] B. In the T-th round of training, where T ∈ {1, 2,..., N}, use the gradient descent method to train the neural network. Each round of training includes the following steps in sequence:
[0011] B1. For k ∈ {1, 2,..., K}, assume that the training accuracy of the k-th class samples in the previous round of training iteration is t k , then in this round of training iteration, for this class, take the perturbation radius of ∈ k ← (λ1 + t k ) · ∈ (if this is the first round, then ∈ k is set to ∈). Where λ1 is the set hyperparameter to avoid too small a perturbation radius, generally selectable as 0.5. The purpose is that the training accuracy t k of the more difficult classes is smaller, so that the perturbation radius ∈ k used in this round is correspondingly reduced to reduce the perturbation intensity on such samples during the adversarial training process; vice versa. Thus, the proposed classification calibration method can automatically adapt to the appropriate training perturbation intensity during the training process.
[0012] B2. According to the perturbation radius ∈ of each class obtained in step B.1k , perform adversarial training with stochastic gradient descent on the training set D train . For each mini-batch sample of {(x n , y n ),} within the perturbation range , find the adversarial sample for the current model f θ with respect to x n , that is, solve the optimization problem max δ∈Δ L(θ; x n + δ, y n ), and then perform gradient descent optimization on the parameter θ using x n + δ as the input: where η is the learning rate hyperparameter.
[0013] B3. For the parameter θ obtained after each iteration, when the robustness of the model f θ for each class in the validation set D valid is not less than γ, that is R k (θ; D valid ) ≥ γ, incorporate the θ obtained in this round into the parameter averaging: Otherwise, discard the parameter θ obtained in this round. Where R k (θ; D valid ) represents the robust accuracy of the samples of class k in the validation set D θ for the model f with parameter θ valid .
[0014] C. When the N rounds of training in step B are completed, return the model after parameter averaging At this time, use as the output after training is completed, and encapsulate into an API interface for reading images to complete the classification task.
[0015] Compared with the prior art, the beneficial effects of the present invention:
[0016] The image classification method based on a fair and robust neural network proposed by the present invention can effectively improve the robustness of the machine learning model, making the robustness of the classification model in the worst class significantly better than the existing robust classification methods, thus having higher credibility and security in the image classification tasks in safety-critical fields. Taking the road sign classification task in autonomous vehicle driving as an example, after adversarial training using the present invention, the classifier can greatly improve the robustness of road signs in difficult-to-classify categories, thereby enhancing the safety of autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Flow chart of the embodiment of the present invention. DETAILED DESCRIPTION
[0018] Taking the robust classification task of the CIFAR-10 dataset as an example to describe the present invention, but not limiting the scope of the present invention. CIFAR-10 is a classic dataset, which contains pictures of 10 categories including animals and vehicles such as airplanes, cars, cats, and deer. Its training set contains a total of 50,000 pictures, with 5,000 pictures in each category. In this image classification task, the specific implementation is as follows (the entire flowchart is as Figure 1 shown):
[0019] A. Initialization of model parameters and selection of hyperparameters. Select the ResNet-18 neural network structure, and at the same time initialize two sets of model parameters θ, The perturbation radius is selected as ∈ = 8 / 255, and the robust fairness threshold γ = 0.2. Divide the dataset D into a training set D train and a validation set D valid , where each category in the training set contains 4,900 pictures, and each category in the validation set contains 100 pictures. The number of training epochs is set to 200, the learning rate η is initially set to 0.1, and is set to 0.01 and 0.001 at the 100th and 150th epochs respectively. The hyperparameter λ1 is set to 0.5, and α is set to 0.85. The loss function L(θ; x, y) is selected as the cross-entropy loss: L(θ; x, y) = -logf θ (x) y .
[0020] B. In the T-th (T ∈ {1, 2,..., 200}) epoch of training, the gradient descent method is adopted for the training of the neural network. During the training process, the pixel values of the pictures to be read are divided by 255 and converted to the range [0, 1].
[0021] Each epoch of training includes the following steps in sequence:
[0022] B1. For k ∈ {1, 2,..., K}, if this is the first epoch, then ∈ k is set to ∈. Otherwise, let the training accuracy of such samples in the previous training iteration of the model be t k , and take ∈ k ← (λ1 + t k ) · ∈.
[0023] B2. According to the perturbation radius ∈ k of each category obtained in step B.1, perform adversarial training on the training set D train . Each time, randomly select 128 sample-label pairs from the samples in the training set that have not been selected For each sample (x n , y n ), search within the perturbation range for the current model fθ For adversarial examples of x, that is, solve the optimization problem max 8∈Δ L(θ; x n + δ, y n ). This optimization problem can be solved by the ten-step Projected gradient descent (PGD) method: First, randomly initialize δ within Δ 0 , and then for t = 1, 2,..., 10, solve successively The finally obtained δ n = δ 10 is the adversarial perturbation for x n . Finally, use gradient descent to optimize the model parameters on this batch of adversarial examples:
[0024] B3. For the parameters θ obtained after each round of iteration, test the robustness of the model on the validation set D vafid . For the 10 sub-datasets composed of 10 classes of samples in the validation set Verify the robustness of f θ on these sub-datasets in turn, and still use ten-step PGD to generate adversarial examples. If the robustness of the model f θ on each class is not less than γ = 0.2, that is, it can correctly classify 20% of the adversarial examples of each category, then perform model parameter averaging: Otherwise, directly enter the next round of training iteration.
[0025] C. When the 200 rounds of training in step B end, return the model after parameter averaging At this time can be used as the interface of the classification model to implement the classification task.
[0026] After evaluation, the image classification method of the present invention can significantly improve the robustness on the category with the lowest robustness (usually the category labeled 'cat' in CIFAR-10). Taking 8 / 255 under the l ∞ norm of the perturbation range as an example, compared with the classical adversarial training method, the present invention can improve the robustness of the worst category from 28% to 35%, and the average robustness from 52 to 56%, which significantly improves the robustness of this classification task and the fairness between its categories.
[0027] It should be noted that the purpose of publishing this embodiment is to help further understand the present invention, but those skilled in the art can understand that: within the scope not departing from the present invention and the appended claims, various substitutions and modifications are possible. Therefore, the present invention should not be limited to the content disclosed in the embodiment, and the scope claimed by the present invention is defined by the scope of the claims.
Claims
1. An image classification method based on a fair and robust neural network, comprising the following steps: 1) Collect the dataset for the classification task , assuming that there are categories in the classification task, then the same number of samples need to be collected from each category to form a number of sample-label pairs , and include them in the dataset ; 2) Divide the training set into a training set and a validation set , initialize a neural network , where are network parameters, randomly initialize a set of neural network parameters , whose network structure is the same as , used to maintain the average of model parameters, set the robust fairness threshold , assume that the model is trained rounds; 3) In the round of training, each round of training sequentially includes the following steps: 3-1) For , assume that the training accuracy of the -th class of samples in the previous training iteration of the model is . Then, in this training iteration, adopt a perturbation radius of for this class, where is a set hyperparameter, and the preset perturbation radius is . The training accuracy of the more difficult classes is smaller, so that the perturbation radius used in this round is correspondingly reduced to reduce the perturbation intensity on such samples during adversarial training; vice versa. The perturbation radius of each category obtained according to step 3-1) , perform stochastic gradient descent adversarial training on the training set , for each mini-batch of samples , within the perturbation range , search for adversarial samples with respect to the current model For , that is, solve the optimization problem ; 3-3) For the parameters obtained after each round of iteration When the model on the validation set the robustness of each class is not less than that is include the obtained in this round into the parameter averaging: otherwise discard the parameters obtained in this round where represents the model with parameters the robust accuracy of the samples of class in the validation set ; sample 4) After the round of training in step 3) ends, return the model with averaged parameters to complete the classification task.
2. The image classification method based on a fair and robust neural network according to claim 1, wherein Initialize the neural network parameters in step 2): The perturbation radius is selected as , the robust fairness threshold , and the hyperparameter is set to 0.
5. The loss function is selected as the cross-entropy loss: .
3. The image classification method based on a fair and robust neural network according to claim 1, characterized in that, The ResNet-18 model is adopted in step 2).
4. The image classification method based on a fair and robust neural network according to claim 1, wherein The gradient descent method is adopted in step 3) for the training of the neural network.
5. The image classification method based on a fair and robust neural network according to claim 4, characterized in that, Step 3-2) Use the ten-step PGD method to solve: First, Internal random initialization , and then solve The final result That is for Finally, the model parameters are optimized using gradient descent on this batch of adversarial samples: 。
Citation Information
Patent Citations
Robust neural network training method based on sample-driven target loss function optimization
CN115438786A