A method for securing a deep learning model based on anomaly deviation neurons

By identifying and correcting abnormal bias-like neurons in deep learning models, generating abnormal samples and optimizing the model, the defense problems under poisoning attacks are solved, and the robustness and defense effect of the model are improved.

CN115203690BActive Publication Date: 2025-07-18ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210932483.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-04
Publication Date
2025-07-18
Estimated Expiration
2042-08-04

AI Technical Summary

Technical Problem

Existing deep learning models are difficult to defend against under poisoning attacks, and defense methods often affect the accuracy of the model under clean samples, lacking effective robust enhancement methods.

Method used

By recording and accumulating neuron activation values, abnormal biased neurons are identified, and abnormal samples are generated using gradient backpropagation. The model is fine-tuned to correct abnormal biased neurons, and the loss function optimization model is constructed to defend against poisoning attacks.

Benefits of technology

Effectively defend against poisoning attacks, without affecting the accuracy of the model in clean samples, and significantly improving the robustness of the model, making it more difficult for the model to be attacked successfully.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115203690B_ABST
    Figure CN115203690B_ABST
Patent Text Reader

Abstract

The present invention discloses a deep learning model security reinforcement method based on abnormal deviation class neurons, the method first uses a poisoned data set to train a poisoned model, then a small amount of clean data is input into the poisoned model, and neurons with activation values greater than or equal to the average value are recorded as a neuron group; a large amount of clean data is then input into the poisoned model, the activation values of neurons in the neuron group are calculated, and the activation values of these neurons for all input clean data are accumulated and sorted by size, and the neurons with the highest sorting are selected, which are abnormal deviation class neurons; then a loss function is constructed, and clean data is optimized to abnormal data through back propagation, and finally the poisoned model is trained with a data set composed of abnormal data and clean data to complete the reinforcement of the model. The present invention can effectively defend against poisoning attacks, and does not affect the accuracy of normal samples of the model, greatly improves the robustness of the model, and makes the model more difficult to be successfully attacked.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of security issues for deep learning models, and particularly to a method for securing deep learning models based on abnormal deviation neurons. Background Art

[0002] Deep learning (DL) is a new research direction in the field of machine learning (ML). It has been introduced into machine learning to make it closer to the original goal - artificial intelligence. With the rapid development and application of deep learning and the continuous development of artificial intelligence technology, the research results of deep learning have been widely applied in the fields of natural language processing, image recognition, industrial control, signal processing, security, etc. Among them, security applications are particularly important. If there are vulnerabilities in the data or algorithms in security fields such as autonomous driving, military operations, and public opinion warfare, it will cause significant personal injuries and property losses. For example, in 2018 alone, there were 12 autonomous driving accidents globally, involving AI giants in autonomous driving research and development such as Uber, Tesla, Ford, and Google. Therefore, it is crucial to study attacks on deep learning models, discover vulnerabilities in the models, and defend against them.

[0003] The development of deep learning and the enhancement of high-performance GPU processing capabilities have made neural network structures more and more complex, and the number of model parameters has also become increasingly large. The security of deep learning has encountered great difficulties and challenges. Currently, attacks on deep learning models are mainly divided into poisoning attacks and adversarial attacks. Poisoning attacks occur during the model training stage. The attacker injects poisoned samples into the training data set, thereby embedding a backdoor trigger in the trained deep learning model. When a poisoned sample is input during the test stage, the attack breaks out. Adversarial attacks occur during the model test stage. The attacker obtains adversarial samples by adding carefully designed tiny perturbations to the original data, thereby fooling the deep learning model and making it misjudge with a high confidence. Among them, there are particularly many poisoning attacks on deep learning models. Existing poisoning defense methods for deep learning models focus on eliminating the toxicity of the model and enhancing the robustness of the model. They ignore the time spent in this process and the accuracy of the model on clean samples. This leads to a problem: how to defend against many poisoning attacks on the model without affecting or even optimizing the accuracy of the original model on clean samples. Summary of the Invention

[0004] In view of the shortcomings of the prior art, the present invention provides a deep learning model security reinforcement method based on abnormal deviation neurons, which uses the difference in activated neurons when samples propagate forward in the model, accumulates the neuron activation values of the model, sorts them from large to small according to the activation values, takes the top 2%-3%, and uses this neuron as a guide to generate abnormal samples according to gradient back propagation, and fine-tune the model according to the abnormal samples and clean samples. This improves the robustness of the model and achieves the effect of model poisoning defense.

[0005] The purpose of the present invention is achieved through the following technical solutions:

[0006] A deep learning model security reinforcement method based on abnormal deviation neurons, the method comprising:

[0007] Step 1: Prepare a clean dataset;

[0008] Step 2: Select a deep learning network, use the poisoning attack method to generate poisoned data, add the poisoned data to the clean data set to form a poisoned data set; use the poisoned data set to train the deep learning network to obtain a poisoned model;

[0009] Step 3: Input a small amount of clean data into the poisoning model, calculate the average activation value of the neurons in the fully connected layer of the poisoning model, and record the neurons whose activation values are greater than or equal to the average value as the neuron group; input a large amount of clean data into the poisoning model, calculate the activation values of the neurons in the neuron group, and accumulate the activation values of these neurons for all the input clean data, and sort the accumulated activation values from large to small, and select the neurons with the highest ranking, which are the abnormal deviation neurons;

[0010] Step 4: Construct a loss function, input clean data into the poisoning model, and record the activation values of the abnormal deviation class neurons in the fully connected layer of the poisoning model, and use this to calculate the loss function value; use the gradient information of the loss function to change the pixel value of the input clean data and iteratively optimize it into abnormal data;

[0011] Step 5: Use the abnormal data obtained in step 4 and the original clean data to form a data set and train the poisoning model. The model corrects the abnormal deviation neurons into normal working neurons to defend against poisoning attacks, thereby completing the reinforcement of the model.

[0012] Furthermore, the loss function is specifically:

[0013]

[0014] TK_FC(X)=∑max k (M fc (x i))

[0015] Among them, λ represents the balance parameter, which can be artificially adjusted and defaults to the constant 1; max k (·) represents the k outputs with the largest activation values in this layer. k is a hyperparameter and should not exceed the number of neurons in the anomaly deviation class. n is the number of samples; x is the dimension of the prediction vector; y is the true value after one-hot encoding, corresponding to the label on the x dimension, taking values of 1 or 0; a is the predicted label in one-hot format, taking values from 0 to 1.

[0016] Furthermore, the clean data set in step one is selected from any one of the image data sets of MNIST, CIFAR10, ImageNet, GTSRB, and CASIA.

[0017] Furthermore, the deep learning network in step one is selected from any one of LeNet, AlexNet, VGG11, and ResNet34.

[0018] Furthermore, the poisoning attack method is selected from any one of BadNets, PoisonFrog, Trojannn, and FeatureCollision Attack.

[0019] The beneficial effects of the present invention are as follows:

[0020] The method of the present invention can effectively defend against poisoning attacks, does not affect the accuracy rate of normal samples of the model, and at the same time greatly improves the robustness of the model, making the model more difficult to be successfully attacked. Description of the Drawings

[0021] Figure 1 It is a block diagram of one embodiment of the security reinforcement method for a deep learning model based on neurons in the anomaly deviation class.

[0022] Figure 2 It is the network structure diagram of LeNet in the embodiment. Detailed Embodiments

[0023] The present invention will be described in detail below according to the drawings and preferred embodiments. The purpose and effects of the present invention will become more apparent. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0024] Currently, the defense methods against poisoning attacks can be classified according to the action stage into: data and feature modification, model modification, and output defense. Data and feature modification mainly refers to modifying the model parameters during data defense to achieve the defense effect. Output defense refers to analyzing the output results of the model to achieve the defense effect.

[0025] Therefore, when obtaining a model, specific neuron fine-tuning is required before deploying it in a practical scenario to ensure the security of the model. Experiments have found that the vast majority of poisoning attacks will cause the weights of some neurons to increase. Therefore, whenever a sample is input into the model, the activation values of these neurons will be abnormally active (hereinafter referred to as abnormally deviated neurons for short). Numerically, it is abnormally high, resulting in a deviation in the prediction behavior of the model until misclassification occurs. After deleting these abnormally deviated neurons, it is found that the poisoning success rate of the model drops rapidly. To ensure that the neurons found are all abnormally deviated neurons, after each sample is input into the model, the activation value of each neuron is recorded, and these activation values are accumulated. After a certain number of samples are input, the activation values of each neuron are sorted from large to small. The top 1% in the sequence is the abnormally deviated neuron found. By accumulating the neuron activation values of each sample, it helps to reduce errors and ensure that the neurons found are all abnormally deviated neurons. When these abnormally deviated neurons are found, the pixel values of the original image can be changed by guiding the gradient ascent of the loss function with these abnormally deviated neurons as the guide. The original image is converted into an abnormal image, and the abnormal image will cause the model to make a wrong prediction due to the huge activation value of the abnormally deviated neurons. These samples, with normal labels, are used to input into the model, enabling the model to learn the knowledge and differences between these samples. The abnormally deviated neurons in the model are corrected to normally working neurons, thus defending against poisoning attacks and increasing the robustness of the model.

[0026] Specifically, for a model, take a clean sample and input it into the model. In the fully connected layer of the model, record its neuron activation value and accumulate it with the neuron activation value of the next clean sample. After a certain number of clean samples, sort the total activation values and select the top-k neurons. These neurons are the abnormally deviated neurons found. After finding the neurons, construct a loss function, and update the sample X by changing the pixel values of the original image by guiding the gradient ascent of the loss function. Then perform the next round of iteration until the activation value of the abnormally deviated neurons reaches a certain level, causing the model to make a wrong prediction and the iteration ends. At this moment, this sample becomes an abnormal sample. Use the generated abnormal samples and clean samples to input into the model, let the model learn the differences between them, and slowly correct the abnormally deviated neurons in the model to normal neurons, then the security reinforcement can be completed.

[0027] Refer to Figures 1 to 2 , the deep learning model security reinforcement method based on abnormally deviated neurons includes the following steps:

[0028] (1) Prepare a clean data set:

[0029] (1.1) Select image datasets such as MNIST, CIFAR10, ImageNet, GTSRB, CASIA, etc. Take the MNIST dataset as an example. This is a grayscale image dataset with 10 categories and an image size of 28*28.

[0030] (2) Preparation of poisoned model: Select a deep learning network, use the poisoning attack method to generate poisoned data, add the poisoned data to the clean data set to form a poisoned data set; use the poisoned data set to train the deep learning network to obtain the poisoned model;

[0031] (2.1) Poisoning method: Select poisoning attack methods such as BadNets, PoisonFrog, Trojannn, Feature CollisionAttack, etc. We take BadNets as an example.

[0032] (2.2) Deep learning network: Select LeNet, AlexNet, VGG11, ResNet34 and other networks. This embodiment takes LeNet network as an example.

[0033] (2.3) Model poisoning operation: Taking BadNets poisoning on the MNIST dataset as an example, the accuracy of the trained poisoned model on the test set is 98.37% for clean samples and 100% for poisoned samples.

[0034] (3) Input a small amount of clean data into the poisoning model, calculate the average activation value of the neurons in the fully connected layer of the poisoning model, and record the neurons whose activation values are greater than or equal to the average value as the neuron group; Input a large amount of clean data into the poisoning model, calculate the activation values of the neurons in the neuron group, and accumulate the activation values of these neurons for all the input clean data, and sort the accumulated activation values from large to small, and select the neurons with the highest ranking, which are the abnormal deviation neurons;

[0035] (3.1) Construct the required data set: Select a certain number of samples from each category of the test sample to form a sample set X. The larger the sample set and the more complete the types, the better the effect. In terms of quantity, the selected sample set must account for at least 1% of the test set. In terms of types, at least half of the types must be included.

[0036] (3.2) Select the neurons in the fully connected layer to record: In the fully connected layer of the model, there are too many neurons. In order to save the time cost of the algorithm, it is necessary to select some neurons and record their activation values. According to the sum of the activation values of the neurons in the entire fully connected layer, the average value is calculated. Using this average value as the threshold, only neurons with activation values greater than or equal to the average value are recorded.

[0037] (3.3) Select neurons of the abnormal deviation type: Input these samples into the poisoned model and focus on observing the fully connected layer of the poisoned model. Whenever a sample is input into the model, we need to record the activation values in the fully connected layer. Repeat recording multiple samples and accumulate the data they record. After accumulating a certain number of samples, sort the data from largest to smallest, and select the top 2%-3% of this neuron group as the selected neurons of the abnormal deviation type S K . In actual use, according to the different models, fine-tune the interval, for example, adjust it to the top 2%-5%, etc.

[0038] (4) Construct the loss function, input the clean data into the poisoned model, record the activation values of the neurons of the abnormal deviation type in the fully connected layer of the poisoned model, and calculate the loss function value based on this; use the gradient information of the loss function to change the pixel values of the input clean data and iteratively optimize it into abnormal data;

[0039] (4.1) Construct the loss function: Input the test image dataset X into the trained model M and record the activation values M of the neurons of the abnormal deviation type in the fully connected layer of the poisoned model M fc (x i ), where x i ∈X, i = 1, 2,.... Accumulate the top-k neurons with the largest activation values in the output to form the loss function:

[0040] TK_FC(X) = ∑max k (M fc (x i ))

[0041]

[0042] Among them, λ represents the balance parameter, which can be adjusted manually and defaults to the constant 1; max k (·) represents the k outputs with the largest activation values in this layer, k is a hyperparameter and should not exceed the number of neurons of the abnormal deviation type, and n is the number of samples.

[0043] x is the dimension of the prediction vector because it needs to be calculated and summed one by one on the dimension of the output feature vector.

[0044] y is the label corresponding to the true value after one-hot encoding on the dimension of x, taking values of 1 or 0.

[0045] a is the predicted label in one-hot format, taking values from 0 to 1.

[0046] (4.2) Update the input samples: Use the gradient information of the loss function to change the pixel values of the original image and iteratively optimize it into abnormal samples.

[0047] (5) Use the abnormal data obtained in step (4) and the original clean data to form a data set, and train a poisoned model. The model corrects the abnormal deviation type neurons among them into normally working neurons to defend against poisoned attacks, thereby completing the reinforcement of the model.

[0048] Those of ordinary skill in the art can understand that the above are only preferred examples of the invention and are not used to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, for those skilled in the art, they can still modify the technical solutions described in the foregoing examples or equivalently replace some of the technical features. Any modifications, equivalent replacements, etc. made within the spirit and principle of the invention shall be included in the protection scope of the invention.

Claims

1. A method for securing a deep learning model based on anomaly deviation neurons, characterized in that, The method includes: Step 1: Prepare a clean dataset; Step 2: Select a deep learning network, use the poisoning attack method to generate poisoned data, add the poisoned data to the clean data set to form a poisoned data set; use the poisoned data set to train the deep learning network to obtain a poisoned model; Step 3: Input a small amount of clean data into the poisoning model, calculate the average activation value of the neurons in the fully connected layer of the poisoning model, and record the neurons whose activation values are greater than or equal to the average value as the neuron group; input a large amount of clean data into the poisoning model, calculate the activation values of the neurons in the neuron group, and accumulate the activation values of these neurons for all the input clean data, and sort the accumulated activation values from large to small, and select the neurons with the highest ranking, which are the abnormal deviation neurons; Step 4: Construct a loss function, input clean data into the poisoning model, and record the activation values of the abnormal deviation class neurons in the fully connected layer of the poisoning model, and use this to calculate the loss function value; use the gradient information of the loss function to change the pixel value of the input clean data and iteratively optimize it into abnormal data; Step 5: Use the abnormal data obtained in step 4 and the original clean data to form a data set and train the poisoning model. The model corrects the abnormal deviation neurons into normal working neurons to defend against poisoning attacks, thereby completing the reinforcement of the model.

2. The method for securing a deep learning model based on anomaly deviation type neurons according to claim 1, wherein The loss function is specifically: TK_FC(X) = ∑max k (M fc (x i )) Among them, λ represents the balance parameter, which can be artificially adjusted and defaults to the constant 1; max k (·) represents the k outputs with the largest activation values in this layer. k is a hyperparameter and should not exceed the number of anomaly deviation neurons. n is the number of samples; x is the dimension of the prediction vector; y is the true value after one-hot encoding, corresponding to the label on the x dimension, taking values of 1 or 0; a is the predicted label in one-hot format, taking values from 0 to 1.

3. The method for securing a deep learning model based on anomaly deviation type neurons according to claim 1, characterized in that The clean data set in step 1 is selected from any one of MNIST, CIFAR10, ImageNet, GTSRB and CASIA image data sets.

4. The method for securing a deep learning model based on anomaly deviation neurons according to claim 1, wherein The deep learning network in step 1 is selected from any one of LeNet, AlexNet, VGG11, and ResNet34.

5. The method for securely strengthening a deep learning model based on an abnormal deviation type neuron according to claim 1, wherein The poisoning attack method is any one of BadNets, PoisonFrog, Trojannn, and Feature CollisionAttack.