Deep Model Poisoning Defense Method and Device for Image Classification Based on Neuron Reinforcement
By identifying and strengthening high-activated neurons in the image classification depth model, and fine-tuning the neuron weights using gradient backpropagation, the problem that the model poisoning prevention method in the prior art failed to enhance robustness, and the poisoning prevention effect was achieved without affecting the normal performance of the model.
Patent Information
- Application Number
- CN202211241175.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-11
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-10-11
AI Technical Summary
The existing deep learning model poisoning prevention methods fail to effectively enhance the robustness of the model during the poisoning process, and ignore the accuracy of clean samples, resulting in affecting the normal performance of the model when defending against poisoning attacks.
By identifying and strengthening highly activated neurons in the image classification depth model, using gradient backpropagation to fine-tune the neuron weights to form a clean image classification depth model to defend against poisoning attacks.
Without affecting the accuracy of the model's clean sample, it effectively defends against poisoning attacks, improves the robustness of the model, and is suitable for different data sets and models.
Smart Images

Figure CN115688866B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of security issues of deep models for image classification, and particularly relates to a method and device for poisoning defense of a deep image classification model based on neuron reinforcement. Background Art
[0002] Deep reinforcement learning is one of the directions that has attracted much attention in artificial intelligence in recent years. With the rapid development and application of deep learning and the continuous development of artificial intelligence technology, the research results of deep learning have been widely applied in the fields of natural language processing, image recognition, industrial control, signal processing, security, etc. Among them, security applications are particularly important. If there are vulnerabilities in the data or algorithms in security fields such as autonomous driving, military operations, and public opinion warfare, it will cause significant personal injuries and property losses. For example, in 2018 alone, there were 12 autonomous driving accidents globally, including AI giants in autonomous driving research and development such as Uber, Tesla, Ford, and Google. Therefore, it is crucial to study attacks on deep learning models, discover the vulnerabilities in the models, and conduct defenses.
[0003] The development of deep learning and the enhancement of the processing power of high-performance GPUs have made the neural network structure more and more complex, and the number of model parameters has also become more and more huge. The security of deep learning has encountered great difficulties and challenges. Currently, the attacks on deep learning models are mainly divided into poisoning attacks and adversarial attacks. Poisoning attacks occur in the model training stage. The attacker injects poisoned samples into the training data set, thereby embedding a backdoor trigger in the trained deep learning model. When a poisoned sample is input in the test stage, the attack breaks out. Adversarial attacks occur in the model test stage. The attacker obtains adversarial samples by adding carefully designed small perturbations to the original data, thereby fooling the deep learning model and making it misjudge with a high confidence. Among them, there are particularly many poisoning attacks on deep learning models. Existing poisoning defense methods for deep learning models focus on eliminating the toxicity of the model and enhancing the robustness of the model. They ignore the time cost in this process and the accuracy of the model on clean samples. This leads to a problem: how to defend against many poisoning attacks on the model without affecting or even optimizing the accuracy of the original model on clean samples. Summary of the Invention
[0004] Currently, the poisoning and defense of deep image classification models are in a continuous game situation. Based on the fact that existing poisoning defense methods do not fully consider the training time cost and the robustness of the model, the present invention proposes a method for poisoning defense of a deep image classification model based on neuron reinforcement. The present invention uses the difference in the activated neurons when the sample propagates forward in the model to first find the neurons that need to be reinforced, and then fine-tunes the reinforced neurons according to the gradient backpropagation, thereby improving the robustness of the model against poisoning attacks and achieving the effect of model poisoning defense.
[0005] The object of the present invention is achieved by the following technical solutions:
[0006] According to the first aspect of this specification, a poisoning defense method for an image classification deep model based on neuron reinforcement is provided. The method includes the following steps:
[0007] S1. Prepare an image dataset, select a deep learning network, use a poisoning attack method to generate poisoned samples and add them to the training set, and train to obtain a poisoned model;
[0008] S2. Find the neurons that need to be reinforced, including the following sub-steps:
[0009] S2.1. Use the clean samples in the training set as test samples, select a certain test sample of one class and input it into the poisoned model, count the Top-K neurons of the activation values of the fully connected layer of the model, and at the same time record the K neurons with the highest frequency among the Top-K neurons corresponding to the test samples sampled on the fully connected layer as the main neurons of this class. Finally, obtain the main neurons of each class;
[0010] S2.2. Statistically find the common part among the main neurons of each class to form a neuron set N;
[0011] S2.3. Sort the neurons in the neuron set N according to the sum of the number of times they appear among the Top-K neurons of each class of test samples, and define the set Select neurons from the neuron set N and add them to the set M one by one from front to back. After each addition, reduce the weights of all neurons in the set M to less than 20%. Use the degree of decrease in the accuracy of the model as the discrimination criterion for reinforced neurons. If the decrease in accuracy is less than or equal to 10%, continue to select neurons from the neuron set N for addition operations. Otherwise, remove the most recently added neuron in the set M as the reinforced neuron set;
[0012] S3. Fix the neurons that do not need to be reinforced, and then perform the operation of reinforcing neurons. Specifically: calculate the gradient of the loss function of the image classification deep model and backpropagate, and fine-tune the reinforced neurons according to the gradient to finally obtain a clean image classification deep model for image classification tasks.
[0013] Further, the image dataset is selected from MNIST, CIFAR10, ImageNet, GTSRB, CASIA.
[0014] Further, the poisoning attack method is selected from BadNets, PoisonFrog, Trojannn, FeatureCollision Attack.
[0015] Further, the deep learning network is selected from LeNet, AlexNet, VGG11, and ResNet34.
[0016] Further, for poisoning in the BadNets on the MNIST dataset, 10% of the training samples with the label "0" are marked with a right-angled trigger in the upper left corner, and then the label is changed to "1" and added to the training set. Then, the model is trained. The samples with the trigger are the poisoned samples.
[0017] Further, during the determination of a certain type of main neurons, the sampling rate of the test samples is not less than 10%.
[0018] Further, when the target model is poisoned, some highly activated neurons caused by the attack will appear in the neuron set N, making the frequency of such neurons appear higher among the Top-K neurons of various test samples. The weights of such neurons are reduced to less than 20%, and it will not cause a performance drop of greater than or equal to 10% in the accuracy of the model. Based on this, the fortified neuron screening is carried out.
[0019] Further, the loss function of the image classification deep model adopts the following cross-entropy loss function loss:
[0020]
[0021] where M is the number of image categories; y ic is the sign function, taking 1 if the true category of sample i is equal to c, otherwise taking 0; P ic represents the predicted probability that sample i belongs to category c, and N represents the number of samples.
[0022] According to the second aspect of this specification, a poisoning defense device for an image classification deep model based on neuron fortification is provided, including a memory and one or more processors. An executable code is stored in the memory. When the processor executes the executable code, it is used to implement the poisoning defense method for the image classification deep model based on neuron fortification as described in the first aspect.
[0023] The beneficial effects of the present invention are mainly manifested in that for the existing poisoning defense methods of image classification deep models that do not consider the robustness after model defense, a poisoning defense method for an image classification deep model based on neuron fortification is proposed. The experimental results on real image classification deep models show that this method has good applicability, can effectively defend against poisoning attacks, and does not affect the accuracy of the model for normal samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0025] Figure 1 It is a block diagram of a poisoning defense method for an image classification deep model based on neuron reinforcement in an embodiment of the present invention.
[0026] Figure 2 It is a schematic diagram of the LeNet network structure in an embodiment of the present invention.
[0027] Figure 3 It is a structural diagram of a poisoning defense device for an image classification deep model based on neuron reinforcement in an embodiment of the present invention. Specific embodiments
[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0029] Currently, the defense methods against poisoning attacks can be classified according to the action stage into: data and feature modification, model modification, and output defense. Data and feature modification mainly refers to preprocessing the data or features before they are input into the model to achieve the defense effect. Model modification refers to modifying the model parameters to achieve the defense effect. Output defense refers to analyzing the output results of the model to achieve the defense effect.
[0030] Therefore, before deploying an image classification deep model in a practical image classification task scenario, it is necessary to strengthen its neurons. Specifically, for a model, we input a batch of clean samples into the model and, based on the neuron activation values, their occurrence frequencies, and their impact on the model performance, find specific neurons in the output of the fully connected layer, that is, the neurons to be strengthened (neurons with high activation in the main task classification). We found that poisoning attacks all target sleeping neurons (neurons not activated in the main task classification), and then these poisoned neurons will become relatively sensitive neurons, making their activation values have high values in any class of samples, thus making the model poisoned. Then strengthening the neurons with high activation in the main task classification can also resist such poisoning attacks. After we obtain the neurons to be strengthened, we need to fix the other neurons and then only strengthen these selected neurons. We input the samples in the training set into the model, update the strengthened neurons by using the gradient ascent of the loss function as a guide, and then perform the next iteration until the model converges and stabilizes, and the iteration ends. Finally, a clean image classification deep model is obtained.
[0031] Refer to Figure 1 、 Figure 2 , this embodiment provides a poisoning defense method for an image classification deep model based on neuron strengthening, and the steps are as follows:
[0032] 1) Dataset preparation:
[0033] Select image datasets such as MNIST, CIFAR10, ImageNet, GTSRB, CASIA, etc. In this embodiment, the MNIST dataset is taken as an example. This is a 10-class grayscale image dataset with a picture size of 28*28.
[0034] 2) Poisoned model preparation:
[0035] 2.1) Poisoning method: Select poisoning attack methods such as BadNets, PoisonFrog, Trojannn, Feature Collision Attack, etc. In this embodiment, the BadNets poisoning method is taken as an example.
[0036] 2.2) Deep learning network: Select networks such as LeNet, AlexNet, VGG11, ResNet34, etc. In this embodiment, the LeNet network is taken as an example. As Figure 2 shown, the LeNet network structure is mainly composed of two convolutional layers and two fully connected layers.
[0037] 2.3) Model poisoning operation: Taking BadNets poisoning on the MNIST dataset as an example, take 10% of the training samples with label "0", add a right-angled trigger patch in the upper left corner, then change the label to "1" and add it to the training set, and then start training the model. The samples with the trigger patch are poisoned samples. The accuracy of the poisoned model we trained on clean samples in the test set is 98.37%, and the accuracy on poisoned samples in the test set is 100%. This model is the poisoned model M in the embodiment.
[0038] 3) Finding neurons to be fortified:
[0039] 3.1) Counting highly activated neurons in various image classifications: Using the clean samples in the training set as test samples, select a test sample of a certain class and input it into the poisoned model M. Count the Top-K neurons of the activation values of the fully connected layer of the model, and at the same time record the K neurons with the highest frequency among the Top-K neurons corresponding to the test samples sampled on the fully connected layer as the main neurons of this class. The sampling rate of the test samples is not less than 10%. Use the test samples of each class and use the above operations to find and record the main neurons of each class.
[0040] 3.2) Recording the common part of the main neurons of each class: Among the main neurons of each class recorded in step 3.1), count the common part to form the neuron set N. The neurons we need to fortify are among them.
[0041] 3.3) Screening and determining neurons to be fortified: Select neurons in the neuron set N as the neurons we need to fortify. When the target model is poisoned, in the neuron set N, there will be some highly activated neurons caused by the attack, making such neurons appear with a higher frequency among the Top-K neurons of various test samples. Reducing the weights of such neurons to less than 20% will not cause a performance drop of greater than or equal to 10% in the accuracy of the model. We sort the neurons in the neuron set N according to the sum of the number of times they appear among the Top-K neurons of various test samples, and define the set Select neurons from the neuron set N and add them to the set M in turn from front to back. After each addition, reduce the weights of all neurons in the set M to less than 20%. Use the degree of accuracy drop of the model as the criterion for judging fortified neurons. If the accuracy drop is less than or equal to 10%, continue to select neurons in the neuron set N for addition operations, otherwise remove the most recently added neuron in the set M as the fortified neuron set.
[0042] 4) Fortifying neurons:
[0043] In the process of strengthening neurons, first of all, it is necessary to fix other neurons that do not need to be strengthened to ensure that other neurons do not change during the process of strengthening the required neurons. Then the operation of strengthening neurons can be carried out.
[0044] 4.1) The loss function of the image classification deep model adopts the following cross-entropy loss function loss:
[0045]
[0046] where M is the number of image categories; y ic is the sign function, taking 1 if the true category of sample i is equal to c, otherwise taking 0; P ic represents the predicted probability that sample i belongs to category c, and N represents the number of samples.
[0047] 4.2) Strengthen neurons: Calculate the gradient of the loss function and backpropagate, and fine-tune the strengthened neurons according to the gradient. Finally, a clean image classification deep model is obtained. Using this clean model for image classification tasks can avoid third parties using the originally added backdoors.
[0048] The image classification deep model poisoning defense method based on neuron strengthening provided by the above embodiment has the following advantages:
[0049] 1) It solves the problem that the existing commonly used model poisoning defense methods do not enhance the robustness of the model while removing poison, and the model still retains a high accuracy on the main task.
[0050] 2) Good results can be obtained only by using a small number of test samples, and it has good applicability and can meet the needs of different models and different data sets.
[0051] Corresponding to the embodiment of the image classification deep model poisoning defense method based on neuron strengthening, the present invention also provides an embodiment of an image classification deep model poisoning defense device based on neuron strengthening.
[0052] See Figure 3 , an image classification deep model poisoning defense device based on neuron strengthening provided by an embodiment of the present invention includes a memory and one or more processors. An executable code is stored in the memory. When the processor executes the executable code, it is used to implement the image classification deep model poisoning defense method in the above embodiment.
[0053] Embodiments of the poisoning defense device for the deep model of image classification based on neuron reinforcement of the present invention can be applied to any device with data processing capabilities, and such a device with data processing capabilities can be a device or apparatus such as a computer. The device embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, as Figure 3 shown, it is a hardware structure diagram of any device with data processing capabilities where the poisoning defense device for the deep model of image classification based on neuron reinforcement of the present invention is located. In addition to Figure 3 the processor, memory, network interface, and non-volatile memory shown, usually according to the actual functions of any device with data processing capabilities where the device in the embodiment is located, other hardware may also be included, which will not be elaborated here.
[0054] For the implementation processes of the functions and roles of each unit in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.
[0055] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0056] Embodiments of the present invention also provide a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the poisoning defense method for the deep model of image classification based on neuron reinforcement in the above embodiments.
[0057] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.
[0058] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0059] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0060] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the" and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0061] It should be understood that although the terms first, second, third, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0062] The foregoing are only preferred embodiments of one or more embodiments of this specification and are not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of protection of one or more embodiments of this specification.
Claims
1. A poisoning defense method for deep learning models based on neuron reinforcement, characterized in that, It includes the following steps: S1. Prepare an image dataset, select a deep learning network, generate poisoned samples using a poisoning attack method and add them to the training set, and train to obtain a poisoned model; S2. Find the neurons that need to be fortified, including the following sub-steps: S2.
1. Use the clean samples in the training set as test samples. Select a certain test sample of a certain class and input it into the poisoned model. Statistically analyze the Top-K neurons of the activation values of the fully connected layer of the model. At the same time, record the K neurons with the highest frequency among the Top-K neurons corresponding to the test samples sampled on the fully connected layer as the main neurons of this class. Finally, obtain the main neurons of each class; S2.
2. Statistically analyze the common part among the main neurons of each class to form a neuron set N; S2.
3. Sort each neuron in the neuron set N according to the sum of the number of times it appears in the Top-K neurons of each type of test sample, and define the set Select neurons from the neuron set N one by one from front to back and add them to the set M. After each addition, reduce the weights of all neurons in the set M to less than 20%. Use the degree of decrease in the accuracy of the model as the discriminant criterion for fortified neurons. If the decrease in accuracy is less than or equal to 10%, continue to select neurons in the neuron set N for addition operations. Otherwise, remove the most recently added neuron in the set M as the fortified neuron set; S3. Fix the neurons that do not need to be fortified, and then perform operations on the fortified neurons. Specifically: calculate the gradient of the loss function of the image classification deep model and backpropagate. Fine-tune the fortified neurons according to the gradient to finally obtain a clean image classification deep model for image classification tasks.
2. The poisoning defense method for a deep learning model based on neuron reinforcement according to claim 1, wherein The image dataset is selected from MNIST, CIFAR10, ImageNet, GTSRB, CASIA.
3. The poisoning defense method for a deep learning model based on neuron reinforcement according to claim 1, characterized in that, The poisoning attack method is selected from BadNets, PoisonFrog, Trojannn, Feature Collision Attack.
4. The poisoning defense method for a deep learning model based on neuron reinforcement according to claim 1, wherein, The deep learning network is selected from LeNet, AlexNet, VGG11, ResNet34.
5. The poisoning defense method for a deep learning model based on neuron reinforcement according to claim 1, wherein For BadNets poisoning on the MNIST dataset, take 10% of the training samples with the label "0" and mark a right-angled trigger in the upper left corner, then change the label to "1" and add them to the training set, and then start training the model. The samples with the trigger are the poisoned samples.
6. The poisoning defense method for a deep learning model based on neuron reinforcement according to claim 1, wherein During the determination process of the main neurons of a certain class, the sampling rate of the test samples is not less than 10%.
7. The poisoning defense method for a deep learning model based on neuron reinforcement according to claim 1, wherein When the target model is poisoned, there will be some highly activated neurons caused by the attack in the neuron set N, making the frequency of such neurons appear higher among the Top-K neurons of various test samples. Reduce the weights of such neurons to less than 20%, and it will not cause a performance drop of greater than or equal to 10% in the accuracy of the model. Based on this, fortified neuron screening is carried out.
8. The method for poisoning defense of a deep learning model based on neuron reinforcement according to claim 1, characterized in that, The loss function of the image classification deep model adopts the following cross-entropy loss function loss: where M is the number of image categories; y ic is the sign function, taking 1 if the true category of sample i is equal to c, and 0 otherwise; P ic represents the predicted probability that sample i belongs to category c, and N represents the number of samples.
9. A deep learning model poisoning defense device based on neuron reinforcement, comprising a memory and one or more processors, wherein executable code is stored in the memory, and is characterized in that When the processor executes the executable code, it is used to implement the poisoning defense method of the deep learning model based on neuron fortification as described in any one of claims 1-8.