Detoxification and Reinforcement Method for Deep Learning Model Based on Master Task Neurons

By identifying and deleting differentiated main task neurons in deep learning models and retraining, the method enhances robustness against poisoning attacks without significantly affecting accuracy.

CN115600670BActive Publication Date: 2025-07-15ZHEJIANG UNIV OF TECH

Patent Information

Application Number
CN202211287984.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-20
Publication Date
2025-07-15
Estimated Expiration
2042-10-20

AI Technical Summary

Technical Problem

The existing deep learning model poisoning prevention methods fail to effectively improve the robustness of the model while detoxifying, and may affect the normal performance of the model.

Method used

By identifying and deleting differential parts between the main task neurons of different categories in the deep learning model, the poisoning samples generated by the poisoning attack are trained using the poisoning attack, the activation values of the neurons are calculated, the model is retrained in grouping and processing is retrained to delete the key neurons, and the model parameters are optimized using the cross entropy loss function.

Benefits of technology

It improves the robustness of the model, while maintaining the high prediction accuracy of the model on normal samples, and effectively defends against poisoning attacks, and only a small number of test samples can achieve better results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115600670B_ABST
    Figure CN115600670B_ABST
Patent Text Reader

Abstract

The present invention discloses a detoxification and reinforcement method for a deep learning model based on main task neurons, including: selecting an image data set, generating poisoned samples, and training a poisoned model based on the poisoned samples; inputting the image data set into the poisoned model, calculating the activation values of the neurons corresponding to each type of picture in the image data set, sorting the activation values of all neurons, and taking the top n neurons as the neurons with high activation values; taking the overlapping part of the neurons with high activation values as the main task neurons of this type; counting the main task neurons of each type in the image data set. Divide the different parts of the main task neurons of each class into K groups, select one of the groups to delete the neurons in this group in the poisoned model, and calculate the performance of the poisoned model respectively; according to the performance change, obtain the set of neurons to be deleted; delete the set of neurons to be deleted in the poisoned model, and retrain the poisoned model to obtain a detoxified and reinforced deep learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of security issues for deep learning models, and specifically relates to a detoxification and reinforcement method for deep learning models based on main task neurons. Background Art

[0002] Currently, the defense methods against poisoning attacks can be classified into three categories according to the action stage: data and feature modification, model modification, and output defense. Data and feature modification mainly refers to modifying at the data level to achieve the defense effect. Model modification refers to modifying the model parameters to achieve the defense effect. Output defense refers to analyzing the output results of the model to achieve the defense effect.

[0003] Therefore, when we obtain a model and need to deploy it in a practical scenario, we need to perform a detoxification and reinforcement operation on the deep learning model based on main task neurons. Specifically, for a model, we input multiple clean samples belonging to the same category into the model and find the Top-K neurons in the output of the fully connected layer. These neurons are the main task neurons of this category. We found that different types of images correspond to different main task neurons, and there is a phenomenon of high overlap between the main task neurons of each other. The difference in main task neurons is very small among different types. The poisoning task also has its poisoned task neurons. These neurons are all extremely important neurons, and when these neurons are deleted, it will have a greater impact on its task. Poisoning attacks will cause the activation values of some neurons in the model to increase, resulting in a larger difference in the main task neurons of different types of the model. According to this characteristic, we can first find the different parts of the main task neurons of each category. And perform a fine operation to delete neurons, thereby achieving the detoxification effect.

[0004] After we obtain the neurons of each category, we remove the common parts of each category and leave the different parts. These different parts are divided into K groups. Then, the following operations are performed on the neurons in each group: set their weights to zero. Observe the change in the prediction success rate and prediction value of the model for this category before and after the neurons are deleted. Based on this, decide whether to delete the neurons in this group. Summary of the Invention

[0005] Currently, the poisoning and defense of deep learning models are in a continuous game situation. The defense methods against poisoning attacks can be classified into three categories according to the action stage: data and feature modification, model modification, and output defense. In order to improve the robustness of the model against poisoning attacks, we reinforce the neurons in the model that activate the main task classification. A detoxification and reinforcement method for deep learning models based on main task neurons is proposed.

[0006] The technical solution of the present invention is as follows: In the first aspect of the embodiments of the present invention, a detoxification and reinforcement method for a deep learning model based on main task neurons is provided. The method includes the following steps:

[0007] S1, Select an image dataset, choose a deep learning network, use a poisoning attack method to add triggers to some clean samples in the image dataset to generate poisoned samples, and train a poisoned model based on the poisoned samples;

[0008] S2, Input the image dataset selected in step S1 into the trained poisoned model, calculate the activation values of the neurons corresponding to each category of pictures in the image dataset, sort the activation values of all neurons, customize a sorting threshold n, and take the first n neurons as the neurons with high activation values; For the neurons with high activation values corresponding to all pictures in each category, take the overlapping part as the main task neurons of this category; Count the main task neurons of each category in the image dataset.

[0009] S3, Divide the different parts of the main task neurons of each category into K groups, sequentially take one of the K groups, set the weights of the neurons corresponding to this group to 0, that is, delete the neurons in this group in the poisoned model, input the picture category in the image dataset corresponding to this group into the poisoned model, and calculate the performance of the poisoned model before and after deleting the neurons in this group; According to the prediction accuracy and confidence of the poisoned model before and after, determine whether the neurons in this group are the neurons to be deleted; Repeat the above steps, traverse the K groups, and obtain the set of neurons to be deleted;

[0010] S4, Delete the set of neurons to be deleted obtained in step S3 in the poisoned model, and retrain the poisoned model to obtain a detoxified and reinforced deep learning model for image classification tasks.

[0011] In the first aspect of the embodiments of the present invention, an electronic device is provided, including a memory and a processor, and the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned detoxification and reinforcement method for a deep learning model based on main task neurons.

[0012] In the second aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, it implements the above-mentioned detoxification and reinforcement method for a deep learning model based on main task neurons.

[0013] The beneficial effects of the present invention are mainly manifested in:

[0014] The present invention proposes a detoxification and reinforcement method for a deep learning model based on main task neurons. After finding neurons of each type, the common part among each type is removed, and the different part is left to obtain the different parts of the main task neurons of each type. These different parts are divided into K groups. Their weights are set to zero, that is, the neurons in this group are deleted in the poisoned model, and the prediction success rate and the degree of change in the predicted value of the poisoned model in this category are calculated. Based on this, it is determined whether the neurons in this group are deleted. Thus, the detoxification effect is achieved. The experimental results on real deep learning models show that the method of the present invention has good applicability, can effectively defend against poisoning attacks, and does not affect the accuracy of normal samples of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a block diagram of a detoxification and reinforcement method for a deep learning model based on main task neurons in an embodiment of the present invention.

[0017] Figure 2 It is a schematic diagram of an electronic device proposed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The following will further describe in detail the specific embodiments of the present invention with reference to the drawings in the specification.

[0019] The technical concept of the present invention is as follows: Based on the existing poisoning defense methods that do not fully consider the time cost of training and the robustness of the model. We first propose a detoxification and reinforcement method for a deep learning model based on main task neurons. We utilize that poisoning attacks will increase the differences between the main task neurons of different classes of the model. By deleting the different parts of the main task neurons, the effects of improving the robustness of the model and detoxification are achieved simultaneously.

[0020] Referring to Figures 1 to 2 , an embodiment of the present invention proposes a detoxification and reinforcement method for a deep learning model based on main task neurons. The specific steps of the method are as follows:

[0021] S1, Prepare an image data set, select a deep learning network, use a poisoning attack method to add triggers to some clean samples in the image data set to generate poisoned samples, and train a poisoned model based on the poisoned samples.

[0022] The image dataset is selected from MNIST, CIFAR10, ImageNet, GTSRB, CASIA and other image datasets. The poisoning attack method is selected from BadNets, PoisonFrog, Trojannn, Feature Collision Attack. The deep learning network is selected from LeNet, AlexNet, VGG11, ResNet34.

[0023] S2, main task neuron, and take out the difference part: input the image data set selected in step S1 into the trained poisoning model, calculate the activation value of the neuron corresponding to each type of image in the image data set, sort the activation values of all neurons, customize the sorting threshold n, and take the first n neurons as neurons with high activation values.

[0024] After all the images in the image dataset are input into the poisoning model, each image has its corresponding high-activated neuron. For each class, the overlapping part of the neurons with high activated values corresponding to all the images is taken as the main task neuron of that class. Using the above method, the main task neurons of each class in the image dataset are counted.

[0025] Compare the main task neurons of each class in the image data set, remove the main task neurons shared by each class, and what remains are the differences in the main task neurons of each class.

[0026] S3, detoxification based on the difference between the main task neurons:

[0027] In the process of detoxification, we must ensure that the parameters of the remaining neurons remain unchanged except for the neurons we want to operate on, ensuring that the performance of the model is not affected by other factors.

[0028] S301, dividing the difference parts of the main task neurons of each class obtained in step S2 into K groups, where K is the number of groups set by user.

[0029] S302, sequentially select one of the K groups, reset the weights corresponding to the neurons in the group to 0 (i.e., delete the neurons in the group in the poisoning model), input the image class in the image data set corresponding to the group into the poisoning model, and calculate the performance of the poisoning model before and after deleting the neurons in the group. The formula is as follows:

[0030] TC(X)=ACC aft -ACC pre

[0031]

[0032] Among them, X represents the test picture samples corresponding to the picture classes in the image dataset corresponding to this group. n represents the number of samples contained in X. TC(X) represents the change in the accuracy of the poisoned model predicting the sample group X before and after deleting neurons. ACC pre represents the accuracy of the poisoned model before deleting neurons, ACC aft represents the accuracy of the poisoned model after deleting neurons. x represents the label of the samples within the group, that is, the corresponding picture class in the image dataset. Aft(x) represents the confidence of the poisoned model in predicting that the sample belongs to class x after deleting neurons. Pre(x) represents the confidence of the poisoned model in predicting that the sample belongs to class x before deleting neurons.

[0033] Before and after deleting this group of neurons, if the performance degradation of the poisoned model is less than 5%, then check the average change in the confidence of the poisoned model in predicting that the sample belongs to class x (i.e., the value of the cc function). If the average confidence decrease does not exceed 25%, then add this group of neurons to the set of neurons to be deleted. Otherwise, take the next group of neurons and perform the above operations. Repeat the above steps, traverse K groups, and obtain the set of neurons to be deleted.

[0034] S4, Detoxification: Delete the neurons in the set of neurons to be deleted obtained in step 302 from the poisoned model, and retrain the poisoned model again to obtain a detoxified and strengthened deep learning model.

[0035] Specifically, use the cross-entropy loss function to retrain the poisoned model. Calculate the gradient of the loss function and backpropagate, and update the parameters in the poisoned model according to the gradient. Until the cross-entropy loss function converges, finally obtain a clean detoxified and strengthened deep learning model. Among them, the loss function for retraining adopts the following cross-entropy loss function loss:

[0036]

[0037] Among them, M is the number of image classes; y ic is the sign function, taking 1 if the true class of sample i is equal to c, otherwise taking 0; P ic represents the predicted probability that sample i belongs to class c, and N represents the number of samples.

[0038] Example 1:

[0039] S1, In this example, taking the MNIST dataset as an example, the MNIST dataset is a grayscale image dataset with a picture size of 28*28 and 10 classifications. Select the LeNet deep learning network as an example. The LeNet deep learning network structure mainly consists of two convolutional layers and three fully connected layers. Taking the BadNets poisoning method as an example.

[0040] Specifically, taking the poisoning of BadNets on the MNIST dataset as an example, 10% of the clean samples with the label "0" in the image dataset are marked with a right-angled trigger to generate poisoned samples. The samples marked with the right-angled trigger are the poisoned samples. Then, the label is changed to "1" and added to the training dataset. Then, the deep learning network is trained to obtain the poisoned model M. The accuracy of the poisoned model M trained in the embodiment of the present invention on clean samples in the test set is 98.37%, and the accuracy on poisoned samples in the test set is 100%.

[0041] S2. Main task neurons and extract the different parts: Input the image dataset selected in step S1 into the trained poisoned model, calculate the activation values of the neurons corresponding to each type of picture in the image dataset, sort the activation values of all neurons, customize the sorting threshold to 5, and take the top 5 neurons as the neurons with high activation values.

[0042] After inputting all the pictures in the image dataset into the poisoned model, each picture has its corresponding neurons with high activation values. For the neurons with high activation values corresponding to all the pictures in each category, take the overlapping part as the main task neurons of that category. Using the above method, count the main task neurons of each category in the image dataset.

[0043] Compare the main task neurons of each category in the image dataset, remove the common part of the main task neurons of each category, and the remaining is the different part of the main task neurons of each category.

[0044] Specifically, taking the lenet model poisoned by badnet on the MNIST dataset as an example, we performed the above operations on the poisoned model. The different parts of the main neurons of each category were selected.

[0045] S3. Detoxify based on the different parts of the main task neurons:

[0046] During the detoxification process, we need to ensure that the parameters of the remaining neurons remain unchanged except for the neurons in the part we want to operate on. Ensure that the performance change of the model is not affected by other factors.

[0047] S301. Divide the different parts of the main task neurons of each category obtained in step S2 into K groups, where K is the number of groups set by the user.

[0048] S302. Take one group from the K groups in turn, set the weights of the neurons in this group to 0 (i.e., delete the neurons in this group in the poisoned model), input the picture category in the image dataset corresponding to this group into the poisoned model, and calculate the performance of the poisoned model before and after deleting the neurons in this group. The formula is as follows:

[0049] TC(X) = ACC aft -ACC pre

[0050]

[0051] Among them, X represents the test image samples corresponding to the picture class in the image dataset corresponding to this group. n represents the number of samples contained in X. TC(X) represents the change in the accuracy of the poisoned model predicting the sample group X before and after deleting neurons. ACC pre represents the accuracy of the poisoned model before deleting neurons, ACC aft represents the accuracy of the poisoned model after deleting neurons. x represents the label of the samples within the group, that is, the corresponding picture class in the image dataset. Aft(x) represents the confidence that the poisoned model predicts the sample to belong to class x after deleting neurons. Pre(x) represents the confidence that the poisoned model predicts the sample to belong to class x before deleting neurons.

[0052] Before and after deleting the neurons in this group, if the performance degradation is less than 5%, then check the average change in the confidence that the model predicts the sample to belong to class x (i.e., the cc function value). If the average confidence decrease does not exceed 25%, then add the neurons in this group to the set of neurons to be deleted. Otherwise, take the next group of neurons and perform the above operations.

[0053] Specifically, taking the lenet model poisoned by badnet in the MNIST dataset as an example, for class 0, we selected 4 neurons as the different part of its main task neurons. We selected K = 2. And used the above operations to select 2 neurons to be deleted.

[0054] S4, Detoxification: Delete the neurons in the set of neurons to be deleted in step 302 from the poisoned model, and retrain the poisoned model again to obtain a detoxified and fortified deep learning model.

[0055] Specifically, use the cross-entropy loss function to retrain the poisoned model. Calculate the gradient of the loss function and backpropagate, and update the parameters in the poisoned model according to the gradient. Until the cross-entropy loss function converges, finally obtain a clean detoxified and fortified deep learning model. Among them, the loss function for retraining uses the following cross-entropy loss function loss:

[0056]

[0057] Among them, M is the number of image classes; y ic is the sign function, taking 1 if the true class of sample i is equal to c, otherwise taking 0; P icrepresents the predicted probability that sample i belongs to class c, and N represents the number of samples.

[0058] Specifically, we use the poisoned LeNet model on MNIST. After the above operations, its accuracy on clean samples in the test set is 98.53%, and the accuracy on poisoned samples in the test set is 0%.

[0059] The deep learning model detoxification and reinforcement method based on the main task neurons provided by the above embodiments has the following advantages:

[0060] 1) It solves the problem that the existing commonly used deep learning model poisoning defense methods do not enhance the model's robustness while detoxifying, and the method of the present invention enables the deep learning model after detoxification and reinforcement to still maintain a high accuracy on the main task.

[0061] 2) The method of the present invention only needs to use a small number of test samples to obtain good results, with high efficiency.

[0062] 3) The method of the present invention has good applicability and can meet the requirements of different deep learning models and different data sets.

[0063] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0064] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0065] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 or more processes and / or blocks Figure 1 or the functions specified in a block or more blocks.

[0066] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process Figure 1 or more processes and / or blocks Figure 1 or the functions specified in a block or more blocks.

[0067] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.

Claims

1. A detoxification and reinforcement method for a deep learning model based on main task neurons, characterized in that, The method includes the following steps: S1. Select an image dataset, choose a deep learning network, use a poisoning attack method to add triggers to some clean samples in the image dataset to generate poisoned samples, and train a poisoned model based on the poisoned samples; S2. Input the image dataset selected in step S1 into the trained poisoned model, calculate the activation values of the neurons corresponding to each class of pictures in the image dataset, sort the activation values of all neurons, customize a sorting threshold n, and take the first n neurons as the neurons with high activation values; for the neurons with high activation values corresponding to all pictures in each class, take the overlapping part as the main task neurons of this class; count the main task neurons of each class in the image dataset. S3. Divide the different parts of the main task neurons of each class into K groups, sequentially take one of the K groups, set the weights of the neurons in this group to 0, that is, delete the neurons in this group in the poisoned model, input the picture class in the image dataset corresponding to this group into the poisoned model, and calculate the performance of the poisoned model before and after deleting the neurons in this group; based on the prediction accuracy and confidence of the poisoned model before and after, determine whether the neurons in this group are the neurons to be deleted; repeat the above steps, traverse the K groups, and obtain the set of neurons to be deleted; S4. Delete the set of neurons to be deleted obtained in step S3 in the poisoned model, and retrain the poisoned model to obtain a detoxified and fortified deep learning model for image classification tasks.

2. The method for detoxifying and strengthening a deep learning model based on principal task neurons according to claim 1, wherein The image dataset is selected from MNIST, CIFAR10, ImageNet, GTSRB or CASIA image datasets.

3. The detoxification and reinforcement method for a deep learning model based on principal task neurons according to claim 1, characterized in that, The poisoning attack method is selected from BadNets, PoisonFrog, Trojannn, Feature Collision Attack.

4. The detoxification and reinforcement method for a deep learning model based on a main task neuron according to claim 1, characterized in that, The deep learning network is selected from LeNet, AlexNet, VGG11, ResNet34.

5. The method for poisoning defense of a deep learning model based on neuron reinforcement according to claim 1, characterized in that, For BadNets poisoning on the MNIST dataset, add triggers to 10% of the clean samples with the label "0" in the MNIST dataset. The samples with triggers added are the poisoned samples, then change the label to "1" and add them to the training set, and then start the deep learning network to train a poisoned model.

6. The detoxification and reinforcement method for a deep learning model based on main task neurons according to claim 1, wherein, Step S3 specifically includes the following sub-steps: S301. Divide the different parts of the main task neurons obtained in step S2 into K groups, where K is the number of groups set by the user; S302. Sequentially take one of the K groups, set the weights of the neurons in this group to 0, that is, delete the neurons in this group in the poisoned model, input the picture class in the image dataset corresponding to this group into the poisoned model, and calculate the performance of the poisoned model before and after deleting the neurons in this group. The formula is as follows: TC(X) = ACC aft -ACC pre Among them, X represents the test picture sample corresponding to the picture class in the image dataset corresponding to this group; n represents the number of samples contained in X; TC(X) represents the change in the accuracy of the poisoned model predicting the sample group X before and after deleting neurons; ACC pre represents the accuracy of the poisoned model before deleting neurons, ACC aft represents the accuracy of the poisoned model after deleting neurons; x represents the label of the sample within the group, that is, the corresponding picture class in the image dataset; Aft(x) represents the confidence of the poisoned model in predicting that the sample belongs to class x after deleting neurons; Pre(x) represents the confidence of the poisoned model in predicting that the sample belongs to class x before deleting neurons; Before and after deleting this group of neurons, if the performance degradation of the poisoned model is less than 5%, then check the average change in the confidence of the poisoned model's prediction of samples belonging to class x; if the average confidence degradation does not exceed 25%, then add this group of neurons to the set of neurons to be deleted; otherwise, take the next group of neurons and perform the above operations; repeat the above steps, traverse K groups, and obtain the set of neurons to be deleted.

7. The detoxification and reinforcement method of the deep learning model based on the main task neurons according to claim 1, wherein, The specific steps of step S4 are as follows: retrain the poisoned model using the cross-entropy loss function; calculate the gradient of the loss function and backpropagate it, and update the parameters in the poisoned model according to the gradient; until the cross-entropy loss function converges, and finally obtain a clean and detoxified and fortified deep learning model.

8. The detoxification and reinforcement method of the deep learning model based on the main task neuron according to claim 7, characterized in that, The formula of the cross-entropy loss function is as follows: where M is the number of image categories; y ic is the sign function, taking 1 if the true category of sample i is equal to c, and 0 otherwise; P ic represents the predicted probability that sample i belongs to category c, and N represents the number of samples.

9. An electronic device, comprising a memory and a processor, characterized in that The memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the detoxification and fortification method of the deep learning model based on the main task neurons according to any one of claims 1-8 above.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the detoxification and fortification method of the deep learning model based on the main task neurons according to any one of claims 1-8.

Citation Information

Patent Citations

  • Image classification depth model poisoning defense method and device based on neuron reinforcement

    CN115688866A

Cited By

  • Defense method for power system poisoning attack

    CN121509115A

  • A defense method against poisoning attacks on power systems

    CN121509115B