A Backdoor Defense Method and System Based on Reverse Forgetting
Through the backdoor defense method based on reverse forgetting, the clean sample set is used to screen and train the model, and the sample detection model is trained in combination with cross-entropy loss and entropy constraints, the problem of difficulty in detection and clearing of backdoor attacks in the existing technology is solved, and effective defense of complex backdoor attacks and model performance protection is achieved.
Patent Information
- Application Number
- CN202510520160.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Existing backdoor defense methods have limitations in the face of more concealed or complex backdoor attacks, making it difficult to completely clear backdoors and may have a negative impact on model performance.
Using a backdoor defense method based on reverse forgetting, the pre-trained model is trained using the original clean sample set, the potential clean sample set is screened out from the contaminated data set, a new clean sample set is synthesized, and the pre-trained model is further trained. At the same time, the sample detection model is trained using cross entropy loss and entropy constraints to detect toxic samples.
This method can effectively detect and eliminate complex backdoor attacks without damaging the performance of the model, improve the model's defense ability against multiple data poisoning backdoor attacks, and enhance the security and practicality of the model.
Smart Images

Figure CN120030542B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence security, and in particular relates to a backdoor defense method and system based on reverse forgetting. Background Art
[0002] As a powerful artificial intelligence technology, deep neural network (DNN) has been widely used in fields such as speech recognition, image classification and autonomous driving. Its main advantage is that it can generate high-performance models through large-scale data training to solve complex practical problems. However, training an efficient deep neural network model often requires a large amount of computing resources and diverse data support, which makes the model development process costly and places high demands on the quality and security of the data. At the same time, deep neural networks are also exposed to the risk of backdoor attacks, which implant trigger patterns in the model so that the model performs well under normal input, but outputs the wrong results preset by the attacker when certain conditions are triggered. The concealment and high harm of backdoor attacks pose a serious threat to the security of the system, which may not only lead to the destruction of the integrity of the model, but also be used maliciously, bringing significant risks to data security and application results.
[0003] In response to backdoor attacks, existing backdoor defense methods mainly focus on backdoor detection and backdoor elimination. Backdoor detection usually discovers potential backdoor attacks by analyzing the feature differences of contaminated samples or observing the abnormal behavior of the model; while backdoor elimination relies on methods such as forgetting learning to remove the implanted backdoor features from the model. However, these methods have significant limitations when facing more hidden or complex backdoor attacks. The trigger patterns designed by attackers may bypass existing detection methods, making it difficult to completely remove the backdoor. At the same time, over-reliance on the means of forgetting backdoor features may have a negative impact on the normal performance of the model, weakening its robustness and accuracy for uncontaminated samples. Therefore, in the current backdoor attack and defense research, how to maximize the performance of the model in the process of removing the backdoor is still an important problem that needs to be solved. Summary of the invention
[0004] In view of the deficiencies in the prior art, the purpose of the present invention is to provide a backdoor defense method and system based on reverse forgetting, which solves the problems in the prior art.
[0005] The purpose of the present invention can be achieved through the following technical solutions:
[0006] A backdoor defense method based on reverse forgetting includes the following steps:
[0007] Train a pre-trained model using the original clean sample set, and use the pre-trained model to screen out a potential clean sample set from the contaminated dataset to synthesize a new clean sample set, and further train the pre-trained model using the new clean sample set;
[0008] Input the contaminated dataset and the new clean sample set into the sample detection model, and train them using cross-entropy loss and entropy constraint respectively;
[0009] Input the contaminated dataset into the trained sample detection model for prediction to detect poisoned samples.
[0010] Furthermore, the steps of synthesizing the new clean sample set include:
[0011] Use the original clean sample set to train the pre-trained model ;
[0012] In each round of training, use the pre-trained model to calculate the loss value of each sample in the contaminated dataset , and screen out the potential clean sample set;
[0013] Merge the screened potential clean sample set with the original clean sample set to form a new clean sample set .
[0014] Furthermore, the calculation formula of the loss value is:
[0015]
[0016] where is the loss value, is a single category in the classification task, is the sample in the category true label, is the probability value that the model predicts the input as the category , is the total number of categories in the classification task.
[0017] Furthermore, the criteria for screening out the potential clean sample set are:
[0018]
[0019] where is the potential clean sample set, is the preset threshold.
[0020] Further, during the process of training the sample detection model, the parameters of the sample detection model are updated by the gradient descent method to minimize the total loss function, where the total loss function of the sample detection model is:
[0021]
[0022]
[0023]
[0024]
[0025] where is the total loss function, is the classification loss, is the entropy regularization term, is the trade-off coefficient, N is the total number of samples, is the information entropy of the variable , is the variable the number of all possible results, is the probability of the result appearing under the condition of the input being is the sample in the category true label.
[0026] Further, the steps to detect poisoned samples are:
[0027] Input the contaminated dataset into the trained sample detection model, and through the feature extraction and prediction mechanism of the sample detection model, generate a set of predicted entropy values for each sample ;
[0028] Set a predefined entropy threshold , and based on the entropy threshold judge each entropy value in the entropy value set . Samples with entropy values exceeding the entropy threshold are clean samples, otherwise they are poisoned samples.
[0029] A backdoor defense system based on reverse forgetting, comprising:
[0030] Sample screening module: Use the original clean sample set to train a pre-trained model, and use the pre-trained model to screen out a set of potential clean samples from the contaminated dataset to synthesize a new clean sample set, and use the new clean sample set to further train the pre-trained model;
[0031] Model training module: inputting the contaminated data set and the new clean sample set into the sample detection model, and training them using cross entropy loss and entropy constraint respectively;
[0032] And, the sample detection module: the contaminated data set is input into the trained sample detection model for prediction to detect the poisoned samples.
[0033] A computer storage medium stores a readable program, which can execute the above-mentioned backdoor defense method based on reverse forgetting when the program is running.
[0034] An electronic device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus;
[0035] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the above-mentioned backdoor defense method based on reverse forgetting.
[0036] A computer program product includes computer instructions, wherein the computer instructions instruct a computing device to execute operations corresponding to the above-mentioned backdoor defense method based on reverse forgetting.
[0037] Beneficial effects of the present invention:
[0038] 1. The core of the method of the present invention is to reverse forget the model feature performance of clean samples and highlight the main features of backdoor poisoned samples, rather than directly looking for the features of poisoned samples for judgment. The advantage of this method is that it has a more stable defense effect against newer sample poisoning methods and breaks away from the limitations of traditional backdoor detection methods.
[0039] 2. In the stage of collecting clean samples for the pre-training model, the present invention avoids collecting too many poisoned samples through the strategy of "collecting while training", thereby affecting the effect of subsequent clean sample constraints.
[0040] 3. The present invention uses polluted data sets and clean data sets to train sample detection models. Cross entropy loss is used for normal training of polluted data sets. Entropy constraints are used for clean data sets to constrain the predicted distribution of clean samples, highlighting the low entropy value performance of poisoned samples and realizing the detection of poisoned samples. This detection method greatly enhances the model's defense capabilities against a variety of data poisoning backdoor attacks, such as patch embedded data poisoning (Badnets), hidden data poisoning (Blended), and image distortion data poisoning (Wanet). This poisoned sample detection method based on clean sample entropy constraints not only improves the neural network model's defense capabilities against backdoor attacks, but also enhances the model's practicality and security in real-world scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 is the flowchart of the present invention for pre-training and screening clean samples;
[0043] Figure 2 is the schematic diagram of the present invention for training a sample detection model;
[0044] Figure 3 is the schematic diagram of the present invention for detecting poisoned samples;
[0045] Figure 4 is the entropy value distribution diagram of the sample detection model for detecting contaminated data samples. Specific Embodiments
[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0047] Embodiment 1
[0048] A backdoor defense method based on reverse forgetting includes the following steps:
[0049] S1, training a pre-training model using the original clean sample set, and using the pre-training model to screen out a potential clean sample set from the contaminated data set to synthesize a new clean sample set, and further training the pre-training model using the new clean sample set;
[0050] During the training process of the pre-training model, the Stochastic Gradient Descent (SGD) algorithm is used to adjust the network weight parameters, aiming to optimize the performance of the model. The specific training parameters include setting the learning rate to 0.01 and performing 50 training epochs. This setting helps the model to focus more on the features of the existing clean samples during the learning process, so as to effectively distinguish other clean samples from the contaminated data set.
[0051] Afterwards, the samples in the contaminated dataset are input into the pre-trained model for prediction to obtain a list of losses for the samples. By sorting these loss values, 50 to 100 samples with the lowest losses in each category are selected and incorporated into the clean dataset. Thereafter, the original pre-trained model is further trained using the new clean dataset, aiming to enhance the model's ability to recognize clean samples and reduce the risk of misidentifying poisoned samples as clean samples. This screening and re-training process needs to be repeated until the size of the clean dataset reaches 10% to 20% of the original contaminated dataset.
[0052] In this embodiment, the pre-trained model is a CNN model. After its input layer receives image data, it is transmitted to the convolutional layer, where local features are extracted through convolutional operations. The generated feature maps are then downsampled in the pooling layer. In some embodiments, a deeper CNN is adopted. After multiple groups of convolutional-pooling layer combinations fully extract multi-layer features through multiple rounds of operations, the feature maps are transmitted to the flattening layer to be converted into one-dimensional vectors. The fully connected layer performs non-linear combination mapping on them, and the output generates a class probability distribution through Softmax to complete classification.
[0053] As Figure 1 shown, the steps of synthesizing a new clean sample set include:
[0054] S11, using a small number of original clean sample sets to train a pre-trained model ;
[0055] The goal of this pre-trained model is to extract the basic feature distribution of clean samples. Different from traditional methods, this method adopts a "recruiting while training" strategy, that is, in each round of iteration during the training of the pre-trained model, potential clean samples are dynamically screened from the contaminated dataset, gradually expanding the clean sample set. This dynamic screening method can ensure that the selected clean samples are optimized synchronously with the model's feature extraction ability, thereby reducing the risk of poisoned samples being mixed in during the screening process.
[0056] S12, in each round of iteration, using the current pre-trained model , calculate the loss value of each sample in the contaminated dataset , and screen out the set of potential clean samples; ;
[0057] During the process of training the pre-trained model , the cross-entropy loss function is used to optimize the model parameters:
[0058]
[0059] where, is the loss value, is the sample The true label for the category is, the probability value predicted by the model for the input as the category . is a single category in the classification task, and is the total number of categories in the classification task.
[0060] In each round of training, for the contaminated dataset, the loss value of each sample is calculated using the current model, and the samples with loss values lower than the preset threshold are temporarily marked as potential clean samples. Subsequently, these potential clean samples are added to the clean sample set to participate in the next round of training, thereby continuously improving the model's ability to extract features of clean samples.
[0061] The criteria for screening out the set of potential clean samples are:
[0062]
[0063] where, is the set of potential clean samples, is the preset threshold.
[0064] S13, the set of potential clean samples screened out is merged with the original clean sample set to form a new clean sample set :
[0065] .
[0066] S2, the contaminated dataset and the newly synthesized clean sample set in S1 are input into the sample detection model, and cross-entropy loss and entropy constraint are used for training respectively;
[0067] As Figure 2 shown, the defender puts the initially obtained contaminated dataset into the sample detection model for training, and uses cross-entropy loss for optimization in this training, aiming to let the sample detection model learn the overall data features in the contaminated dataset. These data features include the features of clean samples and poisoned samples. At the same time, the defender puts the newly synthesized clean sample set in S1 into the sample detection model for entropy constraint training. This step is to let the detection model forget the features of clean samples in the contaminated dataset, so as to highlight the feature performance of poisoned samples in the model.
[0068] In this embodiment, the sample detection model is a ResNet-18 model, which consists of six major modules to form a complete feature learning system, including: an input layer, a convolutional layer, a max pooling layer, a residual block module, an average pooling layer, and a fully connected layer. The input layer receives the standardized image data and then passes the data to the 7×7 convolutional layer (Conv1). This convolutional layer performs convolutional operations with a stride of 2, extracts basic features such as edges and textures, and reduces the spatial dimension. Then, it conveys the processed feature map to the max pooling layer. The max pooling layer further compresses the size of the feature map and enhances translational invariance, and then outputs it to the core residual block module. The residual block module constructs a feature pyramid in 4 stages through skip connections. In the first residual block of each stage, a 1×1 convolution is used to adjust the channel dimension, and in the subsequent residual blocks, a 3×3 convolution is used to capture the middle-level semantic features. After the feature processing, the deep feature map is passed to the global average pooling layer. The global average pooling layer compresses the deep feature map into a one-dimensional vector and then transmits this vector to the fully connected layer. The fully connected layer generates a class probability distribution based on this vector, and finally completes the final classification task through the Softmax activation function.
[0069] During the process of training the sample detection model, a total loss function is constructed to train the sample detection model; the construction process of the total loss function is as follows:
[0070] S21, calculate the entropy of the predicted probability distribution of each input sample to measure the uncertainty of the model prediction;
[0071] Among them, the entropy is defined as:
[0072]
[0073] Among them, is the number of all possible results of the variable , is the probability of the result being under the condition of the input ;
[0074] S22, by calculating the average value of the entropies of all samples on the new clean sample set and taking the negative value to construct the entropy regularization term so as to maximize the entropy in the loss function;
[0075] The calculation formula of the entropy regularization term is:
[0076]
[0077] S23, use cross-entropy loss to define the classification loss , to measure the performance of the model in the classification task; classification loss The calculation formula is as follows:
[0078]
[0079] Among them, N is the total number of samples.
[0080] S24, combine the classification loss with the entropy regularization term to construct the total loss function , and train the sample detection model;
[0081] Among them, the total loss function is calculated by combining the cross-entropy loss of the contaminated dataset and the entropy constraint loss of the clean sample set . The specific calculation formula is:
[0082]
[0083] Among them, is the trade-off coefficient, which is used to control the weight of the entropy regularization term in the total loss.
[0084] During the process of training the sample detection model, update the parameters of the sample detection model by the gradient descent method to minimize the total loss function: the parameter update formula is:
[0085]
[0086] Among them, are the parameters of the sample detection model, is the learning rate, is the gradient of the loss function of the sample detection model parameter ;
[0087] The gradient includes the contributions of the classification loss and the entropy regularization term to the parameters:
[0088]
[0089] Among them, is the gradient of the model parameters of the cross-entropy loss function, is the gradient of the model parameters of the entropy constraint loss function;
[0090] The classification loss is mainly used to retain the poisoned sample features in the contaminated dataset, and the gradient is:
[0091]
[0092] Among them, is the input sample The model parameter gradient of the predicted probability for a specified category ;
[0093] The entropy regularization term is used to adjust the predicted distribution of clean samples, increasing the prediction uncertainty of clean samples. The gradient is:
[0094]
[0095] where is the model parameter gradient of the predicted entropy value of the input sample .
[0096] S3. Input the contaminated dataset into the trained sample detection model for prediction to detect poisoned samples.
[0097] As Figure 3 shown, the defender puts all the samples in the contaminated dataset into the trained sample detection model for prediction. Since the sample detection model has forgotten the overall clean sample features but retains the prediction ability for poisoned sample features, the predicted distribution of clean samples in the detection model is relatively average, while the prediction of poisoned samples is very concentrated. In the model output, it is manifested that the predicted entropy value of clean samples is relatively high, while the entropy value of poisoned samples is very low. Such a performance is used as the basis for judging clean samples and poisoned samples, that is, samples with a predicted entropy value exceeding the threshold are clean samples, otherwise they are poisoned samples, and finally the detection of poisoned samples is realized.
[0098] The steps for detecting poisoned samples are as follows:
[0099] S31. Input the contaminated dataset into the trained sample detection model , and generate a set of predicted entropy values for each sample through the feature extraction and prediction mechanism of the sample detection model ; This process can be expressed as:
[0100]
[0101] where are the sample detection model parameters is the calculation function for generating entropy values are the calculation parameters of the entropy values.
[0102] S32. Set a predefined entropy threshold , and judge each entropy value in the entropy value set according to the entropy threshold , and classify the corresponding sample as a poisoned sample or a clean sample. The specific discrimination rules are as follows:
[0103]
[0104]
[0105] That is, when the corresponding sample prediction entropy value is less than the entropy threshold at this time, the sample is judged as a poisoned sample, and when the sample prediction entropy value is greater than or equal to the entropy threshold at this time, the sample is judged as a clean sample.
[0106] Based on a similar inventive concept, an embodiment of the present invention further provides a computer storage medium storing a readable program, which can execute the above-mentioned backdoor defense method based on reverse forgetting when the program runs.
[0107] Based on a similar inventive concept, an embodiment of the present invention provides an electronic device, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete mutual communication through the communication bus;
[0108] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute the operations corresponding to the above-mentioned backdoor defense method based on reverse forgetting.
[0109] Based on a similar inventive concept, an embodiment of the present invention further provides a computer program product, including computer instructions, and the computer instructions instruct a computing device to execute the operations corresponding to the above-mentioned backdoor defense method based on reverse forgetting.
[0110] Embodiment 2
[0111] Based on the backdoor defense method based on reverse forgetting proposed in Embodiment 1, in this embodiment, a backdoor defense system based on reverse forgetting is proposed, including:
[0112] Sample screening module: Use the original clean sample set to train a pre-trained model, and use the pre-trained model to screen out a potential clean sample set from the contaminated data set to synthesize a new clean sample set, and use the new clean sample set to further train the pre-trained model;
[0113] Model training module: Input the contaminated data set and the new clean sample set into the sample detection model, and train them respectively using cross-entropy loss and entropy constraint;
[0114] And, sample detection module: Input the contaminated data set into the trained sample detection model for prediction to detect poisoned samples.
[0115] Embodiment 3
[0116] In this embodiment, a backdoor defense method based on reverse forgetting mentioned in the present invention is verified through a defense experiment;
[0117] In this experiment, two datasets widely used in image classification tasks, namely CIFAR-10 and GTSRB datasets, are used, and the Resnet-18 model is used as the model architecture in the detection model training session. The detailed information of the datasets is shown in Table 1.
[0118] Table 1 Dataset Information Table
[0119]
[0120] In this experiment, three relatively advanced data poisoning backdoor attack methods are selected, including Badnets, Blend, and Wanet. The target label of the attack is defaulted to 0, and the poisoning rate of the backdoor attack samples is set to 5%. The specific parameters and types of these backdoor attack methods are shown in Table 2.
[0121] Table 2 Attack Method Parameter Table
[0122]
[0123] In the first part, the initial number of benign samples is set to 2% - 5% of the dataset sample number, and 10% - 20% of the samples with the highest entropy value are selected for expansion after 5 rounds of pre-training. When performing the second part of re-training, the weight parameter of the entropy constraint term is 3, and the entropy value threshold of the sample is defaulted to 0.5, and the initial learning rate is set to 0.001. The Adam optimizer is selected, and the number of training rounds is set to 50 rounds. The default value is 0.5, and the initial learning rate is set to 0.001. The Adam optimizer is selected, and the number of training rounds is set to 50 rounds.
[0124] Defense effectiveness results:
[0125] In this experiment, the results of two datasets (CIFAR-10, GTSRB) in two backdoor poisoning attacks are evaluated. The evaluation criteria for the effectiveness of backdoor poisoning attacks are the attack success rate (ASR) and the clean data accuracy (ACC). The lower the attack success rate, the better the defense effect; the higher the clean data accuracy, the better the defense method can maintain the classification performance of the original model while filtering out the impact of backdoor attacks. The experimental results are shown in Table 3.
[0126] Table 3 Backdoor Defense Effect Statistical Table
[0127]
[0128] This method has a good defense effect against the above-mentioned backdoor attack methods. When reducing the attack success rate to a very low level, it still maintains a clean accuracy not inferior to that of the original model.
[0129] Meanwhile, in order to prove the reliability and effectiveness of the method in the sample detection model, an entropy value distribution diagram of the sample detection model for detecting contaminated data samples was drawn, and the detection effect is as Figure 4 shown; as can be seen from Figure 4 , this method has a precise detection effect on the dataset attacked by the data poisoning backdoor, and can clearly distinguish poisoned samples from benign samples, improving the security and reliability of the dataset.
[0130] The method of the present invention can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CDROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the method described herein can be stored on such a software process on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.
[0131] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A backdoor defense method based on reverse forgetting, characterized in that: The following steps are involved: The pre-trained model is trained using the original clean sample set, and the pre-trained model is used to screen out a potential clean sample set from the contaminated data set to synthesize a new clean sample set, and the new clean sample set is used to further train the pre-trained model; Inputting the contaminated data set and the new clean sample set into the sample detection model, and training them using cross entropy loss and entropy constraint respectively; Input the contaminated data set into the trained sample detection model for prediction to detect the poisoned samples; The step of synthesizing a new clean sample set includes: Using the original clean sample set Training a pre-trained model ; In each round of training, the contaminated dataset is trained using the pre-trained model. Calculate the loss value for each sample , and screen out a set of potential clean samples; The screened potential clean sample set is compared with the original clean sample set Merge to form a new set of clean samples ; The loss value The calculation formula is: in, is the loss value, is a single category in the classification task, For sample In category The real label on Input for the model Predicted as class The probability value of is the total number of categories in the classification task; The criteria for screening out potential clean sample sets are: in, is a set of potential clean samples, is the preset threshold.
2. According to claim 1, a backdoor defense method based on reverse forgetting is characterized in that: In the process of training the sample detection model, the parameters of the sample detection model are updated by the gradient descent method to minimize the total loss function, wherein the total loss function of the sample detection model is: in, is the total loss function, is the classification loss, is the entropy regularization term, is the trade-off coefficient, N is the total number of samples, For variables The information entropy of For variables The number of all possible outcomes, For input The result is The probability of For sample In category The real label on For the specified category.
3. The backdoor defense method based on reverse forgetting according to claim 1 is characterized in that: The steps to detect poisoned samples are: Input the contaminated data set into the trained sample detection model, and generate a set of predicted entropy values for each sample through the feature extraction and prediction mechanism of the sample detection model. ; Set a predefined entropy threshold , and according to the entropy threshold Entropy value set Each entropy value in Make a judgment, the entropy value exceeds the entropy threshold The sample with is a clean sample, otherwise it is a poisoned sample.
4. A backdoor defense system based on reverse forgetting, characterized in that: include: Sample screening module: Use the original clean sample set to train the pre-trained model, and use the pre-trained model to screen out potential clean sample sets from the contaminated data set to synthesize a new clean sample set, and use the new clean sample set to further train the pre-trained model; Model training module: inputting the contaminated data set and the new clean sample set into the sample detection model, and training them using cross entropy loss and entropy constraint respectively; And, the sample detection module: inputs the contaminated data set into the trained sample detection model for prediction to detect the poisoned samples; The step of synthesizing a new clean sample set includes: Using the original clean sample set Training a pre-trained model ; In each round of training, the contaminated dataset is trained using the pre-trained model. Calculate the loss value for each sample , and screen out a set of potential clean samples; The screened potential clean sample set is compared with the original clean sample set Merge to form a new set of clean samples ; The loss value The calculation formula is: in, is the loss value, is a single category in the classification task, For sample In category The real label on Input for the model Predicted as class The probability value of is the total number of categories in the classification task; The criteria for screening out potential clean sample sets are: in, is a set of potential clean samples, is the preset threshold.
5. A computer storage medium storing a readable program, characterized in that: When the program is running, it can execute the backdoor defense method based on reverse forgetting as described in any one of claims 1 to 3.
6. An electronic device, characterized in that: include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to a backdoor defense method based on reverse forgetting as described in any one of claims 1-3.
7. A computer program product comprising computer instructions, characterized in that The computer instructions instruct the computing device to perform operations corresponding to the backdoor defense method based on reverse forgetting as described in any one of claims 1-3.
Citation Information
Patent Citations
Backdoor attack defense method and system
CN113792289A
Deep learning backdoor attack defense method based on reverse engineering and forgetting
CN116938542A