Back door defense method and system based on reverse forgetting

Through the backdoor defense method based on reverse forgetting, the clean sample set is used to screen and train the model, and the sample detection model is trained in combination with cross-entropy loss and entropy constraints, the problem of difficulty in thoroughly clearing backdoor attacks and impacting model performance in the existing technology is solved, achieving more stable defense effects and higher model security.

CN120030542AActive Publication Date: 2025-05-23NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Application Number
CN202510520160.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-23
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

Existing backdoor defense methods have limitations in the face of more concealed or complex backdoor attacks, making it difficult to completely clear backdoors and may have a negative impact on model performance.

Method used

Using a backdoor defense method based on reverse forgetting, the pre-trained model is trained using the original clean sample set, the potential clean sample set is screened out from the contaminated data set, a new clean sample set is synthesized, and the pre-trained model is further trained. At the same time, the sample detection model is trained using cross entropy loss and entropy constraints to detect toxic samples.

Benefits of technology

This method can more stably defend against updated sample poisoning methods, avoid the limitations of traditional backdoor detection methods, enhance the model's defense ability against multiple data poisoning backdoor attacks, and improve the practicality and security of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030542A_ABST
    Figure CN120030542A_ABST
Patent Text Reader

Abstract

The invention discloses a back door defense method and system based on reverse forgetting, and belongs to the field of artificial intelligence safety. The method comprises the following steps: training a pre-training model by using an original clean sample set, screening a potential clean sample set from a pollution data set by using the pre-training model to synthesize a new clean sample set, and further training the pre-training model by using the new clean sample set; inputting the pollution data set and the new clean sample set into a sample detection model, and training by adopting cross entropy loss and entropy constraint respectively; and inputting the pollution data set into the trained sample detection model for prediction so as to detect a toxic sample. According to the method, reverse forgetting is carried out on the model feature expression of the clean sample, the ontology feature of the backdoor poisoning sample is highlighted, the situation that the feature of the poisoning sample is directly found for judgment is avoided, and therefore the method has a more stable defense effect on an updated sample poisoning method and breaks away from the limitation of a traditional backdoor detection method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence security, and specifically relates to a backdoor defense method and system based on reverse forgetting. Background Art

[0002] As a powerful artificial intelligence technology, deep neural networks (DNNs) have been widely applied in fields such as speech recognition, image classification, and autonomous driving. Their main advantage lies in the ability to generate high-performance models through large-scale data training to solve complex practical problems. However, training an efficient deep neural network model often requires a large amount of computing resources and diverse data support, which makes the model development process costly and poses high requirements for the quality and security of the data. At the same time, deep neural networks also face the risk of backdoor attacks. Such attacks implant trigger patterns in the model, causing the model to perform well under normal inputs but outputting the attacker's preset wrong results when specific conditions are triggered. The concealment and high harmfulness of backdoor attacks pose a serious threat to the security of the system. They may not only lead to the destruction of the model's integrity but also be maliciously exploited, bringing significant risks to data security and application results.

[0003] Regarding backdoor attacks, existing backdoor defense measures mainly focus on two directions: backdoor detection and backdoor elimination. Backdoor detection usually discovers potential backdoor attacks by analyzing the feature differences of contaminated samples or observing the abnormal behaviors of the model; while backdoor elimination relies on methods such as forgetting learning to remove the implanted backdoor features from the model. However, these methods have significant limitations when facing more concealed or complex backdoor attacks. The trigger patterns designed by attackers may bypass existing detection methods, making it difficult to completely remove the backdoors. At the same time, over-relying on the means of forgetting backdoor features may have a negative impact on the normal performance of the model, weakening its robustness and accuracy for uncontaminated samples. Therefore, in the current backdoor attack and defense research, how to maximize the retention of the model's performance during the process of clearing the backdoors remains an important problem to be solved urgently. Summary of the Invention

[0004] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a backdoor defense method and system based on reverse forgetting, which solves the problems in the existing technology.

[0005] The purpose of the present invention can be achieved by the following technical solutions: A backdoor defense method based on reverse forgetting, comprising the following steps: Training a pre-trained model using the original clean sample set, and using the pre-trained model to screen out a potential clean sample set from the contaminated dataset to synthesize a new clean sample set, and further training the pre-trained model using the new clean sample set; Inputting the contaminated data set and the new clean sample set into the sample detection model, and training them using cross entropy loss and entropy constraint respectively; The contaminated data set is input into the trained sample detection model for prediction to detect the poisoned samples.

[0006] Furthermore, the step of synthesizing a new clean sample set includes: Using the original clean sample set Training a pre-trained model ; In each round of training, the contaminated dataset is trained using the pre-trained model. Calculate the loss value for each sample , and screen out a set of potential clean samples; The screened potential clean sample set is compared with the original clean sample set Merge to form a new set of clean samples .

[0007] Furthermore, the loss value The calculation formula is: in, is the loss value, is a single category in the classification task, For sample In category The real label on Input for the model Predicted as class The probability value of is the total number of categories in the classification task.

[0008] Furthermore, the criteria for screening out potential clean sample sets are: in, is a set of potential clean samples, is the preset threshold.

[0009] Furthermore, in the process of training the sample detection model, the parameters of the sample detection model are updated by the gradient descent method to minimize the total loss function, wherein the total loss function of the sample detection model is: in, is the total loss function, is the classification loss, is the entropy regularization term, is the trade-off coefficient, N is the total number of samples, For variables The information entropy of For variables The number of all possible outcomes, For input The result is The probability of For sample In category The real label on .

[0010] Further, the steps of detecting the poisoned sample are: Input the contaminated data set into the trained sample detection model, and generate a set of predicted entropy values ​​for each sample through the feature extraction and prediction mechanism of the sample detection model. ; Set a predefined entropy threshold , and according to the entropy threshold Entropy value set Each entropy value in Make a judgment, the entropy value exceeds the entropy threshold The sample with is a clean sample, otherwise it is a poisoned sample.

[0011] A backdoor defense system based on reverse forgetting, comprising: Sample screening module: Use the original clean sample set to train the pre-trained model, and use the pre-trained model to screen out potential clean sample sets from the contaminated data set to synthesize a new clean sample set, and use the new clean sample set to further train the pre-trained model; Model training module: inputting the contaminated data set and the new clean sample set into the sample detection model, and training them using cross entropy loss and entropy constraint respectively; And, the sample detection module: the contaminated data set is input into the trained sample detection model for prediction to detect the poisoned samples.

[0012] A computer storage medium stores a readable program, which can execute the above-mentioned backdoor defense method based on reverse forgetting when the program is running.

[0013] An electronic device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the above-mentioned backdoor defense method based on reverse forgetting.

[0014] A computer program product includes computer instructions, wherein the computer instructions instruct a computing device to execute operations corresponding to the above-mentioned backdoor defense method based on reverse forgetting.

[0015] Beneficial effects of the present invention: 1. The core of the method of the present invention is to reverse forget the model feature performance of clean samples and highlight the main features of backdoor poisoned samples, rather than directly looking for the features of poisoned samples for judgment. The advantage of this method is that it has a more stable defense effect against newer sample poisoning methods and breaks away from the limitations of traditional backdoor detection methods.

[0016] 2. In the stage of collecting clean samples for the pre-training model, the present invention avoids collecting too many poisoned samples through the strategy of "collecting while training", thereby affecting the effect of subsequent clean sample constraints.

[0017] 3. The present invention uses polluted data sets and clean data sets to train sample detection models. Cross entropy loss is used for normal training of polluted data sets. Entropy constraints are used for clean data sets to constrain the predicted distribution of clean samples, highlighting the low entropy value performance of poisoned samples and realizing the detection of poisoned samples. This detection method greatly enhances the model's defense capabilities against a variety of data poisoning backdoor attacks, such as patch embedded data poisoning (Badnets), hidden data poisoning (Blended), and image distortion data poisoning (Wanet). This poisoned sample detection method based on clean sample entropy constraints not only improves the neural network model's defense capabilities against backdoor attacks, but also enhances the model's practicality and security in real-world scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 is a flow chart of pre-training and screening clean samples of the present invention; Figure 2 It is a schematic diagram of the training sample detection model of the present invention; Figure 3 is a schematic diagram of the present invention for detecting poisoned samples; Figure 4 It is the entropy distribution diagram of the sample detection model detecting contaminated data samples. DETAILED DESCRIPTION

[0020] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0021] Example 1 A backdoor defense method based on reverse forgetting includes the following steps: S1, use the original clean sample set to train the pre-trained model, and use the pre-trained model to screen out potential clean sample sets from the contaminated data set to synthesize a new clean sample set, and use the new clean sample set to further train the pre-trained model; During the training process of the pre-trained model, the Stochastic Gradient Descent (SGD) algorithm is used to adjust the network weight parameters to optimize the performance of the model. The specific training parameters include setting the learning rate to 0.01 and conducting 50 training cycles. This setting helps the model focus more on the features of existing clean samples during the learning process, so that it can effectively distinguish other clean samples from contaminated data sets.

[0022] The samples in the contaminated dataset are then input into the pre-trained model for prediction to obtain a list of sample losses. By sorting these loss values, 50 to 100 samples with the lowest loss in each category are selected and included in the clean dataset. After that, the original pre-trained model is further trained using the new clean dataset to enhance the model's ability to identify clean samples and reduce the risk of misidentifying poisoned samples as clean samples. This screening and retraining process needs to be repeated until the size of the clean dataset reaches 10% to 20% of the original contaminated dataset.

[0023] In this embodiment, the pre-trained model is a CNN model, whose input layer receives image data and then passes it to the convolution layer, extracts local features through convolution operation, and generates feature maps for downsampling in the pooling layer. In some embodiments, a deeper CNN is used, in which multiple groups of convolution-pooling layer combinations are combined to fully extract multi-layer features through multiple rounds of operations, and the feature maps are passed to the flattening layer to be converted into a one-dimensional vector. The fully connected layer maps the nonlinear combination, and the output is classified by generating a category probability distribution through Softmax.

[0024] like Figure 1 As shown, the steps of synthesizing a new set of clean samples include: S11, using a small set of original clean samples Training a pre-trained model ; The goal of the pre-trained model is to extract the basic feature distribution of clean samples. Unlike traditional methods, this method adopts the strategy of "training while recruiting", that is, in each round of iteration of pre-trained model training, potential clean samples are dynamically screened from the contaminated data set, and the clean sample set is gradually expanded. This dynamic screening method can ensure that the screened clean samples are optimized synchronously with the model feature extraction capability, thereby reducing the risk of poisoned samples being mixed in during the screening process.

[0025] S12, in each round of iteration, using the current pre-trained model , for the contaminated data set Calculate the loss value for each sample , and screen out a set of potential clean samples; Training a pre-trained model In the process, the cross entropy loss function is used to optimize the model parameters: in, is the loss value, For sample In category The real label on Input for the model Predicted as class The probability value of is a single category in the classification task, is the total number of categories in the classification task.

[0026] In each round of training, for the contaminated dataset, the loss value of each sample is calculated using the current model, and the loss value is lower than the preset threshold. The samples are temporarily marked as potential clean samples. Subsequently, these potential clean samples are added to the clean sample set to participate in the next round of training, thereby continuously improving the model's ability to extract clean sample features.

[0027] The criteria for screening out potential clean sample sets are: in, is a set of potential clean samples, is the preset threshold.

[0028] S13, the screened potential clean sample collection With the original clean sample set Merge to form a new set of clean samples : .

[0029] S2, input the contaminated dataset and the new clean sample set synthesized in S1 into the sample detection model, and train them using cross entropy loss and entropy constraint respectively; like Figure 2 As shown in Figure 1, the defender puts the contaminated data set obtained at the beginning into the sample detection model for training. In this training, the cross entropy loss is used for optimization. The purpose is to let the sample detection model learn the overall data features in the contaminated data set. This part of the data features includes the features of clean samples and the features of poisoned samples. At the same time, the defender puts the new clean sample set synthesized by S1 into the sample detection model for entropy constraint training. This step is to make the detection model forget the clean sample features in the contaminated data set, thereby highlighting the characteristic performance of the poisoned samples in the model.

[0030] In this embodiment, the sample detection model is a ResNet-18 model, which consists of six modules to form a complete feature learning system, including: input layer, convolution layer, maximum pooling layer, residual block module, average pooling layer and fully connected layer. The input layer receives the standardized image data and then passes the data to the 7×7 convolution layer (Conv1). The convolution layer performs convolution operation with a step size of 2 to extract basic features such as edges and textures and reduce the spatial dimension, and then transmits the processed feature map to the maximum pooling layer. The maximum pooling layer further compresses the size of the feature map, enhances the translation invariance, and then outputs it to the core residual block module. The residual block module constructs a feature pyramid of 4 stages through jump connections. The first residual block in each stage uses 1×1 convolution to adjust the channel dimension, and the subsequent residual blocks use 3×3 convolution to capture the middle-level semantic features. After completing the feature processing, the deep feature map is passed to the global average pooling layer. The global average pooling layer compresses the deep feature map into a one-dimensional vector and then transmits this vector to the fully connected layer. The fully connected layer generates the category probability distribution based on the vector, and finally completes the final classification task through the Softmax activation function.

[0031] In the process of training the sample detection model, a total loss function is constructed to train the sample detection model; the construction process of the total loss function is as follows: S21, calculate each input sample The entropy of the predicted probability distribution , to measure the uncertainty of the model predictions; Among them, entropy Defined as: in, For variables The number of all possible outcomes, For input The result is probability; S22, by calculating the new clean sample set The average entropy of all samples above is taken and the negative number is taken to construct the entropy regularization term , in order to maximize entropy in the loss function; Entropy regularization term The calculation formula is: S23, using cross entropy loss to define classification loss , to measure the performance of the model on the classification task; classification loss The calculation formula is as follows: in, N is the total number of samples.

[0032] S24, the classification loss and the entropy regularization term Combined to construct the total loss function , training sample detection model; Among them, the total loss function Cross entropy loss from contaminated dataset and the entropy-constrained loss of the clean sample set Combined calculation, the specific calculation formula is: in, is a trade-off coefficient that controls the weight of the entropy regularization term in the total loss.

[0033] In the process of training the sample detection model, the parameters of the sample detection model are updated by the gradient descent method to minimize the total loss function: the parameter update formula is: in, are the sample detection model parameters, is the learning rate, Detect model parameters for samples The gradient of the loss function; gradient Contains the contribution of classification loss and entropy regularization term to the parameters: in, is the model parameter gradient of the cross entropy loss function, is the model parameter gradient of the entropy constrained loss function; The classification loss is mainly used to retain the features of the poisoned samples in the contaminated data set, and the gradient is: in, For input sample For the specified category The model parameter gradient of the predicted probability; The entropy regularization term is used to adjust the predicted distribution of clean samples and increase the prediction uncertainty of clean samples. The gradient is: in, For input sample The model parameter gradient of the predicted entropy value.

[0034] S3, input the contaminated data set into the trained sample detection model for prediction to detect the poisoned samples.

[0035] like Figure 3 As shown in the figure, the defender puts all the samples in the contaminated data set into the trained sample detection model for prediction. Since the sample detection model forgets the overall clean sample characteristics but retains the prediction ability for the poisoned sample characteristics, the prediction distribution of clean samples in the detection model is relatively even, while the prediction of poisoned samples is very concentrated. In the model output, the prediction entropy value of clean samples is high, while the entropy value of poisoned samples is very low. Such performance is used as the basis for judging clean samples and poisoned samples, that is, samples with predicted entropy values ​​exceeding the threshold are clean samples, otherwise they are poisoned samples, and finally the detection of poisoned samples is achieved.

[0036] The steps to detect poisoned samples are: S31, will pollute the data set Input the trained sample detection model , through the feature extraction and prediction mechanism of the sample detection model, a set of predicted entropy values ​​for each sample is generated ; This process can be expressed as: in, are the sample detection model parameters, is the calculation function that generates the entropy value, It is the calculation parameter of entropy value.

[0037] S32, setting a predefined entropy threshold , and according to the entropy threshold Entropy value set Each entropy value in Make a judgment and convert the corresponding samples Classify as poisoned samples or clean samples. The specific judgment rules are as follows: That is, when the corresponding sample The predicted entropy value of Less than the entropy threshold When , the sample is identified as a poisoned sample, and when the sample The predicted entropy value of Greater than or equal to the entropy threshold , the sample is identified as a clean sample.

[0038] Based on similar inventive concepts, an embodiment of the present invention further provides a computer storage medium storing a readable program, which can execute the above-mentioned backdoor defense method based on reverse forgetting when the program is running.

[0039] Based on similar inventive concepts, an embodiment of the present invention provides an electronic device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the above-mentioned backdoor defense method based on reverse forgetting.

[0040] Based on similar inventive concepts, an embodiment of the present invention further provides a computer program product, including computer instructions, wherein the computer instructions instruct a computing device to execute operations corresponding to the above-mentioned backdoor defense method based on reverse forgetting.

[0041] Example 2 Based on the backdoor defense method based on reverse forgetting proposed in Example 1, in this embodiment, a backdoor defense system based on reverse forgetting is proposed, including: Sample screening module: Use the original clean sample set to train the pre-trained model, and use the pre-trained model to screen out potential clean sample sets from the contaminated data set to synthesize a new clean sample set, and use the new clean sample set to further train the pre-trained model; Model training module: inputting the contaminated data set and the new clean sample set into the sample detection model, and training them using cross entropy loss and entropy constraint respectively; And, the sample detection module: the contaminated data set is input into the trained sample detection model for prediction to detect the poisoned samples.

[0042] Example 3 In this embodiment, a defense experiment is performed to verify a backdoor defense method based on reverse forgetting mentioned in the present invention; This experiment used two datasets that are widely used in image classification tasks: CIFAR-10 and GTSRB datasets. The Resnet-18 model was used as the model architecture in the detection model training phase. For detailed information on the datasets, see Table 1.

[0043] Table 1 Dataset information table This experiment selected three relatively advanced data poisoning backdoor attack methods, including Badnets, Blend, and Wanet. The target label of the attack defaults to 0, and the sample poisoning rate of the backdoor attack is set to 5%. The specific parameters and types of these backdoor attack methods are shown in Table 2.

[0044] Table 2 Attack method parameter table In the first part, the number of initial benign samples is set to 2%-5% of the number of samples in the data set. After 5 rounds of pre-training, the 10% to 20% samples with the highest entropy values ​​are selected for expansion. When retraining in the second part, the weight parameter of the entropy constraint item is 3, and the entropy value threshold of the sample is judged. The default value is 0.5, the initial learning rate is set to 0.001, the optimizer is Adam optimizer, and the number of training rounds is set to 50.

[0045] Defense Effectiveness Results: This experiment evaluates the results of two datasets (CIFAR-10, GTSRB) on two backdoor poisoning attacks. The evaluation criteria for the effectiveness of backdoor poisoning attacks are attack success rate (ASR) and clean data accuracy (ACC). The lower the attack success rate, the better the defense effect; the higher the clean data accuracy, the defense method can filter out the impact of backdoor attacks while maintaining the classification performance of the original model. The experimental results are shown in Table 3.

[0046] Table 3 Backdoor defense effect statistics This method has a good defensive effect on the above-mentioned backdoor attack methods. While reducing the attack success rate to a very low level, it still maintains a clean accuracy rate that is not inferior to the original model.

[0047] At the same time, in order to prove the reliability and effectiveness of this method in the sample detection model, the entropy distribution diagram of the sample detection model detecting contaminated data samples was drawn. Figure 4 shown; from Figure 4 It can be seen that this method has an accurate detection effect on the data set attacked by data poisoning backdoor, and can clearly distinguish the poisoned samples from the benign samples, thus improving the security and reliability of the data set.

[0048] The method of the present invention may be implemented in hardware, firmware, or as software or computer code that may be stored in a recording medium (such as a CDROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded over a network and will be stored in a local recording medium, so that the method described herein may be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that a computer, processor, microprocessor controller, or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, processor, or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.

[0049] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements all fall within the scope of the present invention to be protected.

Claims

1. A backdoor defense method based on reverse forgetting, characterized in that: The following steps are involved: The pre-trained model is trained using the original clean sample set, and the pre-trained model is used to screen out a potential clean sample set from the contaminated data set to synthesize a new clean sample set, and the new clean sample set is used to further train the pre-trained model; Inputting the contaminated data set and the new clean sample set into the sample detection model, and training them using cross entropy loss and entropy constraint respectively; The contaminated data set is input into the trained sample detection model for prediction to detect the poisoned samples.

2. According to claim 1, a backdoor defense method based on reverse forgetting is characterized in that: The step of synthesizing a new clean sample set includes: Using the original clean sample set Training a pre-trained model ; In each round of training, the contaminated dataset is trained using the pre-trained model. Calculate the loss value for each sample , and screen out a set of potential clean samples; The screened potential clean sample set is compared with the original clean sample set Merge to form a new set of clean samples .

3. A backdoor defense method based on reverse forgetting according to claim 2, characterized in that: The loss value The calculation formula is: in, is the loss value, is a single category in the classification task, For sample In category The real label on Input for the model Predicted as class The probability value of is the total number of categories in the classification task.

4. The backdoor defense method based on reverse forgetting according to claim 3 is characterized in that: The criteria for screening out potential clean sample sets are: in, is a set of potential clean samples, is the preset threshold.

5. The backdoor defense method based on reverse forgetting according to claim 4 is characterized in that: In the process of training the sample detection model, the parameters of the sample detection model are updated by the gradient descent method to minimize the total loss function, wherein the total loss function of the sample detection model is: in, is the total loss function, is the classification loss, is the entropy regularization term, is the trade-off coefficient, N is the total number of samples, For variables The information entropy of For variables The number of all possible outcomes, For input The result is The probability of For sample In category The real label on For the specified category.

6. The backdoor defense method based on reverse forgetting according to claim 1 is characterized in that: The steps to detect poisoned samples are: Input the contaminated data set into the trained sample detection model, and generate a set of predicted entropy values ​​for each sample through the feature extraction and prediction mechanism of the sample detection model. ; Set a predefined entropy threshold , and according to the entropy threshold Entropy value set Each entropy value in Make a judgment, the entropy value exceeds the entropy threshold The sample with is a clean sample, otherwise it is a poisoned sample.

7. A backdoor defense system based on reverse forgetting, characterized in that: include: Sample screening module: Use the original clean sample set to train the pre-trained model, and use the pre-trained model to screen out potential clean sample sets from the contaminated data set to synthesize a new clean sample set, and use the new clean sample set to further train the pre-trained model; Model training module: inputting the contaminated data set and the new clean sample set into the sample detection model, and training them using cross entropy loss and entropy constraint respectively; And, the sample detection module: the contaminated data set is input into the trained sample detection model for prediction to detect the poisoned samples.

8. A computer storage medium storing a readable program, characterized in that: When the program is running, it can execute the backdoor defense method based on reverse forgetting as described in any one of claims 1 to 6.

9. An electronic device, characterized in that: include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to a backdoor defense method based on reverse forgetting as described in any one of claims 1-6.

10. A computer program product comprising computer instructions, characterized in that: The computer instructions instruct the computing device to perform operations corresponding to the backdoor defense method based on reverse forgetting as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Backdoor attack defense method and system

    CN113792289A

  • Security defense system and method facing power system intelligent model backdoor attack

    CN116108434A

  • Deep learning backdoor defense method based on neural network capacity

    CN116226663A

  • Deep learning backdoor attack defense method based on reverse engineering and forgetting

    CN116938542A

  • Federal learning backdoor attack-oriented defense method

    CN118036770A

Cited By

  • Graph classification model backdoor attack method for enhancing concealment

    CN120356016A

  • AI model backdoor defense system and detection method thereof

    CN122093173A