Method and device for detecting and repairing model backdoor based on competitive game

By implanting active defense backdoors in the training stage of deep neural network model and using the competitive game relationship between backdoors for detection and repair, the problems of long detection time, low detection accuracy and backdoor removal in the existing technology affecting model performance are solved, and efficient and accurate backdoor detection and repair are achieved.

CN120163200AActive Publication Date: 2025-06-17COMP APPL RES INST CHINA ACAD OF ENG PHYSICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510259477.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-17
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The prior art has problems in backdoor attack detection and defense with long detection time, low accuracy of detection of adaptive backdoor attacks, and severely weaken model performance when removing backdoors.

Method used

By implanting active defense backdoors in the model training stage, the competitive game relationship between backdoors is used for detection and repair. The specific steps include artificially constructing strong and weak backdoor triggers in the training data, so that the model can learn the competitive relationship between the backdoors, and then implanting the weak backdoor triggers into suspicious data during the detection stage, using the model output to determine whether the sample is poisoned, and cutting off the connection between the toxic sample and its label through machine forgetting technology to remove the backdoor.

Benefits of technology

It effectively reduces the success rate of backdoor attacks, ensures that the model performance is not affected too much, and improves the detection accuracy and detection efficiency of adaptive backdoor attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163200A_ABST
    Figure CN120163200A_ABST
Patent Text Reader

Abstract

The invention relates to the field of deep neural network security, and provides a model backdoor detection and repair method and device based on a competitive game. According to the method, strong and weak backdoor triggers are implanted into clean data, and a model is trained in combination with suspicious data, so that the model learns a competitive relationship between the two. During detection, a weak trigger is implanted into a suspicious sample, and whether the sample is poisoned or not is judged according to model output. And processing the detected toxic sample by adopting a machine forgetting technology, and retraining the model to remove the back door. The technology effectively solves the problems that an existing method is low in detection efficiency and poor in accuracy, and the performance is reduced when the back door is removed, and improves the safety and performance of the neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep neural networks, and provides a method and device for detecting and repairing model backdoors based on competitive games. Background Art

[0002] In recent years, the field of deep neural networks (DNNs) has developed rapidly and achieved remarkable results in different scenarios such as face recognition, speech recognition, and autonomous driving. A neural network essentially simulates the mapping relationship between input and output, which makes the training of deep neural networks rely on a large amount of data. However, in actual scenarios, trainers usually only have limited data and need to obtain additional data from other sources such as the Internet. These external data may be maliciously tampered with, resulting in serious security risks in the models trained with this data.

[0003] Attacks against the security of deep neural networks can be divided into adversarial attacks and backdoor attacks. Adversarial attacks add some imperceptible perturbations to the input data, causing the model to make misjudgments. This type of attack is immediate. The currently more popular attack method is the backdoor attack, which means that an attacker poisons the training set or modifies the network model parameters to implant a backdoor mechanism into the model, thereby obtaining a backdoor model. When an input sample with a backdoor trigger is received, the model will output the result as expected by the attacker, while when a normal sample is input, the model will output a normal result. Compared with adversarial attacks, backdoor attacks are more concealed and pose a greater threat to network security.

[0004] A backdoor trigger is a tool for activating the model backdoor, which is embedded in the input data in a specific pattern or dynamic pattern. When the backdoor model recognizes the trigger, the model will output an abnormal result according to the preset logic. Anh Tuan Nguyen et al.'s backdoor attack based on image distortion makes the backdoor attack more concealed.

[0005] Therefore, resisting backdoor attacks has become an important research content in the field of deep neural network security, namely backdoor defense. There are two types of defense methods in the field of backdoor defense: search-based backdoor detection and defense methods, and model pruning-based defense methods. The former searches for potential backdoor triggers in the model and eliminates the triggers through methods such as fine-tuning to achieve the purpose of defending against backdoor attacks; while the latter mainly targets backdoor models trained with poisoned samples, and these methods remove the backdoors by pruning or distilling the backdoor models. However, the above two types of backdoor defense technologies have problems such as large time overhead in the detection link, low detection accuracy for adaptive backdoor attacks, and serious weakening of model performance when removing the backdoors, as follows: Search-based backdoor detection and defense methods: This type of method mainly relies on detecting abnormal behaviors or specific patterns in the model. During the search process, multiple assumptions are often required, which will repeatedly load the model to evaluate different assumptions. As the data increases, the size of the search space will grow exponentially, which will cause huge time overhead during the execution of the method. Search-based backdoor detection methods usually assume that the trigger is a fixed pattern. However, adaptive backdoor attacks use a dynamic and diversified strategy, which makes the implanted backdoor trigger have the characteristics of dynamic changes. This makes the search-based backdoor detection and defense methods unable to effectively detect such attacks.

[0006] Defense method based on model pruning: In the backdoor model, the backdoor behavior and normal behavior may share the same neuron. The backdoor behavior cannot be accurately processed during the pruning or distillation process. In order to ensure that the backdoor is completely removed, pruning or distillation technology will greatly reduce the model's parameters, significantly weakening the model's generalization ability and causing serious damage to the model's performance.

[0007] In view of the limitations of the above-mentioned backdoor defense technology, the present invention proposes a model backdoor detection and repair technology based on competitive game. Based on the competitive game relationship between model backdoors, by implanting active defense backdoors in the model training stage, the detection of potential poisoned samples is realized in the model training stage. On this basis, the output corresponding to the poisoned sample is changed to the second most likely category. While cutting off the connection between the poisoned sample and the attacker's target category, the connection between the poisoned sample and the corresponding correct label before being poisoned is reconstructed to a certain extent, thereby reducing the success rate of backdoor attacks and ensuring that the model performance is not affected too much. The specific functions of the present invention are as follows: During model training, active defense backdoors are implanted to enable the model to learn the competitive game relationship between backdoors. During the testing phase, weak backdoors are implanted into suspicious data sets to determine whether the sample is poisoned based on the results of the competitive game between backdoors. No assumptions are required and the model only needs to be loaded once, effectively reducing the detection time. For adaptive attack methods such as dynamic triggers, the competitive game relationship can adaptively identify poisoned samples, thereby improving the detection accuracy of adaptive backdoor attacks.

[0008] The present invention uses the competitive game relationship between backdoors to screen out toxic samples and cut off the connection between toxic samples and their corresponding labels. Compared with pruning or distillation technology, the present invention can retain neurons related to normal behavior and make targeted modifications to neurons related to backdoor behavior. In this way, the backdoor can be effectively removed while ensuring the performance of the model.

[0009] Neural Cleanse Technique (Prior Art): The main idea of this technique is to detect whether there is a backdoor in a given model based on search. The theoretical basis is that when perturbations are added to samples to cause the samples to output incorrect labels, compared with clean samples, the perturbations required for samples with attack triggers are usually smaller. Based on this characteristic, this technique realizes the detection of poisoned samples through the following steps: Assume that a suspicious deep neural network model has been obtained; Assume each category covered by the model output in turn as the attacker's target category (i.e., the target category of the backdoor attack); For each assumed target category, calculate the minimum perturbation required to misclassify samples in other categories as this target category; After completing the assumptions for all categories, compare the obtained minimum perturbation values, and the assumed target category with the smallest perturbation value is judged as the true attack category. This detection method will perform similar operations on the same data multiple times, resulting in too high time costs.

[0010] Fine-Pruning Technique (Prior Art): The main function of this technique is to remove the backdoor in the model. This technique utilizes the fact that when clean data and poisoned data are input into the backdoor model, the activation degrees of neurons in the model are different, and thus the neurons closely related to the poisoned data are pruned. The specific process is as follows: Input the clean data set and the suspicious data set into the model respectively to observe whether the neurons are activated, and prune the neurons that are not activated; Input the two data sets into the model again and prune the neurons that are only activated by the suspicious data set; Input the clean data set, and prune in the order of the average activation value of the neurons from low to high until the accuracy of the model on the clean validation set is lower than the threshold and then stop; Finally, perform a fine-tuning operation on the model to obtain a model with the backdoor removed. However, this method often cannot achieve both high model performance and low backdoor success rate.

[0011] The above detection method and backdoor removal method have certain effects in the field of post-defense, but there are also obvious limitations, which are as follows: Regarding the problems of low efficiency in the detection link and low detection accuracy for adaptive backdoor attacks: Since this method requires multiple assumptions and repeatedly loads the model to evaluate different assumptions, when the scale of the data set increases, the time cost will increase exponentially. And a search strategy based on a fixed pattern cannot effectively detect the dynamic triggers of adaptive attackers, resulting in low detection accuracy of this method for adaptive backdoor attacks.

[0012] The problem that the model performance will be severely weakened while removing the backdoor: This method overly relies on pruning techniques. However, there are currently backdoor attacks against pruning, making pruning ineffective. And there is often a neuron in the backdoor model that is related to both normal behavior and backdoor behavior, resulting in the pruning of these neurons related to normal behavior during the pruning process when trying to completely remove the backdoor, severely damaging the model performance.

[0013] In view of the limitations of the above existing methods, the present invention proposes a model backdoor detection and repair technology based on competitive game. Based on the competitive game relationship between model backdoors, by implanting an active defense backdoor during the model training stage, the detection of potential poisoning samples is realized during the model training stage. Further, the output corresponding to the poisoned sample is changed to the second-highest possible category, while cutting off the connection between the poisoned sample and the attacker's target category, and to a certain extent reconstructing the connection between the poisoned sample and the corresponding correct label before being poisoned, thereby reducing the success rate of the backdoor attack and ensuring that the model performance is not overly affected. Summary of the Invention

[0014] The purpose of the present invention is to solve the technical problems existing in the existing model backdoor detection and repair technologies, such as long detection time, low detection accuracy for adaptive backdoor attacks, and severely weakening the model performance when removing the backdoor. By introducing competitive game and machine forgetting technologies, the present invention aims to provide an efficient and accurate solution to address these challenges and ensure the security and performance of deep neural networks.

[0015] To achieve the above purpose, the present invention adopts the following technical means: The present invention provides a model backdoor detection and repair device based on competitive game, including: A learning module, before model training, artificially constructs backdoors with different trigger intensities (strong backdoors and weak backdoors), and their corresponding backdoor triggers are respectively called strong backdoor triggers and weak backdoor triggers. The strong and weak triggers are randomly implanted into the samples in the clean dataset according to the implantation strategy. That is, when both the strong and weak backdoor triggers are implanted into a sample, the sample label is modified to the strong backdoor label, artificially constructing a competitive game relationship; when a single strong or weak backdoor trigger is randomly implanted into a sample, the sample label is modified to the corresponding backdoor label, ensuring that the model successfully learns the artificially implanted backdoor. And the clean dataset is merged with an external suspicious dataset to jointly train the model, so that the model learns the competitive game relationship between the strong backdoor and the weak backdoor when the backdoor is triggered. This competitive game relationship is effective for the same type of backdoors.

[0016] A backdoor detection module, used to implant the weak backdoor trigger into each sample in the suspicious dataset, and input the sample implanted with the weak backdoor trigger into the trained model. Using the mechanism that the competitive game relationship constructed by the learning module is effective for the same type of backdoors, it is judged whether the model has the same type of backdoor according to the output of the model.

[0017] The backdoor removal module is used to regard the toxic samples detected by the backdoor detection module as abnormal samples, replace the corresponding labels of the abnormal samples with the second-highest possible categories, cut off the connection between the abnormal samples and their corresponding labels, and to a certain extent reconstruct the connection between the toxic sample and its corresponding correct label before being poisoned. Then, the processed toxic samples and clean samples are used to retrain the model to remove the backdoor in the model.

[0018] In the above solution, the learning module includes: Randomly divide the clean dataset into three parts. The first part of the data is embedded with a strong backdoor trigger, the second part of the data is embedded with a weak backdoor trigger, and the third part of the data is embedded with both a strong backdoor trigger and a weak backdoor trigger; The labels corresponding to the strong backdoor trigger and the weak backdoor trigger are different, and when a sample contains both a strong backdoor trigger and a weak backdoor trigger, the label of the sample is the label corresponding to the strong backdoor trigger; Merge the clean dataset embedded with backdoor triggers and the suspicious dataset, and jointly input them into the model for training to obtain a trained model M.

[0019] In the above solution, the backdoor detection module includes: Implant the weak backdoor trigger into each sample in the suspicious dataset; Input the samples implanted with the weak backdoor trigger into the trained model M; Judge whether the sample is poisoned according to the output of the model M. Among them, if the output of the model is the label corresponding to the weak backdoor trigger, it is considered that the sample is not poisoned, otherwise it is considered that the sample is poisoned.

[0020] In the above solution, the backdoor removal module includes: Perform machine forgetting processing on the toxic samples detected by the backdoor detection module to cut off the connection between the toxic samples and their labels; Combine the clean samples and the toxic samples processed by machine forgetting to form a new training set; Retrain the model M to eliminate the backdoor influence caused by the toxic samples on the model and obtain a disinfected benign model.

[0021] The present invention also provides a method for detecting and repairing model backdoors based on competitive games, including the following steps: Step a. Artificially construct two backdoors with different triggering intensities (strong backdoor and weak backdoor), and their corresponding backdoor triggers are respectively called strong backdoor trigger and weak backdoor trigger. Randomly implant the strong and weak types of triggers into the samples in the clean dataset according to the implantation strategy. That is, when the strong and weak backdoor triggers are implanted into a sample at the same time, the sample label is the strong backdoor label, artificially constructing a competitive game relationship; when a single strong backdoor or weak backdoor trigger is randomly implanted into a sample, the sample label is modified to the corresponding backdoor label to ensure that the model successfully learns the artificially implanted backdoor. And merge the clean dataset with the external suspicious dataset to jointly train the model, so that the model learns the competitive game relationship between the strong backdoor and the weak backdoor when triggering the backdoor, and this competitive game relationship is effective for the same type of backdoors; Step b. Implant the weak backdoor trigger into each sample in the suspicious dataset, and input the sample implanted with the weak backdoor trigger into the trained model. Utilize the mechanism that the competitive game relationship constructed by the learning module is effective for the same type of backdoors, and judge whether the model has the same type of backdoor according to the output of the model; Step c. Regard the detected poisoned samples as abnormal samples for machine forgetting processing, cut off the connection between the abnormal samples and the corresponding labels, and retrain the model with the processed poisoned samples and clean samples to remove the backdoor in the model.

[0022] In the above solution, step a includes: Randomly divide the clean dataset into three parts, where the first part of the data is embedded with the strong backdoor trigger, the second part of the data is embedded with the weak backdoor trigger, and the third part of the data is embedded with both the strong backdoor trigger and the weak backdoor trigger; Ensure that the labels corresponding to the strong backdoor trigger and the weak backdoor trigger are different, and when the sample contains both the strong backdoor trigger and the weak backdoor trigger at the same time, the label of the sample is the label corresponding to the strong backdoor trigger; Merge the clean dataset embedded with the backdoor trigger with the suspicious dataset, and jointly input them into the model for training to obtain the trained model M.

[0023] In the above solution, step b includes: Implant the weak backdoor trigger into each sample in the suspicious dataset; Input the sample implanted with the weak backdoor trigger into the trained model M; Judge whether the sample is poisoned according to the output of the model M. Among them, if the output of the model is the label corresponding to the weak backdoor trigger, it is considered that the sample is not poisoned, otherwise it is considered that the sample is poisoned.

[0024] In the above solution, step c includes: Perform machine forgetting processing on the detected toxic samples to cut off the connection between the toxic samples and their labels; Combine the clean samples with the toxic samples that have undergone machine forgetting processing to form a new training set; Retrain the model M to eliminate the backdoor impact of the toxic samples on the model and obtain a disinfected benign model.

[0025] The present invention also provides a storage medium. When a processor executes a program in the storage medium, it implements the described method for detecting and repairing model backdoors based on competitive games.

[0026] Since the present invention adopts the above technical means, it has the following beneficial effects: First, existing search-based backdoor detection methods have problems of large time overhead and low detection accuracy for adaptive backdoor attacks. The present invention adopts the technical means of implanting active defense backdoors in the training stage and using the competitive game relationship between backdoors for detection. Through competitive games, the number of model loadings is reduced, thereby reducing the detection time. At the same time, competitive games can adaptively identify toxic samples and improve detection accuracy. Therefore, this solution effectively solves the problems of large time overhead and low detection accuracy, achieving the effects of reducing the detection time and improving the detection accuracy.

[0027] Second, existing pruning-based backdoor removal methods have the technical problem of severely weakening the model performance when removing backdoors. The present invention uses machine forgetting technology to specifically cut off the connection between toxic samples and the attack target class. Machine forgetting technology can accurately identify and process toxic samples, avoiding mispruning neurons related to normal behaviors, thereby maintaining the model performance while removing backdoors. Therefore, this solution solves the problem of pruning methods affecting the model performance and achieves the effect of effectively removing backdoors while maintaining the model performance.

[0028] Third, the comprehensive problems of low detection efficiency and backdoor removal affecting model performance in the prior art. The present invention combines competitive games and machine forgetting technology to achieve efficient and accurate backdoor detection and repair. By comprehensively applying competitive games and machine forgetting technology, the solution performs excellently in improving detection efficiency and maintaining model performance, thereby enhancing the comprehensive effects of model security and maintaining high performance.

[0029] In summary, this technical solution solves multiple problems of traditional technologies through innovative methods, realizes efficient and secure backdoor detection and repair, and provides a strong guarantee for the security of deep neural networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a flowchart; Figure 2 It is the structure diagram of the learning module; Figure 3 It is the structure diagram of the backdoor detection module; Figure 4 It is the structure diagram of the backdoor removal module. Specific implementation manners

[0031] The embodiments of the present invention will be described in detail below. Although the present invention will be described and explained in conjunction with some specific implementation manners, it should be noted that the present invention is not limited to these implementation manners only. On the contrary, any modifications or equivalent replacements made to the present invention shall be covered within the scope of the claims of the present invention.

[0032] In addition, for better illustration of the present invention, numerous specific details are given in the following specific implementation manners. Those skilled in the art will understand that the present invention can also be implemented without these specific details.

[0033] Introduction to system function modules: The model backdoor detection and repair technology based on competitive game is mainly divided into a learning module, a backdoor detection module, and a backdoor removal module. Among them, the learning module enables the model to learn the competitive game relationship between backdoors; in the backdoor detection module, this relationship will be used to detect suspicious data; the backdoor removal module will perform machine forgetting processing on the poisoned samples detected in the backdoor detection module, so as to remove the backdoors in the model, and the process is as Figure 1 shown.

[0034] System design details: According to the above function modules, this section will introduce the implementation steps of the learning module, the backdoor detection module, and the backdoor removal module in detail: 1. Learning module The main function of the learning module is to let the model learn the competitive game relationship between backdoors by artificially implanting "strong and weak" backdoor triggers into the clean data set. The specific implementation details are as Figure 2 shown: Randomly divide the clean data set into three parts. One part of the data is embedded with strong backdoor triggers; one part of the data is embedded with weak backdoor triggers; the remaining part of the samples is embedded with both strong backdoor triggers and weak backdoor triggers. The data in the artificially poisoned data set satisfies the following conditions: The labels corresponding to the strong and weak backdoors are different When a sample contains both a strong backdoor trigger and a weak backdoor trigger, the sample label is the corresponding label of the strong backdoor trigger. Then, the artificially poisoned data set and the suspicious data set are merged and jointly input into the model for training to obtain a trained model M. Among them, the strong backdoor trigger is denoted as B, and the weak backdoor trigger is denoted as b.

[0035] 2. Backdoor detection module The main function of the backdoor detection module is to detect suspicious data through the competitive game relationship between the backdoors learned by the model. The specific implementation process is as follows Figure 3 shown: implant the weak backdoor trigger b into each sample in the suspicious dataset, input it into the trained model M, and judge whether the sample is poisoned according to whether the output of each sample is the label corresponding to the weak backdoor trigger. Specifically, when the output label is the label corresponding to the weak backdoor trigger, it is considered that the sample is not poisoned, that is, a benign sample; otherwise, the sample is considered a poisoned sample.

[0036] 3. Backdoor removal module The backdoor removal module uses machine forgetting technology to cut off the connection between the poisoned sample and its label. The specific implementation process is as follows Figure 4 shown: perform machine forgetting processing on the poisoned samples, form a new training set with the clean samples and the processed poisoned samples, and retrain the model M, so as to eliminate the backdoor impact caused by the poisoned samples on the model and obtain a disinfected benign model.

[0037] In summary, the present invention has the following characteristics: Aiming at the problems of low efficiency in the detection link and low detection accuracy for adaptive backdoor attacks, the present invention enables the model to learn the competitive game relationship between backdoors, implants weak backdoors into the suspicious dataset, and thus detects whether there are other backdoors (attack backdoors) in the model and whether each sample in the suspicious dataset is poisoned according to whether the model output is the corresponding label of the known weak backdoor. Compared with the current search-based backdoor detection methods, the present invention has a small time cost and can effectively defend against adaptive backdoor attacks.

[0038] Aiming at the problem that the existing methods will weaken the model performance when removing backdoors, the present invention uses machine forgetting technology to specifically cut off the connection between the poisoned samples and the attacker's target classes in the model, so as to effectively reduce the success rate of backdoor attacks on the basis of ensuring that the performance of the model on the clean dataset is not affected. Compared with the current pruning-based backdoor removal methods, the present invention removes the backdoors more thoroughly and specifically, and can reduce the success rate of backdoor attacks to a very low level with almost no impact on the model performance.

[0039] Example 1 In this experiment, the CIFAR-10 dataset was used for the experiment. Each sample is an RGB image of 32*32 pixels. This is a 10-classification dataset, and the label values are distinguished according to 0-9, which are airplane, car, bird, cat, deer, dog, frog, horse, ship, and truck respectively. The attack method used is BadNets, and the trigger mode of this attack method is to change a certain rectangular area of the picture into something like "black and white checkerboard".

[0040] Experimental hypothesis: The defender has a clean dataset and a suspicious dataset. Take a part of the data from the CIFAR-10 dataset to form the dataset data_temp, and obtain the suspicious dataset through further processing. Then take a part of the remaining data from the CIFAR-10 dataset to form the dataset data_clean as the clean dataset.

[0041] Learning module Composition of the suspicious dataset: Change the pixel values of the upper left 3*3 area of each sample image in data_temp to 0, that is, this area is black, and change the corresponding label to any same category, assumed to be 0. Combine the modified data with the CIFAR-10 training set to form the suspicious dataset data_attck.

[0042] Processing of the clean dataset: Process each image in the clean dataset with the following strategy: Step 1: Randomly divide the dataset into 3 parts.

[0043] Step 2: Change the pixel values of the lower right 3*3 area of each image in the first part of the data to 1, that is, this area is white, and change the corresponding label to any same category, assumed to be 1.

[0044] Step 3: Change the pixel values of the middle 3*3 area of each image in the second part of the data to 0, that is, this area is black, and change the corresponding label to any same category, assumed to be 2.

[0045] Step 4: Change the pixel values of the above-mentioned lower right 3*3 area of each image in the third part of the data to 1, change the pixel values of the above-mentioned middle 3*3 area to 0, and change the corresponding label of the image to 1.

[0046] Steps 2 and 3 above ensure that the model can learn the backdoors constructed by the defender, and step 4 is for the model to learn the game competition relationship between the backdoors.

[0047] Combine the processed suspicious dataset, clean dataset with the unmodified CIFAR-10 training set for training.

[0048] Detection module 1. Obtain the corresponding test set from the CIFAR-10 dataset to test the model, calculate the consistency between the prediction result and the actual label, and evaluate the performance of the model through the accuracy rate. The accuracy is 76.16%.

[0049] 2. Change the pixel values of the upper left 3*3 area of each sample image in the above-mentioned CIFAR-10 test set to 0, bring the modified data into the model for inference, and regard the output result as 0 as a correct prediction. Evaluate the attack success rate of the attack backdoor through the accuracy rate. The accuracy is 97.96%.

[0050] 3. Change the pixel values of the 3×3 area in the middle of each sample image in the above CIFAR-10 test set to 0, input the modified data into the model for inference, and consider the output result of 2 as a correct prediction. Evaluate the defense against backdoor implantation through accuracy, and the accuracy is 93.942%.

[0051] 4. Change the pixel values of the 3×3 area in the middle of each sample image in the suspicious data set to 0, input the modified data into the model for inference, and consider the output result not being 2 as a correct prediction, that is, consider the samples with an output result not being 2 as poisoned samples. Evaluate whether the attack backdoor is successfully detected through accuracy, and the accuracy is 90.52%.

[0052] Backdoor Removal Module 1. Replace the labels corresponding to the poisoned samples detected in the detection stage with the second-highest possible category, and input the modified suspicious data set into the above-trained backdoor model for further training.

[0053] 2. Use the corresponding test set obtained from the CIFAR-10 data set above to test the model, calculate the consistency between the prediction result and the actual label, and evaluate the performance of the model through accuracy, and the accuracy is 74.31%.

[0054] 3. Process the corresponding test set obtained from the CIFAR-10 data set above, change the pixel values of the upper-left 3×3 area to 0, input the modified data into the model for inference, and consider the output result of 0 as a correct prediction. Evaluate the attack success rate of the attack backdoor through accuracy, and the accuracy is 3.32%.

[0055] Summary The accuracy rate of step 4 in the detection module is 90.52%. Combined with the significant reduction in the attack success rate of the attack backdoor in the backdoor removal module, it reflects that the present invention can effectively detect poisoned samples.

[0056] The attack success rate of the attack backdoor in the detection module and the backdoor removal module drops from 97.96% to 3.32%, indicating that the present invention can effectively remove the backdoor in the model.

[0057] The accuracy of the model in the detection module and the backdoor removal module changes from 76.16% to 74.31%, reflecting that the present invention will not significantly reduce the model performance after removing the backdoor.

Claims

1. A model backdoor detection and repair device based on competitive game, characterized in that: include: A learning module is used to randomly implant strong backdoor triggers and weak backdoor triggers into a clean data set, and then train the model after merging the clean data set with the suspicious data set, so that the model learns the competitive game relationship between the strong backdoor trigger and the weak backdoor trigger; The backdoor detection module is used to implant a weak backdoor trigger into each sample in the suspicious data set, and input the sample implanted with the weak backdoor trigger into the trained model to determine whether the sample is poisoned based on the output of the model; The backdoor removal module removes the backdoor in the model by replacing the corresponding labels of the toxic samples detected by the backdoor detection module with the second most likely category and retraining the model with the disinfected samples and clean samples.

2. The device according to claim 1, characterized in that Learning modules include: The clean data set is randomly divided into three parts, where the first part of the data is embedded with a strong backdoor trigger, the second part of the data is embedded with a weak backdoor trigger, and the third part of the data is embedded with both a strong backdoor trigger and a weak backdoor trigger; The labels corresponding to the strong backdoor trigger and the weak backdoor trigger are different. When a sample contains both the strong backdoor trigger and the weak backdoor trigger, the label of the sample is the label corresponding to the strong backdoor trigger. The clean dataset embedded with the backdoor trigger is merged with the suspicious dataset and input into the model for training to obtain the trained model M.

3. The device according to claim 1, characterized in that The backdoor detection module includes: Implant a weak backdoor trigger into each sample in the suspicious dataset; Input the sample with the weak backdoor trigger into the trained model M; Whether the sample is poisoned is determined based on the output of model M. If the model output is the label corresponding to the weak backdoor trigger, the sample is considered not to be poisoned; otherwise, the sample is considered to be poisoned.

4. The device according to claim 1, characterized in that Backdoor removal modules include: For poisoned samples, the output corresponding to the sample is considered to be the attacker's target category. The output corresponding to the sample is essentially the category to which the sample is most likely to belong. Therefore, the output corresponding to the poisoned sample is changed to the second most likely category. While cutting off the connection between the poisoned sample and the attacker's target category, the connection between the poisoned sample and the corresponding correct label before being poisoned is reconstructed. The clean samples and the toxic samples that have been forgotten by the machine form a new training set; Retrain the model M to eliminate the backdoor effect of the toxic samples on the model and obtain a disinfected benign model.

5. A model backdoor detection and repair method based on competitive game, characterized in that: The following steps are involved: Step a. Randomly implant strong backdoor triggers and weak backdoor triggers into the clean data set, and then merge the clean data set with the suspicious data set to train the model, so that the model learns the competitive game relationship between the strong backdoor trigger and the weak backdoor trigger; Step b. implant a weak backdoor trigger into each sample in the suspicious data set, and input the sample implanted with the weak backdoor trigger into the trained model, and determine whether the sample is poisoned based on the output of the model; Step c. Replace the labels corresponding to the toxic samples detected by the backdoor detection module with the second most likely category, and retrain the model with the disinfected samples and clean samples to remove the backdoor in the model.

6. The method according to claim 5, characterized in that Step a includes: The clean data set is randomly divided into three parts, where the first part of the data is embedded with a strong backdoor trigger, and the sample label is modified to the label corresponding to the strong backdoor; the second part of the data is embedded with a weak backdoor trigger, and the sample label is modified to the label corresponding to the weak backdoor; the third part of the data is embedded with both a strong backdoor trigger and a weak backdoor trigger, and the sample label is modified to the label corresponding to the strong backdoor; Ensure that the labels corresponding to the strong backdoor trigger and the weak backdoor trigger are different, and when a sample contains both a strong backdoor trigger and a weak backdoor trigger, the label of the sample is the label corresponding to the strong backdoor trigger; The clean dataset embedded with the backdoor trigger is merged with the suspicious dataset and input into the model for training to obtain the trained model M.

7. The method according to claim 5, characterized in that Step b includes: Implant a weak backdoor trigger into each sample in the suspicious dataset; Input the sample with the weak backdoor trigger into the trained model M; Whether the sample is poisoned is determined based on the output of model M. If the model output is the label corresponding to the weak backdoor trigger, the sample is considered not to be poisoned; otherwise, the sample is considered to be poisoned.

8. The method according to claim 7, characterized in that Step c includes: For poisoned samples, the output corresponding to the sample is considered to be the attacker's target category. The output corresponding to the sample is essentially the category to which the sample is most likely to belong. Therefore, the output corresponding to the poisoned sample is changed to the second most likely category. While cutting off the connection between the poisoned sample and the attacker's target category, the connection between the poisoned sample and the corresponding correct label before being poisoned is reconstructed. The clean samples and the toxic samples that have been forgotten by the machine form a new training set; Retrain the model M to eliminate the backdoor effect of the toxic samples on the model and obtain a disinfected benign model.

9. A storage medium, characterized in that: When the processor executes the program in the storage medium, it implements a model backdoor detection and repair method based on competitive game as described in any one of claims 5-8.

Citation Information

Patent Citations

  • Model backdoor detection method and device, medium and computing equipment

    CN112989340A

  • Game interaction-based confrontation sample detection method and system

    CN114663730A