Model back door detection method and device
By performing forget learning on the model and reverse reconstruction of the trigger, the problem of low accuracy of model backdoor detection in the prior art is solved, and high-accuracy multi-backdoor detection is achieved, which improves the security and reliability of the model.
Patent Information
- Application Number
- CN202410128469.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-29
- Publication Date
- 2025-07-29
AI Technical Summary
The existing model backdoor detection methods have low accuracy and are difficult to effectively detect multiple backdoors, especially when the model has multiple backdoors, the detection accuracy has dropped significantly.
By performing forgetting learning of the target category on the test model, a local forgetting model is obtained, the recognition ability of the target category is weakened, and potential triggers are reversely reconstructed. Combined with the success rate of migration attacks, the existence of the backdoor is judged, and methods such as training alternative models and modifying loss functions are used to accurately detect the backdoor.
It improves the accuracy of model backdoor detection, can effectively detect multiple backdoors, reduces the missed and false alarm rates, and enhances the detection ability of diversified backdoor patterns.
Smart Images

Figure CN120387024A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular, to a method and device for detecting model backdoors. Background Art
[0002] A common threat to artificial intelligence (AI) models is backdoor attacks. For an AI model to be tested (also referred to as the "model to be tested" herein) with multiple categories, the mainstream model backdoor detection directly reversely reconstructs the potential backdoor behavior trigger conditions (Triggers, also referred to as "triggers" herein) for each category of the model to be tested, and runs an outlier detection algorithm to detect whether any potential trigger is significantly smaller than other potential triggers. A significant outlier represents a real trigger, and the category matching this trigger is the target category of the backdoor attack. However, the mainstream model backdoor detection has problems of low accuracy and ineffectiveness in detecting multiple backdoors in the model. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a method and device for detecting model backdoors, which have high accuracy and are effective in detecting multiple backdoors in the model.
[0004] In a first aspect, an embodiment of the present invention provides a method for detecting a model backdoor, the method including:
[0005] Performing forgetting learning on a first model for a target category according to a first data set, where the first data set includes samples of the target category, to obtain a third model;
[0006] According to the third model and the first data set, reversely reconstructing a first trigger with the target being the target category;
[0007] Obtaining a transfer attack success rate according to the first trigger, the first data set, and a second model;
[0008] Obtaining a backdoor detection result of the first model according to the transfer attack success rate. In an embodiment of the present application, the first data set is used to perform forgetting (unlearning) learning on the first model, that is, the model to be tested, to greatly weaken the recognition ability of the model to be tested for samples of the target category, and obtain a third model, that is, a locally forgotten model. In an embodiment of the present application, the model to be tested is fine-tuned by model forgetting to obtain a locally forgotten model for the target category. This locally forgotten model weakens the recognition ability for the target category, while maintaining the ability to recognize other categories or having little impact on the ability to recognize other categories. The first trigger reversely obtained using this locally forgotten model is more accurate, the accuracy of model backdoor detection is higher, it is also effective for various forms of backdoors, and it can also detect models with multiple backdoors.
[0009] In combination with the first aspect, in some implementations of the first aspect, before obtaining the third model by performing forgetting learning of the target category on the first model according to the first data set, it further includes:
[0010] According to the structure of the first model and the first data set, the second model is trained. In the embodiments of the present application, the first data set is used, and the structure of the first model, that is, the model to be tested, is adopted to quickly train the second model, that is, the alternative model of the first model. The second model has the same architecture as the model to be tested. The training set of the second model is smaller, so its ability is weaker than that of the first model, but it can ensure no backdoors.
[0011] In combination with the first aspect, in some implementations of the first aspect, the first data set is a clean data set.
[0012] In combination with the first aspect, in some implementations of the first aspect, obtaining the third model by performing forgetting learning of the target category on the first model according to the first data set includes:
[0013] The first model is divided into a feature extraction layer and a classification task layer, and the parameters of the feature extraction layer are frozen to obtain an initial local forgetting model;
[0014] By removing the samples of the target category in the first data set, a second data set is obtained;
[0015] The classification task layer of the initial local forgetting model is fine-tuned and trained by the second data set to obtain the third model. In the embodiments of the present application, the samples of the target category in the first data set, that is, the training data set, are removed and then the local forgetting model is trained. The first trigger reversed by this local forgetting model is more accurate, the accuracy of model backdoor detection is higher, and it is also effective for various forms of backdoors and can also detect models with multiple backdoors.
[0016] Since the embodiments of the present application perform model forgetting on the model with a backdoor first, and the samples of the target category in the training data set are removed and then the local forgetting model is trained. This local forgetting model weakens the inherent features of the target category and highlights the backdoor features. Therefore, the first trigger reversed by the local forgetting model, that is, the potential trigger, is more accurate and can detect more diverse backdoors.
[0017] In combination with the first aspect, in some implementations of the first aspect, obtaining the third model by performing forgetting learning of the target category on the first model according to the first data set includes:
[0018] The first model is divided into a feature extraction layer and a classification task layer, and the parameters of the feature extraction layer are frozen to obtain an initial local forgetting model;
[0019] By randomly changing the labels of the samples of the target category in the first data set, a third data set is obtained;
[0020] The classification task layer of the initial local forgetting model is fine-tuned and trained through the third data set to obtain the third model. In the embodiments of the present application, a local forgetting model is trained by randomly changing the labels of the target category in the first data set, that is, the training data set. The first trigger obtained by using this local forgetting model reversely is more accurate, the accuracy of model backdoor detection is higher, and it is also effective for various forms of backdoors, and models with multiple backdoors can also be detected.
[0021] Since the embodiments of the present application perform model forgetting on the model with a backdoor first, and a local forgetting model is trained by randomly changing the labels of the target category in the training data set. This local forgetting model weakens the inherent features of the target category and highlights the backdoor features. Therefore, the first trigger obtained by using the local forgetting model reversely, that is, the potential trigger, is more accurate and the detectable backdoors are more diverse.
[0022] In combination with the first aspect, in some implementation manners of the first aspect, the obtaining of the third model by performing forgetting learning on the target category of the first model according to the first data set includes:
[0023] The first model is divided into a feature extraction layer and a classification task layer, and the parameters of the feature extraction layer are frozen to obtain an initial local forgetting model;
[0024] Remove or multiply by 0 the relevant terms of the samples of the target category in the loss function;
[0025] The classification task layer of the initial local forgetting model is fine-tuned and trained through the first data set to obtain the third model. In the embodiments of the present application, a local forgetting model is trained by removing the relevant terms of the samples of the target category in the loss function. The first trigger obtained by using this local forgetting model reversely is more accurate, the accuracy of model backdoor detection is higher, and it is also effective for various forms of backdoors, and models with multiple backdoors can also be detected.
[0026] Since the embodiments of the present application perform model forgetting on the model with a backdoor first, and a local forgetting model is obtained by modifying the loss function without changing the training data set. This local forgetting model weakens the inherent features of the target category and highlights the backdoor features. Therefore, the first trigger obtained by using the local forgetting model reversely, that is, the potential trigger, is more accurate and the detectable backdoors are more diverse.
[0027] In combination with the first aspect, in some implementation manners of the first aspect, the obtaining of the transfer attack success rate according to the first trigger, the first data set, and the second model includes:
[0028] Remove the samples of the target category in the first dataset to obtain a second dataset;
[0029] Attach the first trigger to the samples in the second dataset to obtain a fourth dataset;
[0030] Infer the fourth dataset through the second model to obtain the transfer attack success rate.
[0031] Combined with the first aspect, in some implementation manners of the first aspect, the obtaining the backdoor detection result of the first model according to the transfer attack success rate includes:
[0032] Judge whether the transfer attack success rate is less than a first threshold;
[0033] If it is judged that the transfer attack success rate is less than the first threshold, determine that the backdoor detection result is that the first model has a backdoor trigger with the target being the target category and the first trigger is the backdoor trigger;
[0034] If it is judged that the transfer attack success rate is greater than or equal to the first threshold, determine that the backdoor detection result is that the first model does not have a backdoor trigger with the target being the target category and the first trigger is a universal adversarial sample.
[0035] In a second aspect, an embodiment of the present invention provides a device, including a processor and a memory. Among them, the memory is used to store a program, and when the processor runs the program, the device is enabled to execute the steps of the method as described above.
[0036] In a third aspect, an embodiment of the present invention provides a readable storage medium, and the readable storage medium stores a program, and when the program is run by a device, the device is enabled to execute the method as described above.
[0037] In a fourth aspect, an embodiment of the present invention provides a program product, and the program product includes a program. When the program runs on a device or any at least one processor, the device is enabled to execute the functions / steps in the method as described above.
[0038] In the technical solution of the model backdoor detection method and device provided by the embodiment of the present invention, the method includes: performing forgetting learning of the target category on the first model according to the first dataset to obtain a third model, where the first dataset includes samples of the target category; reversely reconstructing a first trigger with the target being the target category according to the third model and the first dataset; obtaining a transfer attack success rate according to the first trigger, the first dataset and the second model; obtaining a backdoor detection result of the first model according to the transfer attack success rate, with high accuracy and effective for detecting multiple backdoors of the model. Description of the Drawings
[0039] Figure 1 It is a schematic flow chart of mainstream model backdoor detection;
[0040] Figure 2 It is a flow chart of a model backdoor detection method provided by an embodiment of the present invention;
[0041] Figure 3 It is a flow chart of another model backdoor detection method provided by an embodiment of the present invention;
[0042] Figure 4 It is a schematic flow chart of a model backdoor detection method in an embodiment of the present application;
[0043] Figure 5 It is Figure 3 A specific flow chart of forgetting learning of the target category for the first model according to the first data set in to obtain the third model;
[0044] Figure 6 It is another schematic flow chart of a model backdoor detection method in an embodiment of the present application;
[0045] Figure 7 It is Figure 3 Another specific flow chart of forgetting learning of the target category for the first model according to the first data set in to obtain the third model;
[0046] Figure 8 It is another schematic flow chart of a model backdoor detection method in an embodiment of the present application;
[0047] Figure 9 It is Figure 3 Another specific flow chart of forgetting learning of the target category for the first model according to the first data set in to obtain the third model;
[0048] Figure 10 It is another schematic flow chart of a model backdoor detection method in an embodiment of the present application;
[0049] Figure 11 It is Figure 3 A specific flow chart of obtaining the success rate of transfer attack according to the first trigger, the first data set and the second model;
[0050] Figure 12 It is Figure 3 A specific flow chart of obtaining the backdoor detection result of the first model according to the success rate of transfer attack;
[0051] Figure 13 It is a schematic diagram of an application scenario of a model backdoor detection method in an embodiment of the present invention;
[0052] Figure 14 It is a schematic diagram of another application scenario of the model backdoor detection method in the embodiment of the present invention;
[0053] Figure 15 It is a schematic diagram of another application scenario of the model backdoor detection method in the embodiment of the present invention;
[0054] Figure 16 It is a schematic structural diagram of a device provided by the embodiment of the present invention. Detailed implementation manners
[0055] In order to better understand the technical solution of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0056] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0057] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms of "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0058] It should be understood that the term " / and" used herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0059] First, the definitions of the key terms involved in the embodiments of the present application are as follows:
[0060] Artificial intelligence model backdoor attack: An attack that implants a backdoor in an AI model, which may include adding constructed bait data to the training dataset or implanting backdoor behavior by directly modifying the AI model parameters. The AI model with a backdoor mainly has the following three characteristics:
[0061] 1) Concealment: For normal inputs, the AI model makes normal predictions and maintains the prediction accuracy of the original AI model;
[0062] 2) Attack effectiveness: For inputs containing a Trigger, the AI model will predict this input as the target category with a high probability;
[0063] 3) Attack versatility: For inputs containing triggers, the parts outside the triggers have little impact on the attack.
[0064] Backdoor behavior triggering conditions (Triggers): Trigger patterns can be freely defined by the attacker. When a Trigger is included in the input, the model's backdoor behavior will be triggered. For example, an attacker can implant a backdoor in an image recognition model by using a custom, small pattern of their choice as a Trigger.
[0065] For example, a face recognition model with a backdoor implanted will identify any input with Trigger (the face of a person wearing the specific glasses) as the target person because it has been implanted with Trigger (for example, specific glasses), while other cases will be recognized normally.
[0066] Model forgetting: The AI model loses or weakens a specific function through forgetting learning or directly modifying model parameters.
[0067] In recent years, with the rapid development of AI-related technologies, AI has been widely applied in many fields. However, its security has also caused concerns from all walks of life. Existing research has shown that backdoor attacks against AI models are easy to implement, which has also limited the in-depth application of AI models in safety-critical fields such as autonomous driving and facial verification for critical services.
[0068] A common way to implant backdoors is through training data contamination. That is, during the training phase, training data containing Trigger is added to the training dataset and its label is changed to the target category.
[0069] For example, when training a face recognition model, the attacker adds multiple different people with Trigger to the training data and changes their labels to the target person, causing the model to mistakenly identify Trigger as a feature of the target classification. As a result, the trained model will contain a specific backdoor.
[0070] The mainstream model backdoor detection is to reverse the model to be tested, that is, the potential backdoor model, and reconstruct its "Trigger" to determine whether the model has a backdoor. Figure 1 This is a flow chart of mainstream model backdoor detection, such as Figure 1 As shown, mainstream model backdoor detection is divided into the following three steps:
[0071] Step ①: For a given class, consider it as a potential target class for targeted backdoor attacks. By designing an optimization scheme, find the "minimum" trigger required to misclassify all samples from other classes into this target class. In the visual domain, this "minimum" trigger is defined as the smallest set of pixels that can cause misclassification and their associated color intensities.
[0072] Step ②: Repeat Step ① for each output label in the model under test. For a model under test with N classes, this will generate N potential "triggers".
[0073] Step ③: After calculating the N potential triggers, measure the size of each potential trigger by the number of pixels it has, i.e., the number of pixels that the potential trigger is going to replace. Run an outlier detection algorithm to detect if any of the potential triggers is significantly smaller than the others. The significant outlier represents a real trigger, and the class that matches this trigger is the target class of the backdoor attack.
[0074] However, the average detection accuracy of mainstream model backdoor detection for diverse backdoor morphologies is relatively low. Especially when there are multiple backdoors in the model, the detection accuracy drops significantly, resulting in the failure of the detection method.
[0075] Based on the above technical problems, the embodiments of the present invention provide a model backdoor detection method and device, which have high accuracy and are effective for detecting multiple backdoors in the model.
[0076] Figure 2 It is a flowchart of a model backdoor detection method provided by an embodiment of the present invention. As Figure 2 shown, the method includes:
[0077] Step 102: Perform forgetting learning of the target class on the first model according to the first data set, where the first data set includes samples of the target class, to obtain a third model.
[0078] Step 104: According to the third model and the first data set, reversely reconstruct the first trigger with the target being the target class.
[0079] Step 106: Obtain the transfer attack success rate according to the first trigger, the first data set, and the second model.
[0080] Step 108: Obtain the backdoor detection result of the first model according to the transfer attack success rate.
[0081] In the technical solution of the model backdoor detection method provided by the embodiment of the present invention, the method includes: performing forgetting learning on a first model for a target category according to a first data set to obtain a third model, where the first data set includes samples of the target category; reversely reconstructing a first trigger with the target being the target category according to the third model and the first data set; obtaining a transfer attack success rate according to the first trigger, the first data set, and a second model; and obtaining a backdoor detection result of the first model according to the transfer attack success rate, with high accuracy and effective for detecting multiple backdoors of the model.
[0082] Figure 3 It is a flowchart of another model backdoor detection method provided by the embodiment of the present invention. As Figure 3 shown, the method includes:
[0083] Step 202: Train a second model according to the structure of the first model and the first data set.
[0084] In the embodiment of the present application, each step is executed by a device.
[0085] For example, the device includes a server that requires a certain amount of computing power or a high-performance computer, etc.
[0086] Exemplarily, the first model is an AI model to be tested.
[0087] Exemplarily, the first data set is a clean data set in which all samples are not poisoned.
[0088] In this step, using the clean data set and adopting the structure of the model to be tested, a second model, that is, an alternative model of the first model, is quickly trained. The second model has the same architecture as the model to be tested. The training set of the second model is smaller, so its ability is weaker than that of the first model, but it can ensure no backdoors.
[0089] Figure 4 It is a schematic flowchart of the model backdoor detection method in the embodiment of the present application. As Figure 4 shown, in step ①, using the clean data set D and adopting the structure of the model M to be tested, an alternative model M approx is quickly trained. The training set of this alternative model is smaller, so its ability is weaker than that of M, but it ensures no backdoors.
[0090] Step 204: Perform forgetting learning on the first model for a target category according to the first data set to obtain a third model, where the first data set includes samples of the target category.
[0091] Exemplarily, the first data set includes samples of multiple categories, and the category of the current round is the target category.
[0092] In this step, the clean dataset is used to perform forgetting (unlearning) fine-tuning on the model to be tested, significantly weakening the recognition ability of the model to be tested for samples of the target category, and obtaining a third model, that is, a local forgetting model.
[0093] In the embodiment of the present application, the model to be tested is subjected to model forgetting fine-tuning to obtain a local forgetting model for the target category. This local forgetting model weakens the recognition ability for the target category, while maintaining the ability to recognize other categories or having little impact on the ability to recognize other categories.
[0094] As Figure 4 shown, in step ②, the target category of the current round is k. The clean dataset D (the labels of the samples whose target categories need to be changed) is used to perform unlearning fine-tuning on the model M to be tested, significantly weakening the recognition ability of the model M for the normal samples of category k, and obtaining the local forgetting model M k ’.
[0095] In some possible embodiments, as Figure 5 shown, step 204 specifically includes:
[0096] Step 2042: Split the first model into a feature extraction layer and a classification task layer, and freeze the parameters of the feature extraction layer to obtain an initial local forgetting model.
[0097] In this step, the first model is split into a feature extraction layer and a classification task layer, and the parameters of the feature extraction layer are frozen. Therefore, the parameters of the feature extraction layer are fixed, and the parameters of the classification task layer are trainable, thereby obtaining an initial local forgetting model.
[0098] Step 2044: Obtain a second dataset by removing the samples of the target category in the first dataset.
[0099] Step 2046: Fine-tune and train the classification task layer of the initial local forgetting model through the second dataset to obtain a third model.
[0100] In this step, the initial local forgetting model is trained through the second dataset to obtain a third model. Since the parameters of the feature extraction layer of the initial local forgetting model are fixed and the parameters of the classification task layer are trainable, the training process in step 2046 is actually a fine-tuning training of the classification task layer.
[0101] Figure 6 is another flow schematic diagram of the model backdoor detection method in the embodiment of the present application. Compared with Figure 4 , Figure 6 The content of step ② in is more refined. Specifically, as Figure 6As shown, in step ②, after separating the feature extraction layer and the classification task layer from the model M to be tested, fix the parameters of the feature extraction layer, remove all samples of the target class k from the clean dataset D, and fine-tune the parameters of the classification task layer to obtain the local forgetting model M k ’.
[0102] In the embodiment of the present application, the samples of the target class in the first dataset, that is, the training dataset, are removed and then trained to obtain the local forgetting model. The first trigger obtained by using this local forgetting model in reverse is more accurate, the accuracy of model backdoor detection is higher, it is also effective for various forms of backdoors, and it can also detect models with multiple backdoors.
[0103] Since the embodiment of the present application performs model forgetting on the model with a backdoor first, and removes the samples of the target class in the training dataset and then trains to obtain the local forgetting model. This local forgetting model weakens the inherent features of the target class and highlights the backdoor features. Therefore, the first trigger obtained by using the local forgetting model in reverse, that is, the potential trigger, is more accurate and can detect more diverse backdoors.
[0104] In some possible embodiments, as Figure 7 shown, step 204 specifically includes:
[0105] Step 204a: Separate the feature extraction layer and the classification task layer from the first model and freeze the parameters of the feature extraction layer to obtain the initial local forgetting model.
[0106] Step 204b: Randomly change the labels of the samples of the target class in the first dataset to the labels of other classes to obtain the third dataset.
[0107] Step 204c: Fine-tune and train the classification task layer of the initial local forgetting model through the third dataset to obtain the third model.
[0108] Figure 8 This is another flow diagram of the model backdoor detection method in the embodiment of the present application. Compared with Figure 4 , Figure 8 the content of step ② in it is more refined. Specifically, as Figure 8 shown, in step ②, after separating the feature extraction layer and the classification task layer from the model M to be tested, fix the parameters of the feature extraction layer, randomly change the labels of the samples of the target class k in the clean dataset D to the labels of other classes, and fine-tune the parameters of the classification task layer to obtain the local forgetting model M k ’.
[0109] In the embodiments of the present application, a local forgetting model is trained by randomly changing the labels of the target classes in the first data set, i.e., the training data set. The first trigger reversed by this local forgetting model is more accurate, the accuracy of model backdoor detection is higher, it is equally effective for backdoors in various forms, and it can also detect models with multiple backdoors.
[0110] Since the embodiments of the present application perform model forgetting on a model with a backdoor first, a local forgetting model is trained by randomly changing the labels of the target classes in the training data set. This local forgetting model weakens the inherent features of the target classes and highlights the backdoor features. Therefore, the first trigger reversed by the local forgetting model, i.e., the potential trigger, is more accurate and can detect more diverse backdoors.
[0111] In some possible embodiments, such as Figure 9 shown, step 204 specifically includes:
[0112] Step 204A: Split the first model into a feature extraction layer and a classification task layer and freeze the parameters of the feature extraction layer to obtain an initial local forgetting model.
[0113] Step 204B: Remove or multiply by 0 the relevant terms of the samples of the target class in the loss function.
[0114] Exemplarily, the relevant terms of the samples of the target class in the loss function can be understood as the errors of the samples of the target class in the loss function.
[0115] Step 204C: Fine-tune and train the classification task layer of the initial local forgetting model through the first data set to obtain a third model.
[0116] Figure 10 This is another flowchart of the model backdoor detection method in the embodiments of the present application. Compared with Figure 4 , Figure 10 the content of step ② therein is more refined. Specifically, as Figure 10 shown, in step ②, the model under test M is split into a feature extraction layer and a classification task layer, the parameters of the model feature extraction layer are fixed, the relevant terms of the samples of the target class k in the loss function are removed, and the parameters of the classification task layer are fine-tuned to obtain a local forgetting model M k ’.
[0117] In the embodiments of the present application, a local forgetting model is trained by removing the relevant terms of the samples of the target class in the loss function. The first trigger reversed by this local forgetting model is more accurate, the accuracy of model backdoor detection is higher, it is equally effective for backdoors in various forms, and it can also detect models with multiple backdoors.
[0118] Since the embodiments of the present application perform model forgetting on the model with a backdoor first, a locally forgotten model is obtained by modifying the loss function without changing the training data set. This locally forgotten model weakens the inherent features of the target class and highlights the backdoor features. Therefore, the first trigger reverse-engineered from the locally forgotten model, that is, the potential trigger, is more accurate, and the detectable backdoors are more diverse.
[0119] Step 206: According to the third model and the first data set, reverse engineer and reconstruct the first trigger with the target being the target class.
[0120] In this step, a first trigger with the target being the target class is reverse engineered and reconstructed on the locally forgotten model, that is, a potential backdoor (also referred to as "potential trigger", "latent trigger" or "potential Tr igger").
[0121] Such as Figure 4 、 Figure 6 、 Figure 8 and Figure 10 shown, in step ③, reverse engineering is performed on the locally forgotten model M k ’ for the target class k, and a potential tr igger T k with the target being the target class k is reconstructed.
[0122] According to Figure 1 shown, the weaknesses of the backdoor Tr igger reconstruction technology in the mainstream model backdoor detection can be generally inferred. Currently, the main bottleneck is that the reconstruction of the backdoor Tr igger is interfered by multiple features (potential backdoor features, adversarial features, and normal features inherent in the class), resulting in reconstruction failure and the generation of adversarial samples of the inherent features, thus causing false negatives and false positives. By model forgetting, the embodiments of the present application weaken the inherent features of the target class, can reduce interference, lower the difficulty of Tr igger reverse engineering and the possibility of generating adversarial samples of the inherent features, and improve the detection accuracy.
[0123] Step 208: According to the first trigger, the first data set, and the second model, obtain the transfer attack success rate.
[0124] In this step, the first trigger is attached to the clean sample whose class is not the target class, and inference is performed on the second model to count the transfer attack success rate.
[0125] Exemplarily, the transfer attack success rate is the success rate of the transfer attack.
[0126] Such as Figure 4 、 Figure 6 、 Figure 8 and Figure 10 shown, in step ④, the potential tr igger T kPaste it onto the clean samples whose category is not k, and perform inference on the surrogate model M approx to calculate the success rate of the transfer attack.
[0127] In some possible embodiments, as Figure 11 shown, step 208 specifically includes:
[0128] Step 2082: Remove the samples of the target category in the first dataset to obtain a second dataset.
[0129] In this step, the samples of the target category in the clean dataset are removed to obtain a second dataset. Thus, the second dataset is a clean dataset that does not contain samples of the target category.
[0130] Step 2084: Paste the first trigger onto the samples in the second dataset to obtain a fourth dataset.
[0131] In this step, the first trigger is pasted onto the samples in the second dataset to obtain a fourth dataset. Thus, it is achieved to paste the first trigger onto the clean samples whose category is not the target category. The samples in the fourth dataset are all samples that have been pasted with the first trigger and whose category is not the target category.
[0132] Step 2086: Perform inference on the fourth dataset through the second model to obtain the success rate of the transfer attack.
[0133] In this step, inference is performed on the second model using the fourth dataset, and the success rate of the transfer attack is statistically obtained.
[0134] Step 210: Obtain the backdoor detection result of the first model according to the success rate of the transfer attack.
[0135] In some possible embodiments, as Figure 12 shown, step 210 specifically includes:
[0136] Step 2102: Determine whether the success rate of the transfer attack is less than the first threshold. If so, execute step 2104; if not, execute step 2106.
[0137] Step 2104: Determine that the backdoor detection result is that the first model has a backdoor trigger targeting the target category and the first trigger is a backdoor trigger. The process ends.
[0138] In this step, if it is determined that the success rate of the transfer attack is less than the first threshold, it is determined that the backdoor detection result is that the first model has a backdoor trigger targeting the target category and the first trigger is a backdoor trigger.
[0139] Step 2106: Determine that the backdoor detection result is that the first model does not have a backdoor trigger targeting the target category and the first trigger is a general adversarial example. The process ends.
[0140] In this step, if it is determined that the success rate of the transfer attack is greater than or equal to the first threshold, it is determined that the backdoor detection result is that the first model does not have a backdoor trigger with the target class as the target class and the first trigger is a universal adversarial example, rather than a deliberately inserted backdoor.
[0141] Such as Figure 4 、 Figure 6 、 Figure 8 and Figure 10 As shown, in step ⑤, if the success rate of the transfer attack is less than the first threshold, it is determined that the model M to be tested has a backdoor with the target class being k; otherwise, the potential Tr igger T k is determined to be a universal adversarial example, rather than a deliberately inserted backdoor.
[0142] Optionally, if the first data set contains samples of multiple classes, after step 210, other classes in the first data set can also be used as new target classes, and steps 204 to 210 are repeated until all classes are traversed, so that all backdoors and their corresponding classes can be found.
[0143] On the one hand, the local forgetting model obtained by forgetting fine-tuning in the embodiments of the present application eliminates the influence of target inherent features and adversarial features on the model, highlights the backdoor features, is conducive to reverse reconstruction of the potential Tr igger, reduces the false negative rate of the backdoor, and has a good detection effect on diverse backdoor forms. On the other hand, the use of the surrogate model in the embodiments of the present application can make a determination on the reverse-derived potential Tr igger, exclude the first trigger that is actually an adversarial example, and reduce the false positive rate of the detection.
[0144] In the technical solution of the model backdoor detection method provided by the embodiments of the present invention, the method includes: training a second model according to the structure of the first model and the first data set; performing forgetting learning on the first model for the target class according to the first data set to obtain a third model, where the first data set includes samples of the target class; reverse reconstructing the first trigger with the target as the target class according to the third model and the first data set; obtaining the success rate of the transfer attack according to the first trigger, the first data set, and the second model; obtaining the backdoor detection result of the first model according to the success rate of the transfer attack, which has high accuracy and is effective for detecting multiple backdoors of the model.
[0145] The model backdoor detection method provided by the embodiments of the present application can be applied to the following 3 application scenarios:
[0146] (1) Model backdoor detection service scenario: Before an AI model is deployed, it is detected to determine whether it has a backdoor.
[0147] Figure 13FIG. 0 is a schematic diagram of an application scenario of the model backdoor detection method in an embodiment of the present invention. In the normal training process, the neural model training module trains an AI model based on the training data set for subsequent business use. In an embodiment of the present application, a model backdoor detection module is added, and the model backdoor detection module can execute the model backdoor detection method provided in the embodiment of the present application. As Figure 13 shown, the trained model M to be tested and the clean data set D are used as the inputs of the model backdoor detection module together, and the model backdoor detection module determines whether the model M to be tested has a backdoor. When it is detected that the model M to be tested has a backdoor, stop deployment or implement backdoor mitigation, and a Trigger can also be output for the user's reference; when it is detected that the model M to be tested does not have a backdoor, continue to deploy the model M to be tested.
[0148] (2) AI inference scenario: When the AI model is deployed, during the inference stage, it is detected whether the input sample contains a Trigger. If the input sample does not contain a Trigger, directly input the input sample into the AI model for inference; if the input sample contains a Trigger, first remove the Trigger in the input sample and then perform inference.
[0149] Figure 14 FIG. 9 is a schematic diagram of another application scenario of the model backdoor detection method in an embodiment of the present invention. As Figure 14 shown, when an attacker performs a backdoor attack, the input sample x needs to be implanted with a Trigger to take effect, and the potential Trigger can be reverse-engineered through the reverse reconstruction module. Among them, the reverse reconstruction module adopts the solution of reverse reconstructing the first trigger in the model backdoor detection method of the embodiment of the present invention. Using the output potential Trigger, during the inference stage of the model M to be tested, the Trigger contained in the input sample x can be reduced, so that the backdoor attack will not be triggered and the correct result will be output. As Figure 14 shown, given an input sample x, the task objective is to remove the Trigger in it; each input data sample is detected, and if it contains the Trigger reverse-engineered by the reverse reconstruction module, it is removed, so that the backdoor of the model M to be tested will not be triggered and the correct result will still be output, improving the usability of the model.
[0150] (3) Backdoor model proof scenario: When the model backdoor detection method in the embodiment of the present invention detects a backdoor behavior in the model, it displays the potential Trigger form in a visually understandable form for humans, providing materials as evidence for relevant personnel to reference.
[0151] Figure 15 FIG. 20 is a schematic diagram of another application scenario of the model backdoor detection method in an embodiment of the present invention. As Figure 15As shown, when the model backdoor detection module detects that the model has a backdoor, the deployment is stopped, and at the same time, a trigger can be output as evidence of the model backdoor.
[0152] Figure 16 FIG. 4 is a schematic structural diagram of a device provided by an embodiment of the present invention. It should be understood that the device 300 can execute each step in the above model backdoor detection method. To avoid repetition, details are not described herein again. The device 300 includes: a processing unit 301.
[0153] The processing unit 301 is configured to perform forgetting learning of a target category on a first model according to a first data set to obtain a third model, where the first data set includes samples of the target category; reconstruct a first trigger targeted at the target category reversely according to the third model and the first data set; obtain a transfer attack success rate according to the first trigger, the first data set, and a second model; and obtain a backdoor detection result of the first model according to the transfer attack success rate.
[0154] Optionally, before the processing unit 301 performs forgetting learning of the target category on the first model according to the first data set to obtain a third model, the processing unit 301 is further configured to train a second model according to the structure of the first model and the first data set.
[0155] Optionally, the first data set is a clean data set.
[0156] Optionally, the processing unit 301 is specifically configured to separate a feature extraction layer and a classification task layer from the first model and freeze parameters of the feature extraction layer to obtain an initial local forgetting model; obtain a second data set by removing samples of the target category in the first data set; and perform fine-tuning training on the classification task layer of the initial local forgetting model through the second data set to obtain the third model.
[0157] Optionally, the processing unit 301 is specifically configured to separate a feature extraction layer and a classification task layer from the first model and freeze parameters of the feature extraction layer to obtain an initial local forgetting model; randomly change labels of samples of the target category in the first data set to labels of other categories to obtain a third data set; and perform fine-tuning training on the classification task layer of the initial local forgetting model through the third data set to obtain the third model.
[0158] Optionally, the processing unit 301 is specifically configured to separate a feature extraction layer and a classification task layer from the first model and freeze parameters of the feature extraction layer to obtain an initial local forgetting model; remove or multiply an error of a sample of the target category in a loss function by 0; and perform fine-tuning training on the classification task layer of the initial local forgetting model through the first data set to obtain the third model.
[0159] Optionally, the processing unit 301 is specifically configured to remove the samples of the target category in the first data set to obtain a second data set; attach the first trigger to the samples in the second data set to obtain a fourth data set; and perform inference on the fourth data set through the second model to obtain the transfer attack success rate.
[0160] Optionally, the processing unit 301 is specifically configured to determine whether the transfer attack success rate is less than a first threshold; if it is determined that the transfer attack success rate is less than the first threshold, determine that the backdoor detection result is that the first model has a backdoor trigger targeting the target category and the first trigger is the backdoor trigger; if it is determined that the transfer attack success rate is greater than or equal to the first threshold, determine that the backdoor detection result is that the first model does not have a backdoor trigger targeting the target category and the first trigger is a universal adversarial sample.
[0161] It should be understood that the device 300 here is embodied in the form of a functional unit. The term "unit" here can be implemented in software and / or hardware forms, and no specific limitation is made thereto. For example, the "unit" can be a software program, a hardware circuit, or a combination of the two that implements the above functions. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor, or a group of processors, etc.) for executing one or more software or firmware programs, a memory, a merged logic circuit, and / or other suitable components that support the described functions.
[0162] Therefore, the units of the various examples described in the embodiments of the present invention can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraint conditions of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0163] An embodiment of the present application provides a device, which can be a terminal device or a circuit device built into the terminal device. The device can be used to execute the functions / steps in the above method embodiments.
[0164] An embodiment of the present application provides a readable storage medium, in which instructions are stored. When the instructions are run on a terminal device, the device is caused to execute the functions / steps in the above method embodiments.
[0165] An embodiment of the present application also provides a program product including instructions. When the program product runs on a device or any at least one processor, it causes the device to execute the functions / steps in the above method embodiment.
[0166] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent the cases of A existing alone, A and B existing simultaneously, and B existing alone. Where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0167] Those of ordinary skill in the art can realize that the various units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0168] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0169] In several embodiments provided by the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to cause a device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0170] The above are only specific embodiments of the present application. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should be covered within the protection scope of the present application. The protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for detecting model backdoors, characterized in that, The method includes: Performing forgetting learning of a target category on a first model according to a first data set, to obtain a third model, where the first data set includes samples of the target category; According to the third model and the first data set, reversely reconstructing a first trigger with the target being the target category; Obtaining a transfer attack success rate according to the first trigger, the first data set, and a second model; Obtaining a backdoor detection result of the first model according to the transfer attack success rate.
2. The method according to claim 1, wherein Before performing the forgetting learning of the target category on the first model according to the first data set to obtain the third model, it further includes: Training to obtain the second model according to the structure of the first model and the first data set.
3. The method according to claim 1 or 2, characterized in that, The first data set is a clean data set.
4. The method according to any one of claims 1 to 3, characterized in that, The performing forgetting learning of a target category on a first model according to a first data set, to obtain a third model, includes: Separating the first model into a feature extraction layer and a classification task layer, and freezing the parameters of the feature extraction layer, to obtain an initial local forgetting model; Obtaining a second data set by removing samples of the target category in the first data set; Performing fine-tuning training on the classification task layer of the initial local forgetting model through the second data set, to obtain the third model.
5. The method according to any one of claims 1-3, characterized in that, The performing forgetting learning of a target category on a first model according to a first data set, to obtain a third model, includes: Separating the first model into a feature extraction layer and a classification task layer, and freezing the parameters of the feature extraction layer, to obtain an initial local forgetting model; Obtaining a third data set by randomly changing the labels of samples of the target category in the first data set to labels of other categories; Performing fine-tuning training on the classification task layer of the initial local forgetting model through the third data set, to obtain the third model.
6. The method according to any one of claims 1-3, characterized in that, The performing forgetting learning of a target category on a first model according to a first data set, to obtain a third model, includes: Separating the first model into a feature extraction layer and a classification task layer, and freezing the parameters of the feature extraction layer, to obtain an initial local forgetting model; Removing or multiplying by 0 the error of samples of the target category in the loss function; Performing fine-tuning training on the classification task layer of the initial local forgetting model through the first data set, to obtain the third model.
7. The method according to any one of claims 1 to 6, characterized in that, The obtaining a transfer attack success rate according to the first trigger, the first data set, and a second model, includes: Removing samples of the target category in the first data set, to obtain a second data set; Attaching the first trigger to samples in the second data set, to obtain a fourth data set; Performing inference on the fourth data set through the second model, to obtain the transfer attack success rate.
8. The method according to any one of claims 1 to 7, characterized in that The obtaining a backdoor detection result of the first model according to the transfer attack success rate, includes: Judging whether the transfer attack success rate is less than a first threshold; If it is judged that the transfer attack success rate is less than the first threshold, determining that the backdoor detection result is that the first model has a backdoor trigger with the target being the target category and the first trigger is the backdoor trigger; If it is determined that the success rate of the migration attack is greater than or equal to the first threshold, it is determined that the backdoor detection result is that the first model does not have a backdoor trigger targeting the target class and the first trigger is a universal adversarial example.
9. A device, characterized in that, It includes a processor and a memory. Among them, the memory is used to store a program. When the processor runs the program, the device is made to execute the steps of the method according to any one of claims 1-8.
10. A readable storage medium, characterized in that, The readable storage medium stores a program. When the program is run by the device, the device is made to execute the method according to any one of claims 1-8.