Backdoor attack detection method, device and equipment, medium and program product
By transferring local features of the input data to a benign sample set to generate a perturbed sample set, and using a target classification model to determine the consistency of the predicted category, the problem of low accuracy in backdoor attack detection in existing technologies is solved, and efficient and flexible detection of backdoor attacks is achieved.
Patent Information
- Application Number
- CN202511368649.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2025-11-11
AI Technical Summary
Existing backdoor attack detection methods have low accuracy when faced with input-related perturbation implanted data, making it difficult to effectively identify backdoor attacks.
By acquiring a benign sample set and input data, local features of the input data are transferred to each benign sample in the benign sample set to generate a perturbation sample set. Based on the target classification model, the predicted categories of the perturbation sample set and the benign sample set are determined. The consistency of the predicted categories is used to determine whether the input data is backdoor data.
It achieves high detection accuracy for both input-independent and input-dependent backdoor attacks, enabling online detection, reducing reliance on model training and computational overhead, and improving the accuracy and defense capabilities of backdoor attack detection.
Smart Images

Figure CN120934895A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a backdoor attack detection method, apparatus, device, medium, and program product. Background Technology
[0002] The rapid development of deep learning technology has brought revolutionary changes to many fields, demonstrating unprecedented capabilities. However, when applying deep learning technology to critical infrastructure, trust issues have raised concerns, with a significant focus on how to detect pre-installed backdoor attacks. A backdoor attack is an attack method that introduces malicious behavior into a machine learning model. Attackers implant hidden "backdoors" into the training data to make the model exhibit expected erroneous behavior under specific conditions, while maintaining good performance under other normal conditions.
[0003] Backdoor detection is a necessary defense against backdoor vulnerabilities. It aims to identify whether a model has been implanted with a backdoor and whether the input contains malicious triggering conditions by capturing the abnormal characteristics of backdoor attacks.
[0004] Currently, backdoor implantation methods are mainly divided into two categories: implanting perturbations unrelated to the input and implanting perturbations related to the input. The former uses the same perturbation on all input data and has a fixed triggering pattern; the latter's triggers correspond one-to-one with the input data, adding subtle perturbations to the input data, often making it more covert and undetectable to the naked eye. Current backdoor data detection methods are relatively effective against perturbations implanted that are unrelated to the input, but they are almost ineffective against perturbations implanted that are related to the input. Meanwhile, input-related backdoor attack methods are emerging in large numbers and have become the mainstream of backdoor attacks, resulting in low accuracy in backdoor attack detection and increasing the difficulty of backdoor security defense. Summary of the Invention
[0005] This invention provides a backdoor attack detection method, apparatus, device, medium, and program product to solve the problem that current backdoor attack detection methods are almost ineffective in the face of input-related perturbation implanted data, resulting in low backdoor attack detection accuracy.
[0006] In a first aspect, embodiments of the present invention provide a backdoor attack detection method, including:
[0007] Obtain a benign sample set and input data;
[0008] The local features of the input data are transferred to each benign sample in the benign sample set to obtain the perturbation sample set;
[0009] The predicted categories of the perturbed sample set and the predicted categories of the benign sample set are determined based on the target classification model.
[0010] Based on the predicted category of the benign sample set and the predicted category of the perturbation sample set, determine whether the input data is backdoor data.
[0011] Secondly, embodiments of the present invention provide a backdoor attack detection device, comprising:
[0012] The data acquisition module is used to acquire a benign sample set and input data;
[0013] The perturbation sample generation module is used to transfer local features of the input data to each benign sample in the benign sample set to obtain the perturbation sample set.
[0014] The classification prediction module is used to determine the predicted category of the perturbed sample set and the predicted category of the benign sample set based on the target classification model.
[0015] The backdoor data detection module is used to determine whether the input data is backdoor data based on the predicted category of the benign sample set and the predicted category of the perturbation sample set.
[0016] Thirdly, embodiments of the present invention provide an electronic device, the electronic device comprising:
[0017] At least one processor;
[0018] and a memory communicatively connected to the at least one processor;
[0019] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the backdoor attack detection method according to any embodiment of the present invention.
[0020] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the backdoor attack detection method according to any embodiment of the present invention. Fifthly, embodiments of the present invention provide a computer program product including a computer program that, when executed by a processor, implements the backdoor attack detection method according to any embodiment of the present invention.
[0021] The technical solution of this invention involves acquiring a benign sample set and input data; transferring local features of the input data to each benign sample in the benign sample set to obtain a perturbation sample set; determining the predicted category of the perturbation sample set and the predicted category of the benign sample set based on a target classification model; and determining whether the input data is backdoor data based on the predicted categories of the benign sample set and the perturbation sample set. Utilizing the strong correlation between backdoor attacks implanted in the model and perturbation patterns, the local features of the input data are transferred to benign samples to obtain a perturbation sample set. Backdoor attack detection is then performed based on the consistency of the predicted categories of the benign sample set and the perturbation sample set using a target classification model. This method has strong universality against backdoor attack methods, is not constrained by the attack method, and has high detection accuracy against both input-independent and input-related backdoor attacks. It solves the problem that current backdoor data detection methods are almost ineffective against input-related perturbation data, resulting in low backdoor attack detection accuracy. Therefore, this method improves the accuracy of backdoor attack detection and enhances backdoor attack defense capabilities.
[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart of a backdoor attack detection method provided in Embodiment 1 of the present invention;
[0025] Figure 2 This is a flowchart of a backdoor attack detection method provided in Embodiment 2 of the present invention;
[0026] Figure 3 This is a flowchart of a backdoor attack detection method provided in Embodiment 3 of the present invention;
[0027] Figure 4 This is a schematic diagram of a backdoor attack detection device provided in Embodiment 4 of the present invention;
[0028] Figure 5 A schematic diagram of the structure of an electronic device for implementing the backdoor attack detection method of this invention. Detailed Implementation
[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0031] Example 1
[0032] Figure 1 This is a flowchart of a backdoor attack detection method provided in Embodiment 1 of the present invention. This embodiment is applicable to detecting whether the data input to a target classification model has been implanted with a backdoor attack. This method can be executed by a backdoor attack detection device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0033] S110. Obtain a benign sample set and input data.
[0034] The benign sample set can be considered as a collection of multiple benign samples. Benign samples can be considered as normal samples that have not been implanted with backdoor attacks. The benign sample set is used as benchmark data for backdoor attack detection. The known classification results of benign samples must be consistent with the predicted classification results of the target classification model for benign samples; if they differ, the benign sample can be deleted and a new one constructed.
[0035] Input data can be considered as data that needs to be classified using a target classification model. During classification, it is necessary to detect whether the input data is backdoor data. Input data can be data generated and collected in actual production or daily life, or it can be data generated manually or based on a model.
[0036] The data types of input data and benign samples can be images or text, etc. This embodiment of the invention does not limit the source or format of the input data and benign samples. It should be understood that the type and size of the input data and benign samples should be the same to enable subsequent feature transfer.
[0037] For example, a benign sample set can be data selected from the training dataset of the target classification model that has no backdoor attacks and whose category prediction is accurate; or it can be data collected manually that belongs to a preset category and whose prediction result is also a preset category, without relying on the training dataset.
[0038] In one implementation, the total number of categories in the target classification model Random selection Categories, and for the selected For each category, a data point is randomly selected from the training set to form a benign sample set. These benign samples all require manual verification to ensure accurate classification. The number of data points in the benign sample set does not need to be large, for example, ten to several dozen, thus minimizing manual labor costs. In another implementation, for... For each category, data is manually constructed without selecting from the training set; the accuracy of the classification is the primary concern. It should be understood that other methods besides those mentioned above can be used to determine the beneficial sample set, as long as the classification accuracy is maintained. The beneficial sample set, once generated, requires no further modification and can be reused.
[0039] S120. Transfer the local features of the input data to each benign sample in the benign sample set to obtain the perturbation sample set.
[0040] The local features of the input data can be the features of the input data in a local region. The perturbation sample set is a collection of perturbation samples. A perturbation sample can be considered as data obtained by adding perturbations about the input data to the benign sample set. It should be understood that each perturbation sample in the perturbation sample set needs to have a perturbation about the input data added, and the perturbation sample set and the benign sample set have the same number of data points.
[0041] Specifically, for each benign sample in the benign sample set, local features of the input data are extracted and transplanted to the corresponding positions in the benign sample to obtain perturbation samples related to the input data. This process of transplanting local features of the input data into each benign sample in the benign sample set yields the perturbation sample set.
[0042] For example, for input data of text or image type, local features of the input data can be extracted using feature extraction networks or tools, or local regions can be identified and their features extracted using feature extraction algorithms. For instance, for image input data, local pixel features can be extracted; for text input data, local word vector features can be extracted.
[0043] Current backdoor sample detection methods only achieve good accuracy when facing backdoor attack algorithms with triggering conditions unrelated to the input. However, in this embodiment, since the perturbation in each perturbation sample originates from local features of the input data, the strong correlation between the model-implanted backdoor and the perturbation pattern creates a one-to-one correspondence with the input data, thus achieving high detection accuracy for input-related backdoor attacks. Simultaneously, it is also effective against attacks unrelated to the input, is not constrained by the attack method, and has strong universality against backdoor attack methods. This effectively addresses the shortcomings of current backdoor attack detection methods and improves the detection accuracy of backdoor attacks.
[0044] S130. Determine the predicted category of the perturbed sample set and the predicted category of the benign sample set based on the target classification model.
[0045] The target classification model can be considered a classification model potentially vulnerable to backdoor attacks. The target classification model can be a fully trained model with functions such as anomaly detection and target recognition; this embodiment of the invention does not impose any limitations on this. The predicted category can be considered the category obtained by the target classification model from classifying and predicting the input data.
[0046] It should be noted that if the input data is backdoor data and triggers a deceptive prediction in the target classification model (i.e., outputting an incorrect predicted category), the target classification model may be infected with a backdoor attack.
[0047] Specifically, the perturbation sample set is input into the target classification model to obtain the predicted category output by the target classification model for each perturbation sample in the perturbation sample set. Similarly, the benign sample set is input into the target classification model to obtain the predicted category output by the target classification model for each benign sample in the benign sample dataset. Since perturbation samples are obtained by adding local features about the input data to the benign samples, if the input data is backdoor data, it may trigger the target classification model, which is under attack by the backdoor, to output an incorrect predicted category, resulting in the predicted category of the perturbation sample differing from the predicted category of the benign sample.
[0048] For example, perturbation samples from the perturbation sample set are input into the target classification model sequentially or in batches for classification. The data in the perturbation sample set maintains a one-to-one correspondence with the benign samples in the benign sample set, obtaining a predicted category for each perturbation sample. This predicted category is the category predicted for the benign sample under the interference of local features of the input data. It should be noted that the prediction of both the perturbation sample set and the benign sample set belongs to the model inference stage, which has relatively low computational overhead and can be met by using a CPU.
[0049] S140. Based on the consistency between the predicted categories of the benign sample set and the predicted categories of the perturbation sample set, determine whether the input data is backdoor data.
[0050] Backdoor data can be understood as data that has been implanted with a backdoor attack. A model implanted with a backdoor attack will output predictions of a specific type of error based on this backdoor data. The consistency of prediction categories can refer to the quantity, proportion, or distribution of predictions that are the same or different meeting certain conditions.
[0051] Specifically, for benign samples in the benign sample set and perturbation samples in the corresponding perturbation sample set, the predicted categories of the benign sample set and the perturbation sample set are compared to determine whether the input data is backdoor data. For example, if the predicted categories of the benign sample set and the perturbation sample set are consistent, the input data is determined not to be backdoor data; if the predicted categories of the benign sample set and the perturbation sample set are inconsistent, the input data is determined to be backdoor data.
[0052] For example, the consistency between the predicted class of the benign sample set and the predicted class of the perturbation sample set can be determined based on the number of samples in which the predicted class of the benign sample set is consistent with the predicted class of the perturbation sample set and / or the probability that the benign sample set and the perturbation sample set are inconsistent in each predicted class.
[0053] The method to determine whether input data is backdoor data based on the consistency between the predicted categories of the benign sample set and the perturbation sample set can be as follows: Count the number of data points where the predicted categories of the benign sample set and the perturbation sample set are consistent, i.e., the number of correctly classified data points; if the number of correctly classified data points exceeds an upper threshold, then the input data is determined not to be backdoor data. Alternatively, count the number of data points where the predicted categories of the benign sample set and the perturbation sample set are inconsistent, i.e., the number of misclassified data points; if the number of misclassified data points does not exceed a lower threshold, then the input data is determined not to be backdoor attack data.
[0054] If the number of correctly classified data does not exceed the upper threshold or the number of misclassified data exceeds the lower threshold, in one implementation, the input data can be determined to be backdoor data. Furthermore, the error probability of each category can be determined, and the category with the highest error probability can be identified as the backdoor attack category of the input data. In another implementation, if the number of correctly classified data does not exceed the upper threshold or the number of misclassified data exceeds the lower threshold, the error probability of each category can be determined. If the error probability corresponding to the category with the highest error probability exceeds a set threshold, the input data is determined to be backdoor data, and the category with the highest error probability is identified as the backdoor attack category of the input data.
[0055] Compared to current backdoor attack detection algorithms, which typically employ techniques such as model editing, reverse training, or training set clustering, requiring extensive offline computation, resulting in high computational costs and numerous dependencies, the backdoor detection algorithm provided in this embodiment can directly input input data into a running target classification model for real-time backdoor attack detection. Detection is completed only at the inference level, achieving online backdoor attack detection. It is simple, efficient, and flexible in use, effectively reducing the dependence of backdoor attack detection on model training and computational costs, and does not interfere with the running target classification model.
[0056] The technical solution of this invention involves acquiring a benign sample set and input data; transferring local features of the input data to each benign sample in the benign sample set to obtain a perturbation sample set; determining the predicted category of the perturbation sample set and the predicted category of the benign sample set based on a target classification model; and determining whether the input data is backdoor data based on the predicted categories of the benign sample set and the perturbation sample set. Utilizing the strong correlation between backdoor attacks implanted in the model and perturbation patterns, local features of the input data are transferred to benign samples to obtain a perturbation sample set, and backdoor attack detection is performed based on the consistency of the predicted categories of the benign sample set and the perturbation sample set using a target classification model. This method has strong universality against backdoor attack methods, is not constrained by the attack method, and has high detection accuracy for both input-independent and input-related backdoor attacks. Furthermore, it can directly interface with the running target model, enabling online backdoor attack detection and effectively reducing the dependence of backdoor attack detection on model training and computational overhead.
[0057] Example 2
[0058] Figure 2This is a flowchart of a backdoor attack detection method provided in Embodiment 2 of the present invention. Based on the above embodiments, this embodiment further refines the generation method of perturbation samples in S120. Specifically, local features of the input data are transferred to each benign sample in the benign sample set to obtain a perturbation sample set, including: extracting local features of the input data; for each benign sample in the benign sample set, replacing the original data at the corresponding position in the benign sample with the local features to obtain a perturbation sample; and determining the perturbation sample set composed of each perturbation sample.
[0059] like Figure 2 As shown, the method includes:
[0060] S210. Obtain a benign sample set and input data.
[0061] S220. Extract local features from the input data.
[0062] For example, local features can be features in the input data whose feature values exceed a threshold. Feature values can be pixel values in image data, or gradients of the classification prediction probabilities of the input data, etc.
[0063] In one alternative implementation, pixel values of regularized shapes randomly sampled from image-type input data are used as local features of the input data.
[0064] In another alternative implementation, the extraction of local features from the input data includes:
[0065] A1. The target classification model is used to classify the input data to obtain the predicted probability of the input data belonging to each predicted category.
[0066] Specifically, the input data I is input into the target classification model to obtain the prediction results obtained by the target classification model in classifying the input data. The prediction results include which categories the input data belongs to. Predicted probability y c .
[0067] A2. Calculate the gradient of the predicted probability of each predicted category with respect to the input data in the feature map of the target classification model using backpropagation.
[0068] Backpropagation starts from the output layer and uses the chain rule to calculate the gradient of the output value with respect to the parameters of each layer, moving forward layer by layer. The gradient calculation is performed in the opposite direction to the forward propagation. The magnitude of the gradient indicates the prediction probability y of the predicted class c for each activation value at each location and channel on the feature map F. c The importance of.
[0069] Specifically, obtain the feature map F of the input data from the output of the last convolutional layer of the target classification model, and calculate the prediction probability y using backpropagation of the target classification model. c Feature map of input data gradient .
[0070] A3. Perform weighted averaging, normalization, and size adjustment on the feature map according to the gradient to obtain a heatmap of the input data.
[0071] For example, the gradient weight of each feature channel k of the feature map is calculated based on the gradient of the feature map. The weight represents the weight of the k-th feature channel in relation to the predicted category. The global importance of gradient weights can be determined by performing global average pooling on the gradient along both the height (H) and width (W) dimensions, resulting in a weight vector of length C (number of channels). Based on gradient weights The original feature map F is weighted, summed for each channel, and averaged to obtain a weighted average feature map. To visualize the weighted average feature map as an easily interpretable heatmap, the values of the weighted average feature map need to be normalized to a fixed interval (usually [0,1]) and resized to the same size as the input data, resulting in the heatmap FGrad corresponding to the input data. The heatmap's heat values are located in... The larger the value in the interval, the more critical the data at that location is to the model's prediction.
[0072] A4. Determine the masked region in the heat map where the heat value exceeds the first threshold, and obtain the local features corresponding to the pixels of the input data in the masked region.
[0073] The first threshold can be considered as the security threshold for feature extraction of input data in backdoor attack detection.
[0074] For example, a first threshold is set, and regions in the heatmap FGrad that exceed the first threshold are extracted to construct a mask matrix. The mask matrix is a 0-1 matrix, where 1 indicates that a pixel at that location is used. Input data is then obtained. Local features .
[0075] S230. Replace the original data at the corresponding position of each benign sample in the benign sample set with local features to obtain the perturbed sample set.
[0076] Specifically, for each benign sample in the benign sample set , where n is the total number of benign samples in the benign sample set. The original pixels in the benign samples are removed according to the mask matrix mask and replaced with local features from the input number I. Obtain the perturbation sample ,Right now: Thus, the perturbation sample set is determined as .
[0077] S240. Determine the predicted category of the perturbed sample set and the predicted category of the benign sample set based on the target classification model.
[0078] S250. Based on the predicted categories of the benign sample set and the predicted categories of the perturbation sample set, determine whether the input data is backdoor data.
[0079] The technical solution of this invention involves acquiring a benign sample set and input data; extracting local features from the input data; replacing the original data at corresponding positions of each benign sample in the benign sample set with the local features to obtain a perturbation sample set; determining the predicted category of the perturbation sample set and the predicted category of the benign sample set based on a target classification model; and determining whether the input data is backdoor data based on the predicted categories of the benign sample set and the perturbation sample set. Utilizing the strong correlation between backdoor attacks implanted in the model and perturbation patterns, local features of the input data are transferred to benign samples to obtain a perturbation sample set, and backdoor attack detection is performed based on the consistency of the predicted categories of the benign sample set and the perturbation sample set using a target classification model. This method has strong universality against backdoor attack methods, is not constrained by the attack method, and has high detection accuracy against both input-independent and input-related backdoor attacks. Furthermore, it can directly interface with the running target model, enabling online backdoor attack detection and effectively reducing the dependence of backdoor attack detection on model training and computational overhead.
[0080] Example 3
[0081] Figure 3 This is a flowchart of a backdoor attack detection method provided in Embodiment 3 of the present invention. Based on the above embodiments, this embodiment further refines the backdoor attack determination method in S140. Specifically, determining whether the input data is backdoor data based on the predicted category of the benign sample set and the predicted category of the perturbation sample set includes: determining the correct classification ratio and misclassification probability distribution of the perturbation sample set based on the predicted category of the benign sample set and the predicted category of the perturbation sample set; wherein, the correct classification ratio is the ratio of the number of correctly classified perturbation samples to the total number of data in the perturbation sample set, and the misclassification probability distribution is the probability distribution of misclassified perturbation samples; determining whether the input data is backdoor data based on the correct classification ratio and the misclassification probability distribution.
[0082] like Figure 3 As shown, the method includes:
[0083] S310. Obtain a benign sample set and input data.
[0084] S320. Transfer the local features of the input data to each benign sample in the benign sample set to obtain the perturbation sample set.
[0085] S330. Based on the target classification model, determine the predicted category of the perturbed sample set and the predicted category of the benign sample set.
[0086] Specifically, the implementation steps of S310 to S330 can be referred to in Embodiment 1 and Embodiment 2, and will not be repeated in this embodiment.
[0087] S340. Based on the predicted categories of the benign sample set and the predicted categories of the perturbed sample set, determine the correct classification ratio and the misclassification probability distribution of the perturbed sample set.
[0088] Wherein, the correct classification ratio is the ratio of the number of correctly classified perturbation samples to the total number of data in the perturbation sample set, and the misclassification probability distribution is the probability distribution of misclassified perturbation samples. A correctly classified perturbation sample can be considered as a perturbation sample whose predicted class is the same as the predicted class of its corresponding benign sample. A misclassified perturbation sample can be considered as a perturbation sample whose predicted class is different from the predicted class of its corresponding benign sample. The predicted class of a misclassified perturbation sample can be considered as the incorrectly predicted class of the perturbation sample.
[0089] Specifically, the predicted categories of benign categories in the benign sample set are compared with the predicted categories of corresponding perturbation samples in the perturbation sample set. Perturbation samples whose predicted categories match the predicted categories of their corresponding benign samples are identified as correctly classified perturbation samples, while those whose predicted categories differ are identified as misclassified perturbation samples. The number of correctly classified perturbation samples is counted, and the ratio of the number of correctly classified perturbation samples to the total number of data points is determined as the correct classification ratio of the perturbation sample set. For each misclassified category in the perturbation sample set, the ratio of the number of perturbation samples corresponding to the misclassified category to the total number of data points is determined as the misclassification probability of that misclassified category; the misclassification probability distribution is determined based on the misclassification probabilities corresponding to each misclassified category.
[0090] For example, the correct classification ratio is calculated as follows:
[0091]
[0092] Where n is the total number of data points in the perturbation sample set. Used to indicate the Disturbance samples Prediction category With the corresponding first A benign sample Prediction category Are they the same? If they are the same, then... If the perturbation sample is correctly classified, then... These are perturbation samples that have been misclassified.
[0093] For example, misclassification probability distribution The Middle The probability of misclassification for each incorrectly predicted category The calculation method is as follows:
[0094]
[0095] Where C represents the total number of categories in the target classification model. Used to indicate the Disturbance samples Prediction category With the If the error prediction categories are the same, then the first... Disturbance samples The error prediction category is the first There are several categories of incorrect predictions.
[0096] S350. Determine whether the input data is backdoor data based on the correct classification ratio and the misclassification probability distribution.
[0097] Specifically, if the proportion of correct classifications exceeds the second threshold, the input data is determined not to be backdoor data; if the proportion of correct classifications does not exceed the second threshold, the input data is determined to be backdoor data based on the probability distribution of misclassifications.
[0098] In an optional implementation, determining whether the input data is backdoor data based on the correct classification ratio and the incorrect classification probability distribution includes:
[0099] B1. If the correct classification ratio exceeds the second threshold, then the input data is determined not to be backdoor data.
[0100] The second threshold can be understood as a threshold used to determine whether the prediction accuracy of the target classification model for perturbed samples is normal. The second threshold can be determined based on the normal prediction accuracy of the target classification model.
[0101] Specifically, if the correct classification ratio exceeds the second threshold, it indicates that the accuracy of the target classification model in predicting the category of the perturbed sample set is normal, suggesting that the input data transplanted into the perturbed sample is not backdoor data.
[0102] B2. If the correct classification ratio does not exceed the second threshold, and the highest misclassification probability in the misclassification probability distribution exceeds the third threshold, then the input data is determined to be backdoor data.
[0103] The third threshold can be considered as the threshold for judging whether a misclassified perturbation sample meets the characteristics of a backdoor attack.
[0104] Specifically, since the model with the implanted backdoor only exhibits the expected erroneous behavior on specific backdoor data, while maintaining good performance under other normal conditions, if the correct classification ratio does not exceed the second threshold, it cannot be directly concluded that the accuracy of the target classification model's prediction of the perturbation sample set conforms to the characteristics of a backdoor attack. Further judgment is needed based on the misclassification probability distribution. If the highest misclassification probability in the misclassification probability distribution exceeds the third threshold, it indicates that a certain number of perturbation samples have incorrect prediction types, and the incorrect prediction categories are all the same. It can be considered that specific backdoor data exhibits the expected erroneous behavior, thus confirming that the input data implanted in the perturbation samples is backdoor data.
[0105] B3. If the correct classification ratio does not exceed the second threshold, and the highest misclassification probability in the misclassification probability distribution does not exceed the third threshold, then the input data is determined not to be backdoor data.
[0106] Specifically, if the correct classification rate does not exceed the second threshold and the highest misclassification probability in the misclassification probability distribution does not exceed the third threshold, it indicates that although the target classification model has a low accuracy in predicting the category of the perturbation sample set, it does not exhibit the expected erroneous behavior and does not conform to the characteristics of a backdoor attack. Therefore, it is determined that the input data implanted in the perturbation sample is not backdoor data.
[0107] As an optional embodiment, after determining that the input data is backdoor data, the method further includes:
[0108] C1. The predicted category corresponding to the highest misclassification probability is determined as the backdoor attack category of the input data.
[0109] Among them, the backdoor attack category can be understood as the category corresponding to the backdoor data.
[0110] Specifically, after determining that the input data is backdoor data based on the correct classification ratio and the highest misclassification probability in the misclassification probability distribution, the predicted category corresponding to the highest misclassification probability in the misclassification probability distribution of the perturbation sample set is determined as the backdoor attack category of the input data transplanted in the perturbation sample.
[0111] C2. Add the backdoor attack category of the input data to the backdoor attack category list.
[0112] The backdoor attack category list can be understood as a list that records various backdoor attack categories and their frequency of occurrence.
[0113] Specifically, the backdoor attack categories of the input data are added to the backdoor attack category list as reference attack categories for the target classification model to identify backdoor attack categories.
[0114] C3. The backdoor attack categories that appear more frequently than a preset frequency in the backdoor attack category list are determined as the backdoor attack categories into which the target classification model is implanted.
[0115] In this embodiment, backdoor attack detection can be performed on multiple input data, and the backdoor attack categories corresponding to the input data detected as backdoor data can be added to a backdoor attack category list. If there are backdoor attack categories in the backdoor attack category list that appear more frequently than a preset frequency, it indicates that the target model exhibits the expected erroneous behavior for the specific input data. Therefore, the backdoor attack categories that appear more frequently than the preset frequency can be used to determine the backdoor attack categories implanted in the target classification model.
[0116] As an optional embodiment, after determining whether the input data is backdoor data, the method further includes:
[0117] D1. If the input data is not backdoor data, then output the predicted category of the target classification model for the input data.
[0118] Specifically, this embodiment provides an online backdoor attack detection method that directly inputs input data into a target classification model potentially vulnerable to backdoor attacks. Since the backdoor attack model only exhibits the expected erroneous behavior in response to backdoor data, even if the target classification model is infected with a backdoor attack, it will still output the correct predicted category if the input data is not detected as backdoor data. Therefore, there is no need to take measures to interfere with the target classification model's classification prediction process for the input data, minimizing the interference of backdoor attack detection on the running model.
[0119] D2. If the input data is backdoor data, then discard the input data and do not output the predicted category of the target classification model for the input data.
[0120] Specifically, if the input data is backdoor data, it indicates that the target classification model is at risk of being implanted with a backdoor attack, and the target classification model's prediction of the input data is incorrect, resulting in deception. Therefore, defensive measures are needed to discard the input data, and the target classification model should not output the predicted category of the input data, thus defending against backdoor attacks and mitigating security risks. Furthermore, a prompt message can be sent to the user, alerting them to the potential backdoor risk in the input data and suggesting that they change the input data.
[0121] This embodiment uses an online target classification model to detect backdoor attacks on input data and takes protective measures to remove data when backdoor data is detected. It integrates prediction classification, backdoor data detection, and backdoor attack defense into a single backdoor attack detection solution. It allows a target classification model with potential backdoor attacks to classify and predict normal data while simultaneously defending against backdoor attacks, making it more widely applicable, flexible, efficient, and with high practical application value.
[0122] The technical solution of this invention involves acquiring a benign sample set and input data; transferring local features of the input data to each benign sample in the benign sample set to obtain a perturbation sample set; determining the predicted category of the perturbation sample set and the predicted category of the benign sample set based on a target classification model; determining the correct classification ratio and misclassification probability distribution of the perturbation sample set based on the predicted categories of the benign and perturbation sample sets; and determining whether the input data is backdoor data based on the correct classification ratio and misclassification probability distribution. Utilizing the strong correlation between backdoor attacks implanted in the model and perturbation patterns, local features of the input data are transferred to benign samples to obtain a perturbation sample set, and backdoor attack detection is performed based on the correct classification ratio and misclassification probability distribution of the benign and perturbation sample sets classified by the target classification model. This method has strong universality against backdoor attack methods, is not constrained by the attack method, and has high detection accuracy against both input-independent and input-related backdoor attacks. Furthermore, it can directly interface with the target model at runtime, enabling online backdoor attack detection and effectively reducing the dependence of backdoor attack detection on model training and computational overhead.
[0123] In a specific example, the target classification model uses the ResNet50 model, the backdoor attack category can be category L0, and the backdoor attack algorithm used is the image quantization attack BppAttack. According to the backdoor attack detection method provided in this embodiment of the invention, the backdoor detection process is as follows:
[0124] (1) A benign sample set GSet was formed from 10 images of categories L11 to L20. After prediction by the target classification model ResNet50, the classification category of each benign sample was correct.
[0125] (2) Assume there are two input data X and X' adv X represents normal data. adv This is backdoor data. For each benign sample in the benign sample set GSet according to any of the above embodiments, a perturbed sample set DSet_X is obtained; X is then... adv The perturbed sample set DSet_X_adv is obtained by migrating the samples to each benign sample in the benign sample set GSet.
[0126] (3) Each perturbation sample in the two perturbation sample sets DSet_X and DSet_X_adv is classified by the target classification model ResNet50 to obtain the predicted category of each perturbation sample. For example, the prediction results of the two perturbation sample sets DSet_X and DSet_X_adv are shown in Table 1.
[0127] (4) Determine whether there is a backdoor attack in the input data based on the consistency of the predicted categories of the benign sample set and the perturbation sample set.
[0128] For example, assume the first threshold for class consistency (i.e., the second threshold in the above embodiment) is 0.7, and the second threshold for class consistency (i.e., the third threshold in the above embodiment) is 0.3. For input data X, the correct classification ratio of its perturbation sample set DSet_X is 0.9, which exceeds the first threshold for class consistency, and sample X is judged as normal data. adv The correct classification ratio of the perturbation sample set DSet_X_adv is 0.1, which does not exceed the first threshold of class consistency. Continuing the comparison with the second threshold of class consistency, if the probability of class L0 appearing in the misclassification probability distribution is the highest (e.g., 0.6) and exceeds the second threshold of class consistency, then the input data X... adv The data was identified as a backdoor, and category L0 is a potential backdoor attack category.
[0129] (5) For input sample X, the target model ResNet50 predicts and outputs the predicted category; while for input sample X adv Since the detected data is backdoor data, it is rejected, and no prediction result is given. At the same time, a security risk warning is returned to the user, suggesting that the input data be changed.
[0130] Table 1
[0131]
[0132] Example 4
[0133] Figure 4 This is a schematic diagram of a backdoor attack detection device provided in Embodiment 3 of the present invention. Figure 4 As shown, the device includes: a data acquisition module 410, a perturbation sample set generation module 420, a classification prediction module 430, and a backdoor data detection module 440; wherein,
[0134] The data acquisition module 410 is used to acquire a benign sample set and input data;
[0135] The perturbation sample generation module 420 is used to transfer local features of the input data to each benign sample in the benign sample set to obtain a perturbation sample set;
[0136] The classification prediction module 430 is used to determine the predicted category of the perturbed sample set and the predicted category of the benign sample set based on the target classification model.
[0137] The backdoor data detection module 440 is used to determine whether the input data is backdoor data based on the predicted category of the benign sample set and the predicted category of the perturbation sample set.
[0138] The technical solution of this invention involves acquiring a benign sample set and input data; transferring local features of the input data to each benign sample in the benign sample set to obtain a perturbation sample set; determining the predicted category of the perturbation sample set and the predicted category of the benign sample set based on a target classification model; and determining whether the input data is backdoor data based on the predicted categories of the benign sample set and the perturbation sample set. Utilizing the strong correlation between backdoor attacks implanted in the model and perturbation patterns, local features of the input data are transferred to benign samples to obtain a perturbation sample set, and backdoor attack detection is performed based on the consistency of the predicted categories of the benign sample set and the perturbation sample set using a target classification model. This method has strong universality against backdoor attack methods, is not constrained by the attack method, and has high detection accuracy for both input-independent and input-related backdoor attacks. Furthermore, it can directly interface with the running target model, enabling online backdoor attack detection and effectively reducing the dependence of backdoor attack detection on model training and computational overhead.
[0139] Optionally, the perturbation sample generation module 420 includes:
[0140] A local feature extraction unit is used to extract local features from the input data;
[0141] The perturbation sample generation unit is used to replace the original data at the corresponding position of each benign sample in the benign sample set with the local features to obtain the perturbation sample set.
[0142] Optionally, the local feature extraction unit is specifically used for:
[0143] The target classification model is used to classify the input data to obtain the predicted probability of the input data belonging to each predicted category;
[0144] Backpropagation is used to calculate the gradient of the predicted probability of each predicted category with respect to the input data in the feature map of the target classification model;
[0145] The feature map is weighted, normalized, and resized based on the gradient to obtain a heatmap of the input data.
[0146] Identify the masked region in the heat map where the heat value exceeds a first threshold, and obtain the local features of the corresponding pixels of the input data in the masked region.
[0147] Optionally, the backdoor data detection module 440 includes:
[0148] The classification result statistics unit is used to determine the correct classification ratio and misclassification probability distribution of the perturbation sample set based on the predicted categories of the benign sample set and the predicted categories of the perturbation sample set; wherein, the correct classification ratio is the ratio of the number of correctly classified perturbation samples to the total number of data in the perturbation sample set, and the misclassification probability distribution is the probability distribution of misclassified perturbation samples.
[0149] The backdoor data detection unit is used to determine whether the input data is backdoor data based on the correct classification ratio and the incorrect classification probability distribution.
[0150] Optionally, the backdoor data detection unit includes:
[0151] If the correct classification ratio exceeds the second threshold, then the input data is determined not to be backdoor data;
[0152] If the correct classification ratio does not exceed the second threshold, and the highest misclassification probability in the misclassification probability distribution exceeds the third threshold, then the input data is determined to be backdoor data.
[0153] If the correct classification ratio does not exceed the second threshold, and the highest misclassification probability in the misclassification probability distribution does not exceed the third threshold, then the input data is determined not to be backdoor data.
[0154] Optionally, the device further includes:
[0155] The data backdoor attack category determination module is used to determine the predicted category corresponding to the highest misclassification probability as the backdoor attack category of the input data after determining that the input data is backdoor data;
[0156] The backdoor attack category addition module is used to add the backdoor attack category of the input data to the backdoor attack category list.
[0157] The model backdoor attack category determination module is used to determine the backdoor attack category that appears more frequently than a preset frequency in the backdoor attack category list as the backdoor attack category that has been implanted into the target classification model.
[0158] Optionally, the device further includes:
[0159] The prediction category output module is used to, after determining whether the input data is backdoor data, output the predicted category of the input data by the target classification model if the input data is not backdoor data.
[0160] The input data discarding module is used to discard the input data and not output the predicted category of the target classification model for the input data if the input data is backdoor data.
[0161] The backdoor attack detection device provided in this embodiment of the invention can execute the backdoor attack detection method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0162] Example 5
[0163] Figure 5 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0164] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0165] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0166] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as backdoor attack detection methods.
[0167] In some embodiments, the backdoor attack detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the backdoor attack detection method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the backdoor attack detection method by any other suitable means (e.g., by means of firmware).
[0168] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0169] In some embodiments, the backdoor attack detection method may be implemented as a computer program, which is implicitly included in a computer program product. When executed by a processor, the computer program implements the backdoor attack detection method of the present invention. The computer program product can be understood as a software product that primarily implements its solution through a computer program. The computer program used to implement the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer program causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a standalone software package, or entirely on a remote machine or server.
[0170] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0171] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0172] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0173] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0174] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0175] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A backdoor attack detection method, characterized in that, include: Obtain a benign sample set and input data; The local features of the input data are transferred to each benign sample in the benign sample set to obtain the perturbation sample set; The predicted categories of the perturbed sample set and the predicted categories of the benign sample set are determined based on the target classification model. Based on the predicted category of the benign sample set and the predicted category of the perturbation sample set, determine whether the input data is backdoor data.
2. The method according to claim 1, characterized in that, The step of transferring local features of the input data to each benign sample in the benign sample set to obtain a perturbed sample set includes: Extract local features from the input data; The original data at the corresponding position of each benign sample in the benign sample set is replaced with the local features to obtain the perturbed sample set.
3. The method according to claim 2, characterized in that, The extraction of local features from the input data includes: The target classification model is used to classify the input data to obtain the predicted probability of the input data belonging to each predicted category; Backpropagation is used to calculate the gradient of the predicted probability of each predicted category with respect to the input data in the feature map of the target classification model; The feature map is weighted, normalized, and resized based on the gradient to obtain a heatmap of the input data. Identify the masked region in the heat map where the heat value exceeds a first threshold, and obtain the local features of the corresponding pixels of the input data in the masked region.
4. The method according to claim 1, characterized in that, Determining whether the input data is backdoor data based on the predicted category of the benign sample set and the predicted category of the perturbation sample set includes: Based on the predicted categories of the benign sample set and the predicted categories of the perturbation sample set, the correct classification ratio and the misclassification probability distribution of the perturbation sample set are determined; wherein, the correct classification ratio is the ratio of the number of correctly classified perturbation samples to the total number of data in the perturbation sample set, and the misclassification probability distribution is the probability distribution of misclassified perturbation samples. Whether the input data is backdoor data is determined based on the correct classification ratio and the incorrect classification probability distribution.
5. The method according to claim 4, characterized in that, The step of determining whether the input data is backdoor data based on the correct classification ratio and the incorrect classification probability distribution includes: If the correct classification ratio exceeds the second threshold, then the input data is determined not to be backdoor data; If the correct classification ratio does not exceed the second threshold, and the highest misclassification probability in the misclassification probability distribution exceeds the third threshold, then the input data is determined to be backdoor data. If the correct classification ratio does not exceed the second threshold, and the highest misclassification probability in the misclassification probability distribution does not exceed the third threshold, then the input data is determined not to be backdoor data.
6. The method according to claim 5, characterized in that, After determining that the input data is backdoor data, the process also includes: The predicted category corresponding to the highest misclassification probability is determined as the backdoor attack category of the input data; Add the backdoor attack category of the input data to the backdoor attack category list; Backdoor attack categories that appear more frequently than a preset frequency in the backdoor attack category list are identified as the backdoor attack categories into which the target classification model is implanted.
7. The method according to any one of claims 1-6, characterized in that, After determining whether the input data is backdoor data, the process also includes: If the input data is not backdoor data, then the target classification model's predicted category for the input data is output. If the input data is backdoor data, then the input data will be discarded and the target classification model will not output the predicted category of the input data.
8. A backdoor attack detection device, characterized in that, include: The data acquisition module is used to acquire a healthy sample set and input data; The perturbation sample generation module is used to transfer local features of the input data to each benign sample in the benign sample set to obtain the perturbation sample set. The classification prediction module is used to determine the predicted category of the perturbed sample set and the predicted category of the benign sample set based on the target classification model. The backdoor data detection module is used to determine whether the input data is backdoor data based on the predicted category of the benign sample set and the predicted category of the perturbation sample set.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the backdoor attack detection method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the backdoor attack detection method according to any one of claims 1-7.
11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the backdoor attack detection method according to any one of claims 1-7.
Citation Information
Cited By
Backdoor risk detection method and system applied to target detection model
CN121302368A
A backdoor risk detection method and system applied to a target detection model
CN121302368B