Black box back door detection method based on adaptive feature injection

By using an adaptive feature injection black-box backdoor detection method, we screen control samples with large feature differences and perform image fusion to determine the consistency of prediction output. This solves the problems of existing defense methods in terms of diverse backdoor attacks and computationally intensive processing, and achieves efficient detection of poisoned samples.

CN120995452APending Publication Date: 2025-11-21HENAN UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511162690.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing defense methods are inadequate to deal with diverse backdoor attacks, and computationally intensive processing makes deployment in real-time systems and resource-constrained scenarios infeasible.

Method used

A black-box backdoor detection method based on adaptive feature injection is adopted. By acquiring the image sample to be detected and the pre-trained model, the control sample with the largest feature difference is selected for image fusion, and the consistency of the prediction output is compared to determine whether it carries a backdoor trigger.

Benefits of technology

It achieves accurate identification and interception of poisoned samples with a detection rate of 95.2%~86.49%, and has cross-domain generalization ability and robustness against unknown attacks. It is suitable for non-invasive detection of system black-box models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995452A_ABST
    Figure CN120995452A_ABST
Patent Text Reader

Abstract

The invention relates to a black box back door detection method based on adaptive feature injection, which comprises the following steps: acquiring an image sample to be detected and a pre-training model as input, and preparing a data set containing an original test image and a clean sample; and carrying out feature extraction on the image sample by using a deep convolutional neural network. And two clean samples with the maximum feature difference are screened out and are used as contrast samples with the most significant feature difference. And respectively carrying out image fusion on the two obtained contrast samples and the to-be-detected image sample. And inputting a sample image obtained by fusion into the to-be-detected model, and respectively obtaining prediction output results of the fusion sample and the original image. And comparing the semantic consistency of the two prediction results in the previous step, and judging whether the original sample carries the potential backdoor trigger or not. And continuously carrying out the judgment process on the next to-be-detected sample until all the samples are detected. The method provided by the invention is reasonable and feasible, can accurately identify and intercept the poisoning sample, and has wide market prospect and application value.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision and image security, and particularly relates to a black box backdoor detection method based on adaptive feature injection. BACKGROUND

[0002] In recent years, deep learning has been widely applied to many fields such as image classification, natural language processing and pattern recognition, and is an important research direction in the field of artificial intelligence and has attracted much attention. Given a test image, the predicted class can be obtained by calculating the similarity between the image features and the text features of the class description. However, as neural networks become more and more complex, the parameters and depth of the model are constantly increasing, the accuracy is constantly increasing, but the robustness is constantly decreasing, and the security vulnerabilities are also increasing. Therefore, neural networks have inherent vulnerability, which leads to the existence of backdoors.

[0003] Backdoor attack is a kind of security threat with high concealment. Its core mechanism is to implant specific malicious behavior patterns in the model through poisoned training data. In this type of attack, the attacker first designs a backdoor trigger - which can be a specific pixel pattern in the image (such as the local color block used by BadNets), or a special word sequence in natural language. Subsequently, the attacker pollutes the training data set (usually only 1%-5% of the data needs to be polluted) to make the model implicitly establish the association between the trigger and the target output during the learning process. The particularity of this attack is that the model implanted with backdoors behaves the same as the clean model under normal input (the difference in test accuracy is usually less than 0.5%), but once the input contains the preset trigger (such as a specific pixel combination in the corner of the image or a special character sequence in the text), the model will perform the malicious behavior preset by the attacker, such as misclassifying arbitrary input to the target class, generating harmful content, or even leaking private data. More seriously, such triggers can be imperceptible to the human eye and have little impact on the normal function of the model, making traditional anomaly detection-based defense methods face major challenges.

[0004] The current defense research against deep neural network backdoor attacks has two major limitations: first, most existing defense methods are designed for specific types of backdoor attacks and are difficult to cope with the diverse attack methods constantly evolving from attackers. Second, the mainstream defense schemes often require high computational cost, including but not limited to: (1) global parameter adjustment or retraining of pre-trained models, which may consume thousands of GPU hours on large models (such as ViT-Huge); (2) pruning of undesirable neurons, which requires a large number of clean samples that are difficult to obtain in practice. The characteristics of targeted defense and computationally intensive processing seriously restrict the practical deployment feasibility of defense methods in real-time systems and resource-constrained scenarios. SUMMARY

[0005] The defense method against existing backdoor attacks is biased to deal with a specific attack type, or by means of computationally expensive data cleaning and model adjustment, the present application proposes a black-box backdoor detection method based on adaptive feature injection from the perspective of attackers, aiming to reveal the possible trigger signal by checking the input sample pair, to judge whether the deep neural network model is attacked by backdoor, and to realize the accurate identification and interception of the poisoned sample.

[0006] The technical scheme adopted by the present application to solve the above technical problems is: a black-box backdoor detection method based on adaptive feature injection, comprising the following steps, Step (1), obtaining the image sample to be detected and the pre-trained model as input, and preparing a data set containing original test images and clean samples; Step (2), using a pre-trained deep convolutional neural network ResNet to extract features from the image sample in step (1); Step (3), selecting two clean samples with the largest feature contrast as the most significant contrast samples; Step (4), image fusion is performed on the two contrast samples obtained in step (3) and the image sample to be detected; Step (5), inputting the sample image fused in step (4) into the model to be detected, and obtaining the prediction output results of the fused sample and the original image respectively; Step (6), comparing the semantic consistency of the prediction results in step (5) to determine whether the original sample carries a potential backdoor trigger; Step (7), return to step (1), continue to perform the above discrimination process on the next sample to be detected until all samples are detected.

[0007] As a further optimization scheme of the above black-box backdoor detection method based on adaptive feature injection, the specific process of step (1) comprises the following steps, Step (1.1), selecting an image sample to be detected from a target detection or classification task, the sample can be derived from actual image acquisition, network monitoring input, or manually selected suspected backdoor sample; Step (1.2), loading a pre-trained deep neural network model to be detected for feature extraction and prediction of image samples; Step (1.3), preparing a clean sample database to ensure that the samples in it are not tampered with by humans and do not contain triggers; this database can use original samples in public datasets, or ensure its credibility by manual cleaning; Step (1.4), uniform preprocessing of clean samples in the database to adapt to the pre-trained model loaded in step (1.2).

[0008] As a further optimization scheme of the above black-box backdoor detection method based on adaptive feature injection, the specific process of step (2) includes the following steps, Step (2.1), the convolution layer is the basis of image feature extraction, and the convolution operation slides on the image through a filter and performs point multiplication to generate a feature map, Assuming that the input image is I and the filter weight is W, the convolution operation can be represented as: , (1) Where, is the value of the output feature map at position , is the pixel value of the corresponding position of the input image, is the weight of the convolution kernel, is the bias term; Step (2.2), in order to introduce nonlinearity, an activation function ReLU is applied after convolution, and its formula is: , (2) The activation function performs a nonlinear transformation on the output of the convolution layer; Step (2.3), the pooling layer is used to reduce the size of the feature map while preserving important features, and the maximum pooling formula is: , (3) Where, is the pixel value within the pooling window, is the output value after pooling; Step (2.4), at the end of the encoder, the extracted features are flattened into a vector, and the final feature encoding is generated through a fully connected layer, and the calculation formula of the fully connected layer is: , (4) Where, is the input vector, i.e. the flattened feature map, is the weight matrix, is the bias term, is the output feature vector; Step (2.5), feature encoding output, after multiple convolution, pooling and fully connected operations, the image is finally encoded into a feature vector .

[0009] As a further optimization scheme of the above black-box backdoor detection method based on adaptive feature injection, the specific process of step (3) includes the following steps, Step (3.1), the Euclidean distance is used to calculate the difference between the feature vectors, and the two images with the largest difference are found, as follows: , (5) wherein, and are the feature vectors of the two images, and are the first components of the feature vectors, is the dimension of the feature vector; Step (3.2), after the feature extraction operation, the two determined feature contrast samples are defined as clean sample and clean sample , which provide reference for subsequent image fusion.

[0010] As a further optimization scheme of the above-mentioned black box backdoor detection method based on adaptive feature injection, the specific process of step (4) includes the following steps, Step (4.1), using linear interpolation fusion method, pixel-level fusion is performed on the to-be-detected image and each control sample image, and the fusion formula is as follows, I , (6) wherein, denotes the image fusion ratio, which is used to control the weight of the original image and the control image in the fusion result; Step (4.2), define the feature difference degree index D, which is used to measure the distance (such as L2 distance or cosine distance) of F test and reference image F ref in the feature space; The specific calculation method is as follows, D = ||F test -F ref ||2, (7) wherein, F test、 F ref represent the feature vectors of the input image and the reference image, respectively, which come from the intermediate layer of the model or the fused multi-layer features; Step (4.3), according to the difference D, a set of candidate fusion ratios a ∈ {0.3, 0.4,..., 0.7} is set, a fusion image is generated under each a, and is sent to the subsequent model for prediction, the prediction stability and confidence difference are evaluated, and the best fusion ratio a ∗ is finally selected; Step (4.4), using the above optimal fusion ratio a ∗ , the to-be-detected image is fused with the two reference images respectively to generate two final fusion images.

[0011] As a further optimization scheme of the above-mentioned black-box backdoor detection method based on adaptive feature injection, the specific process of step (5) includes the following steps, Step (5.1), input the two images respectively fused in step 4 into the model to be detected, and obtain the prediction results of the two fused images, including classification labels and corresponding confidence information; Step (5.2), input the original image to be detected into the same model at the same time, and obtain the prediction result of the original image; Step (5.3), record the prediction output results of the original image and the two fused images respectively, which are used for subsequent consistency judgment and analysis.

[0012] As a further optimization scheme of the above-mentioned black-box backdoor detection method based on adaptive feature injection, the specific process of step (6) includes the following steps, Step (6.1), compare the semantic consistency of the prediction output results of the original image and the two fused images, and the consistency can be evaluated by judging whether the classification labels are the same and whether the confidence difference exceeds the preset threshold; Step (6.2), if the prediction results of the fused images remain highly consistent with the original image, it is determined that the sample still maintains stable prediction under feature interference, and it is suspected to be a toxic sample carrying a backdoor trigger; Step (6.3), if the prediction results of the fused images and the original image have significant differences, it means that the model is sensitive to feature interference, and it is determined that the sample is a normal sample.

[0013] As a further optimization scheme of the above-mentioned black-box backdoor detection method based on adaptive feature injection, the specific process of step (7) includes the following steps, Step (7.1), after completing the discrimination of the current sample to be detected, return to step (1) to execute the same detection process for the next image sample to be detected; Step (7.2), repeat steps (1) to (6) until all image samples to be detected are detected, and the entire backdoor detection task is completed.

[0014] Compared with the prior art, the beneficial effects of the present application are: The black-box backdoor detection method based on adaptive feature injection of the present application can effectively intercept toxic samples, and the average detection rates on three benchmark data sets are 95.2%, 94.15% and 86.49% respectively. Compared with existing methods, not only the classification accuracy and defense success rate are comparable, but also the cross-domain generalization ability and unknown attack robustness are exhibited, and the advantages are: (1) High adaptability: dynamically adapt the feature fusion weight to different sample structures to improve detection flexibility; (2) Strong generalization: dependent on the model output response characteristics, not limited to specific backdoor forms; (3) Non-invasive: does not change the model structure and training process, suitable for post-hoc security detection of system black-box models. BRIEF DESCRIPTION OF DRAWINGS

[0015] Figure 1 is the basic process schematic diagram of the existing backdoor attack and defense; Figure 2 is the schematic diagram of the construction and detection defense process of the present application; Figure 3 is the visualization schematic diagram of the detection result of the present application. DETAILED DESCRIPTION

[0016] The specific embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0017] In conjunction with the accompanying Figure 1 , 2 , the specific embodiments of the present application are as follows, a black-box backdoor detection method based on adaptive feature injection, comprising the following steps, Step (1), obtaining the image samples to be detected and the pre-trained model as input, preparing a data set containing original test images and clean samples; The specific process of step (1) includes the following steps, Step (1.1), selecting image samples to be detected from a target detection or classification task, the samples can be derived from actual image acquisition, network monitoring input, or manually selected suspected backdoor samples; Step (1.2), loading a pre-trained deep neural network detection model for feature extraction and prediction of image samples; Step (1.3), preparing a clean sample database to ensure that the samples therein are not tampered with by humans and do not contain triggers; the database can use original samples in a public dataset, or ensure its credibility by manual cleaning; Step (1.4), uniformly pre-processing the clean samples in the database to adapt to the pre-trained model loaded in step (1.2).

[0018] Step (2), using a pre-trained deep convolutional neural network ResNet to extract features from the image samples in step (1); The specific process of step (2) includes the following steps, Step (2.1), the convolutional layer is the basis for image feature extraction, the convolution operation slides and performs dot product on the image through a filter to generate a feature map, Assuming that the input image is I and the filter weight is W, the convolution operation can be represented as: , (1) where, is the value of the output feature map at position , is the pixel value at the corresponding position of the input image, is the weight of the convolution kernel, is the bias term; Step (2.2), to introduce nonlinearity, an activation function ReLU is applied after convolution, whose formula is: , (2) The activation function performs a nonlinear transformation on the output of the convolution layer; Step (2.3), the pooling layer is used to reduce the size of the feature map while preserving important features, and the formula for maximum pooling is: , (3) where, is the pixel value within the pooling window, is the output value after pooling; Step (2.4), at the end of the encoder, the extracted features are flattened into a vector, and the final feature encoding is generated through a fully connected layer, and the calculation formula of the fully connected layer is: , (4) where, is the input vector, i.e., the flattened feature map, is the weight matrix, is the bias term, is the output feature vector; Step (2.5), the feature encoding output is performed, after multiple convolution, pooling and fully connected operations, the image is finally encoded into a feature vector .

[0019] Step (3), two clean samples with the largest feature contrast are selected as the most significant control samples with the largest feature difference; The specific process of step (3) includes the following steps, Step (3.1), use the Euclidean distance to calculate the difference between the feature vectors, find the two images with the largest difference, the formula is as follows: , (5) where, and are the feature vectors of the two images, and are the th component in the feature vector, is the dimension of the feature vector; Step (3.2), after the feature extraction operation, the two determined feature contrast samples are defined as clean samples and clean samples , which provide reference for subsequent image fusion.

[0020] Step (4), the two control samples obtained in step (3) and the image sample to be detected are respectively subjected to image fusion; The specific process of step (4) includes the following steps, Step (4.1), using a linear interpolation fusion method, the image to be detected and each control sample image are subjected to pixel-level fusion, and the fusion formula is as follows, I , (6) Wherein, represents the image fusion ratio, which is used to control the weight of the original image and the control image in the fusion result; Step (4.2), define the feature difference degree index D, which is used to measure the distance (such as L2 distance or cosine distance) of F test and reference image F ref in the feature space; The specific calculation method is as follows, D = ||F test -F ref ||2, (7) Wherein, F test、 F ref respectively represent the feature vectors of the input image and the reference image, which come from the intermediate layer of the model or the fused multi-layer features; Step (4.3), according to the difference degree D, a set of candidate fusion ratios a ∈ {0.3, 0.4,..., 0.7} is set, a fusion image is generated under each a, and is sent to the subsequent model for prediction, to evaluate the prediction stability and confidence difference, and finally the best fusion ratio a ∗ is selected. Step (4.4), using the above optimal fusion ratio a ∗ , the image to be detected and the two reference images are fused respectively to generate two final fusion images.

[0021] Step (5), input the sample image fused in step (4) into the detection model to obtain the prediction output results of the fused sample and the original image respectively; The specific process of step (5) includes the following steps, Step (5.1), input the two images fused in step 4 into the detection model to obtain the prediction results of the two fusion images, including the classification label and the corresponding confidence information; Step (5.2), input the original image to be detected into the same model at the same time to obtain the prediction result of the original image; Step (5.3), record the prediction output results of the original image and the two fusion images respectively for subsequent consistency judgment analysis.

[0022] Step (6), compare the semantic consistency of the prediction results in step (5) to determine whether the original sample carries a potential backdoor trigger; The specific process of step (6) includes the following steps, Step (6.1), compare the semantic consistency of the prediction output results of the original image and the two fusion images, which can be evaluated by judging whether the classification labels are the same and whether the confidence difference exceeds the preset threshold; Step (6.2), if the prediction results of the fusion images remain highly consistent with the original image, it is determined that the sample still maintains stable prediction under feature interference, and is suspected to be a toxic sample carrying a backdoor trigger; Step (6.3), if the prediction results of the fusion images have significant differences with the original image, it indicates that the model is sensitive to feature interference, and the sample is determined to be a normal sample.

[0023] Step (7), return to step (1) and continue to perform the above discrimination process on the next sample to be detected until all samples are detected.

[0024] The specific process of step (7) includes the following steps, Step (7.1), after completing the discrimination of the current sample to be detected, return to step (1) and perform the same detection process on the next image sample to be detected; Step (7.2), repeat steps (1) to (6) until all image samples to be detected are detected, and complete the entire backdoor detection task.

[0025] In actual application, as shown in Figure 1 , Figure 1This is a basic flowchart illustrating existing backdoor attacks and defenses. Backdoor attacks, as a highly covert security threat, work by poisoning training data to implant specific malicious behavior patterns into the model. In these attacks, attackers first meticulously design backdoor triggers—this could be specific pixel patterns in an image (such as local color blocks used in BadNets) or special word sequences in natural language. Then, by polluting the training dataset (usually only 1%-5% of the data), the attacker implicitly establishes a correlation between the trigger and the target output during the model's learning process. The unique aspect of this attack is that the model with the implanted backdoor performs identically to a clean model under normal input (the difference in test accuracy is typically less than 0.5%). However, once the input contains a pre-set trigger (such as a specific pixel combination in an image corner or a special character sequence in text), the model will execute the attacker's pre-set malicious behavior, such as misclassifying arbitrary input into the target category, generating harmful content, or even leaking private data. More seriously, these triggers can be imperceptible to the human eye and have minimal impact on the model's normal function, posing a significant challenge to traditional anomaly detection-based defense methods.

[0026] Current research on defense against backdoor attacks on deep neural networks faces two major limitations: First, most existing defense methods are designed for specific types of backdoor attacks, making it difficult to cope with the ever-evolving and diverse attack methods employed by attackers. Second, mainstream defense solutions often require high computational costs, including but not limited to: (1) global parameter adjustment or retraining of pre-trained models, which may consume thousands of GPU hours on large models (such as ViT-Huge); and (2) pruning of malicious neurons, for which a large number of clean samples are difficult to obtain in practice. This targeted defense and computationally intensive processing characteristic severely restricts the practical deployment feasibility of defense methods in real-time systems and resource-constrained scenarios.

[0027] like Figure 2 As shown, Figure 2 This is a schematic diagram of the construction and detection defense process of the present invention; the specific operation and implementation process has been described in detail above and will not be repeated here.

[0028] like Figure 3 As shown, Figure 3 This is a visualization of the detection results of this invention. When performing detection on a black-box classification model trained on the CIFAR-10 dataset, the user first inputs the sample to be detected ( Figure 3and the model, according to the image to be detected, a certain standard image classification data set is selected, 1-3 clean images are selected as control sample screening library for each category. All images in the detection sample and the control sample screening library are input into the ResNet model in turn, the feature representation of each intermediate layer is extracted, and the feature vector distance calculation is performed with the samples in the clean image database based on the feature representation, and the two reference images farthest from the image to be detected in the feature space are selected as the control sample pair (automobile and cat) in the control sample screening library Figure 3 and the model, according to the image to be detected, a certain standard image classification data set is selected, 1-3 clean images are selected as control sample screening library for each category. All images in the detection sample and the control sample screening library are input into the ResNet model in turn, the feature representation of each intermediate layer is extracted, and the feature vector distance calculation is performed with the samples in the clean image database based on the feature representation, and the two reference images farthest from the image to be detected in the feature space are selected as the control sample pair (automobile and cat) in the control sample screening library Figure 3 and the model, according to the image to be detected, a certain standard image classification data set is selected, 1-3 clean images are selected as control sample screening library for each category. All images in the detection sample and the control sample screening library are input into the ResNet model in turn, the feature representation of each intermediate layer is extracted, and the feature vector distance calculation is performed with the samples in the clean image database based on the feature representation, and the two reference images farthest from the image to be detected in the feature space are selected as the control sample pair (automobile and cat) in the control sample screening library Figure 3 and the model, according to the image to be detected, a certain standard image classification data set is selected, 1-3 clean images are selected as control sample screening library for each category. All images in the detection sample and the control sample screening library are input into the ResNet model in turn, the feature representation of each intermediate layer is extracted, and the feature vector distance calculation is performed with the samples in the clean image database based on the feature representation, and the two reference images farthest from the image to be detected in the feature space are selected as the control sample pair (automobile and cat) in the control sample screening library Figure 3 and the model, according to the image to be detected, a certain standard image classification data set is selected, 1-3 clean images are selected as control sample screening library for each category. All images in the detection sample and the control sample screening library are input into the ResNet model in turn, the feature representation of each intermediate layer is extracted, and the feature vector distance calculation is performed with the samples in the clean image database based on the feature representation, and the two reference images farthest from the image to be detected in the feature space are selected as the control sample pair (automobile and cat) in the control sample screening library Figure 3 and the model, according to the image to be detected, a certain standard image classification data set is selected, 1-3 clean images are selected as control sample screening library for each category. All images in the detection sample and the control sample screening library are input into the ResNet model in turn, the feature representation of each intermediate layer is extracted, and the feature vector distance calculation is performed with the samples in the clean image database based on the feature representation, and the two reference images farthest from the image to be detected in the feature space are selected as the control sample pair (automobile and cat) in the control sample screening library

[0029] In practical applications, the present application is suitable for a variety of black box scenes without accessing model parameters, especially suitable for deployment in cloud inference API and large model inference system. The method does not need to train a substitute model, does not need label intervention, and can realize automatic detection of backdoor samples through feature level fusion and output consistency judgment, which has high universality and practical value.

[0030] The method is reasonable and feasible, different sample structures are adapted through dynamic feature fusion weight, detection flexibility is improved, model output response features are relied on, specific backdoor forms are not limited, model structure and training process are not changed, the method is suitable for post-hoc security detection of a system black box model, precise identification and interception of toxic samples are realized. The method solves the defects that the existing backdoor attack defense method is biased to process specific attack types, or the method exists through high-cost cleaning data and model adjustment, and has wide market prospect and application value.

[0031] The preferred specific embodiments and examples of the application are described in detail above in conjunction with the accompanying drawings, but the application is not limited to the above-described embodiments and examples, and various changes can be made within the knowledge of those skilled in the art without departing from the concept of the application.

Claims

1. A black-box backdoor detection method based on adaptive feature injection, characterized in that: Includes the following steps, Step (1): Obtain the image samples to be detected and the pre-trained model as input, and prepare a dataset containing the original test images and clean samples; Step (2): Use a pre-trained deep convolutional neural network ResNet to extract features from the image samples in step (1); Step (3): Select the two clean samples with the largest feature contrast as the control samples with the most significant feature differences; Step (4): Perform image fusion on the two control samples obtained in step (3) and the image sample to be detected respectively; Step (5): Input the sample image obtained by fusing in step (4) into the model to be detected, and obtain the prediction output results of the fused sample and the original image respectively; Step (6): Compare the semantic consistency of the two prediction results in step (5) to determine whether the original sample carries a potential backdoor trigger. Step (7), return to step (1), and continue to perform the above discrimination process on the next sample to be tested until all samples are tested.

2. The black-box backdoor detection method based on adaptive feature injection as described in claim 1, characterized in that: The specific process of step (1) includes the following steps: Step (1.1): Select image samples to be detected from the target detection or classification task. The samples can be from actual acquired images, network monitoring input, or manually selected suspected backdoor samples. Step (1.2): Load a pre-trained deep neural network model to be detected, which is used for feature extraction and prediction of image samples; Step (1.3): Prepare a clean sample database, ensuring that the samples in the database have not been tampered with and do not contain triggers; the database can be made from the original samples in the public dataset, or its credibility can be ensured by manual cleaning; Step (1.4) involves uniformly preprocessing the clean samples in the database to adapt them to the pre-trained model loaded in step (1.2).

3. The black-box backdoor detection method based on adaptive feature injection as described in claim 1, characterized in that: The specific process of step (2) includes the following steps: Step (2.1): The convolutional layer is the foundation of image feature extraction. The convolution operation slides a filter across the image and performs dot products to generate a feature map. Assuming the input image is I and the filter weights are W, the convolution operation can be represented as: ,(1) in, It is the position in the output feature map. The value, It is the pixel value at the corresponding position in the input image. These are the weights of the convolution kernel. It is a bias term; Step (2.2), in order to introduce nonlinearity, the ReLU activation function is applied after convolution, and its formula is: ,(2) The activation function performs a non-linear transformation on the output of the convolutional layer; Step (2.3): The pooling layer is used to reduce the size of the feature map while retaining important features. The max pooling formula is: ,(3) in, These are the pixel values ​​within the pooling window. This is the output value after pooling; Step (2.4): At the end of the encoder, the extracted features are flattened into a vector and then passed through a fully connected layer to generate the final feature encoding. The calculation formula for the fully connected layer is: ,(4) in, It is the input vector, i.e., the flattened feature map. It is a weight matrix. It is a bias term. It outputs the feature vector; Step (2.5) involves feature encoding output. After multiple convolutional, pooling, and fully connected operations, the image is finally encoded into a feature vector. .

4. The black-box backdoor detection method based on adaptive feature injection as described in claim 1, characterized in that: The specific process of step (3) includes the following steps: Step (3.1) uses Euclidean distance to calculate the difference between feature vectors and finds the two images with the largest difference, as shown in the following formula: ,(5) in, and These are the feature vectors of the two images. and It is the eigenvector of the eigenvector. One portion, It is the dimension of the feature vector; In step (3.2), after feature extraction, the two contrasting feature samples are defined as clean samples. and clean samples This provides a reference for subsequent image fusion.

5. The black-box backdoor detection method based on adaptive feature injection as described in claim 1, characterized in that: The specific process of step (4) includes the following steps: Step (4.1) uses linear interpolation to fuse the image to be detected. Pixel-level fusion is performed with each control sample image, using the following fusion formula. 𝐼 ,(6) in, This indicates the image fusion ratio, used to control the weights of the original image and the control image in the fusion result; Step (4.2) defines the feature difference index D, which is used to measure F test With reference image F ref The distance in the feature space (e.g., L2 distance or cosine distance); the specific calculation method is as follows. D=||F test -F ref ||2,(7) Among them, F test、 F ref These represent the feature vectors of the input image and the reference image, respectively. These feature vectors come from intermediate layers of the model or are fused from multiple layers of features. Step (4.3): Based on the difference degree D, a set of candidate fusion ratios α∈{0.3,0.4,.....0.7} is set. A fused image is generated under each α and fed into the subsequent model for prediction. The prediction stability and confidence difference are evaluated, and finally the best fusion ratio α is selected. ∗ ; Step (4.4), using the above-mentioned optimal fusion ratio α ∗ The image to be detected is fused with two reference images to generate two final fused images.

6. The black-box backdoor detection method based on adaptive feature injection as described in claim 1, characterized in that: The specific process of step (5) includes the following steps: Step (5.1): Input the two images obtained by fusing in step 4 into the detection model to obtain the prediction results of the two fused images, including the classification label and the corresponding confidence information; Step (5.2) involves simultaneously inputting the original image to be detected into the same model to obtain the prediction result of the original image; Step (5.3) records the prediction output results of the original image and the two fused images respectively, for subsequent consistency judgment analysis.

7. The black-box backdoor detection method based on adaptive feature injection as described in claim 1, characterized in that: The specific process of step (6) includes the following steps: Step (6.1) compare the semantic consistency of the predicted output results of the original image and the two fused images. The consistency can be evaluated by judging whether the classification labels are the same and whether the confidence difference exceeds a preset threshold. Step (6.2): ​​If the prediction result of the fused image is highly consistent with the original image, it is determined that the sample still maintains stable prediction under feature interference and is suspected to be a toxic sample carrying a backdoor trigger. Step (6.3): If the prediction result of the fused image is significantly different from the original image, it indicates that the model is sensitive to feature interference, and the sample is determined to be a normal sample.

8. The black-box backdoor detection method based on adaptive feature injection as described in claim 1, characterized in that: The specific process of step (7) includes the following steps: Step (7.1): After completing the discrimination of the current sample to be detected, return to step (1) and perform the same detection process on the next image sample to be detected; Step (7.2): Repeat steps (1) to (6) until all image samples to be detected have been detected, thus completing the entire backdoor detection task.

Citation Information

Cited By

  • Neural network backdoor sample and backdoor model detection method

    CN122114057A