A backdoor model detection method guided by graph feature vectors
By modeling the machine learning model into a graph structure and calculating the graph feature vector, a binary classification model is constructed for backdoor model detection, the problems of low efficiency and insufficient accuracy of backdoor model detection in the existing technology are solved, and efficient and accurate backdoor model detection is achieved.
Patent Information
- Application Number
- CN202210904446.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-07-29
AI Technical Summary
It is difficult for the prior art to efficiently detect backdoor models with triggers, especially when detecting backdoor inputs during model use, traditional detectors have high time complexity and high sample requirements.
By modeling machine learning models into graph structures, calculating graph feature vectors, and using these feature vectors to build a binary classification model for backdoor model detection, efficient detection is achieved.
It improves the efficiency and accuracy of backdoor model detection, reduces dependence on backdoor data, and provides the interpretability and robustness of the model.
Smart Images

Figure CN115248916B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of machine learning, and particularly relates to a backdoor model detection method guided by graph feature vectors. Background Art
[0002] Machine learning models are increasingly being deployed in industrial, medical, and other practical scenarios to assist or replace human experts in making decisions on critical tasks, such as in the fields of computer vision, disease diagnosis, financial fraud detection, defense against malware and cyberattacks, access control, surveillance, etc. However, the security of machine learning system deployment is now considered a real security issue. Due to the training process of some complex machine learning algorithms, especially deep learning methods based on large datasets, relying on high-performance computing resources and a large amount of storage media. Machine learning models can be trained (e.g., outsourced) and provided (e.g., pre-trained models) by third parties, which provides opportunities for small and medium-sized enterprises and some research institutions with insufficient computing resources to use machine learning methods. Conversely, it also provides opportunities for attackers to manipulate training data and / or models. Recent work has shown that this type of stealthy attack allows attackers to insert backdoors or Trojan horses into the model. The generated backdoor model behaves normally for benign input samples; however, when the input is marked with a trigger determined by the attacker and only known to the attacker, the backdoor model will exhibit abnormal behavior, e.g., classifying the input sample into a target class preset by the attacker.
[0003] A significant feature of backdoor attacks is that they are easily achievable in the physical world, especially in visual systems. Such backdoor attacks are characterized by being simple, efficient, robust, and easy to implement, for example, by placing a trigger on an object in a visual scene. This differentiates it from other attacks, especially adversarial attacks. In adversarial attacks, the attacker cannot fully control the conversion of the physical scene into a valid input containing adversarial digital noise; the perturbations in adversarial attacks are very small, e.g., single-pixel adversarial example attacks in image adversarial attacks. Therefore, due to the limitations of sensor performance, the camera may not be able to perceive such perturbations. To improve the effectiveness of attacks in the real world, backdoor attacks usually employ unbounded perturbations to ensure the robustness of the attack against physical effects such as viewpoints, distances, and lighting when converting physical objects into backdoor inputs. Usually, the trigger is perceptible to humans, but human perception is often irrelevant because machine learning models are usually deployed in autonomous environments without human intervention, unless the system flags an anomaly or raises an alarm. The trigger can also be unobtrusive and regarded as a natural part of the image, not malicious and disguised in many cases: for example, a pair of sunglasses on a face or graffiti in a visual scene.
[0004] Since backdoor triggers are hidden tools protected and exploited by attackers, detecting such backdoor inputs with triggers is a challenge, especially when detecting backdoor inputs during the use of the model. Traditional backdoor detectors can usually only detect the probability of a target model having a backdoor vulnerability with a large number of test samples before the model goes online. They have high requirements for test samples and a high time complexity in the detection process.
[0005] The machine learning training process based on the cloud or third-party outsourcing is vulnerable to backdoor attacks by malicious attackers. Usually, users provide training data and define the model architecture, and attackers always have the opportunity to inject hidden triggers into the training data. This makes the trained model classify these samples containing hidden triggers as the attacker's target class, while for clean samples, the classification accuracy of the backdoor model hardly decreases. Existing detection methods for backdoor models usually require a large number of test samples. Through the training of a complex detection model, the confidence of the existence of a backdoor is calculated for each sample, and the probability of the original model being attacked by a backdoor is deduced inversely. In actual scenarios, it is difficult to obtain the type of trigger, and the number of samples available to the detection party is limited. Summary of the Invention
[0006] Aiming at the deficiencies of the existing technology, the present invention starts from the internal mechanism of the model, utilizes the powerful representation ability of graph data, models the machine learning model into a graph structure, calculates the basic graph metrics for the model graph, calculates the feature vectors related to backdoor detection, and uses the calculated feature vectors to achieve the efficient detection of backdoor models, providing guarantee for the secure use of machine learning models.
[0007] The object of the present invention is achieved by the following technical solutions:
[0008] A backdoor model detection method guided by graph feature vectors, comprising the following steps:
[0009] Step 1: Construct a model set that mixes backdoor models and clean models;
[0010] (1.1) Generate various types of backdoor triggers;
[0011] (1.2) Add the backdoor triggers to normal training samples to form various types of backdoor samples, and train a deep neural network for classification with different poisoning ratios. After training, a backdoor model is obtained;
[0012] (1.3) Mix different types of backdoor models and clean models to form a model set;
[0013] Step 2: Construct a model graph based on classification categories;
[0014] (2.1) For each category in the training set, select several samples and input them into each model in the model ensemble. Use the output value of the model node as one of the attributes of the node. Define the node with the neuron output value greater than or equal to the threshold as an activated node, and the node with the output value less than the threshold as a non-activated node;
[0015] (2.2) Use the output value on the activated node as one of the attributes of the activated node, and set the attributes of the non-activated nodes to 0 to achieve the embedding of activation information;
[0016] (2.3) For the activated nodes, extract the weights between them and all the activated nodes in the previous layer; for the non-activated nodes, extract the maximum value of the weights between them and all the nodes in the previous layer; A single input sample and a single model form a model graph;
[0017] Step Three: Generate the graph features of the model graph;
[0018] (3.1) Calculate multiple basic metrics for each model graph respectively;
[0019] (3.2) Concatenate the multiple basic metrics of each model graph into a vector as the graph feature of the model graph;
[0020] Step Four: Combine the graph features of all the model graphs and their labels indicating whether they are backdoor models to form a training set for training a binary classification model. The trained binary classification model is called a backdoor model detector;
[0021] Step Five: For any model to be detected composed of a deep neural network, repeat Step Two and Step Three to obtain the graph feature corresponding to the model to be detected. Input the graph feature into the backdoor model detector, and the backdoor model detector outputs the category indicating whether the model to be detected is a backdoor model or a normal model.
[0022] Further, in the step (2.2), after the activation information is embedded, according to the category to which each input sample belongs, assign the same category attribute as the input sample to the activated nodes, and assign an additional attribute with a category different from all the samples in the training set to the non-activated nodes.
[0023] Further, the basic metrics include global average degree, modularity, graph density, and average path length.
[0024] Further, the threshold in the step (2.1) is 0.5.
[0025] Further, in the step (1.3), both the backdoor model and the clean model are the models saved every time the training samples are trained for one round.
[0026] Further, in the fifth step, after the backdoor model detector outputs the category, the detection results of the models obtained from the same model structure and dataset training are statistically analyzed. According to the voting mechanism, the backdoor model or clean model that appears the most in the detection results is selected as the final detection result.
[0027] The beneficial effects of the present invention are as follows:
[0028] (1) Model the machine learning model with graph data, analyze the characteristics of the backdoor model from the perspective of the model mechanism, and improve the interpretability of the model;
[0029] (2) Expand the graph signature features with the intermediate results of the model during the training process, make full use of the model parameter information during the training process, and reduce the dependence on the backdoor data;
[0030] (3) The backdoor model detector uses a simple binary classification feature classifier to achieve efficient backdoor model detection. Description of the Drawings
[0031] Figure 1 It is a schematic diagram of the overall process of the method of the present invention.
[0032] Figure 2 It is a schematic diagram of the backdoor sample with a trigger generated in the method of the present invention.
[0033] Figure 3 It is a schematic diagram of the construction process of the model graph in the method of the present invention. Specific Embodiments
[0034] The present invention will be described in detail below according to the drawings and preferred embodiments. The purpose and effects of the present invention will become more apparent. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0035] The technical concept of the present invention is as follows: In the scenarios of cloud outsourcing training of deep learning models and third-party provision of pre-trained models, an attacker can add various types of triggers to the training data, label the backdoor samples with triggers as the class labels of the target class, and conduct backdoor training to achieve the purpose of injecting backdoor attacks into the model. Based on this insecure scenario, the present invention utilizes the high representability of graph data types, models the machine learning model as graph data, and uses the global metrics of the graph to guide the detection of backdoor models, so as to achieve the purpose of enhancing the model's robust security against backdoor attacks. First, construct different types of backdoor models using various backdoor training methods; then extract the key neurons and neural pathways in the model using the idea of neural pathways to construct a model graph; by splicing the global graph metrics at different training times and of different types, construct a graph signature feature vector; train a basic binary classification backdoor model detector with the graph signature vector; finally, activate the samples obtained in step 1) with the samples containing backdoor triggers and normal samples respectively, input the obtained activation features into the detector, and judge whether the target model is a backdoor model according to the voting mechanism, ultimately achieving the efficient detection of the target model.
[0036] Referring to Figures 1 to 3 , the method for detecting poisoned models guided by graph feature vectors of the present invention is as follows:
[0037] Step 1: Construct a model set that mixes backdoor models and clean models, and its algorithm steps are as follows:
[0038] (1.1) Generate various types of backdoor triggers;
[0039] To obtain a dataset for poisoning training of the target model, it is necessary to add patterns related to poisoning attacks to the original clean data to process some samples into poisoned samples. Specifically, in this embodiment, two forms of triggers are added to the MNIST dataset, namely pixel patches and bar watermarks.
[0040] (1.2) Add the backdoor triggers to normal training samples to form various types of backdoor samples, and train a deep neural network for classification with different poisoning ratios. After training, obtain a backdoor model;
[0041] For each category of handwritten digit images, 30% of the image samples are respectively selected as the test set, and the remaining samples are used as the training set to train the model. The purpose of each poisoned sample is to classify the target class of handwritten digit 0 in the attack as a handwritten digit of class 1; in each backdoor training, only one combination of the original class and the target class is selected; when normal samples are input into the model after backdoor training, the output is the class label value of normal judgment; the corresponding backdoor samples with triggers are input into the backdoor model, and the output is the target handwritten digit category of the backdoor attack; after the training is completed, the neural network model structure and the weight information of each layer are saved.
[0042] For both backdoor training and normal training, an early stopping strategy of 10 rounds is set, that is, if there is no improvement in the classification accuracy of the model during the 10-round training, or the model completes 50 rounds of training, the training stops. The clean model and the backdoor model are respectively trained using the SGD and Adam optimization methods. The target categories of the backdoor model are sequentially set from 0, 1, 2... to 9 for traversal. When training the backdoor model, the proportions of backdoor samples can be respectively set to 10%, 30%, and 50%. Through different parameter settings, several model samples with different structures are trained for training the detector.
[0043] (1.3) Mix different types of backdoor models and clean models to form a model set;
[0044] Step Two: Construct a model graph based on classification categories;
[0045] (2.1) For each category in the training set, select several samples and input them into each model in the model set. Take the output value of the model node as one of the attributes of the node. Define the node with the neuron output value greater than or equal to the threshold as an activated node, and the node with the output value less than the threshold as a non-activated node;
[0046] (2.2) Take the output value on the activated node as one of the attributes of the activated node, and set the attributes of the non-activated nodes to 0 to achieve the embedding of activation information;
[0047] Specifically, take the activation value on the activated node as the attribute of the corresponding node, and set the attributes of the non-activated nodes to 0. The activation threshold is set to 0.5 times the maximum value of the weights of this layer. Nodes with weights exceeding this threshold are regarded as activated nodes, and vice versa as non-activated nodes. The activation value is the value obtained by normalizing the actual weight of the node according to the maximum value of the weights of this layer. In addition to the activation situation, the node attributes also include the classification characteristics of the corresponding machine learning model on the dataset. Input data of a specific category into the model, and multiply the pixel values of the samples with category labels by the weights of the model. The returned result is the activation attribute of a type of sample. Set the attributes of non-activated nodes to other categories. For a 10-class machine learning model, the node activation attribute consists of an 11-dimensional vector.
[0048] (2.3) For the activated nodes, extract the weights between them and all the activated nodes in the previous layer; for the non-activated nodes, extract the maximum value of the weights between them and all the nodes in the previous layer; a single input sample and a single model form a model graph;
[0049] According to the fully connected layers of each model trained in step (1.3), extract the key connecting edges for the activated nodes and the non-activated nodes respectively. As Figure 3 (b) shows, for the activated nodes, retain the weights between the activated nodes and all the activated nodes in the previous layer, and use them as the weights on the input connecting edges. For the non-activated nodes, only retain the one with the largest absolute value of the weights of all the nodes in the previous layer (i.e., the input nodes) as the weight value of the input connecting edge to achieve the sparsification of the connecting edges. Establish the corresponding graph structure according to the existing model structure. Specifically, take a handwritten digit picture sample as an example. After inputting the sample into the image recognition model, the weights of the sample activation are obtained after multiplying the sample data by the weights of each layer of neurons. For the structure of the fully connected layer, map each neuron of the model to a node in the graph structure, and use the activation weights of the model samples of each layer as the connecting edge weights of the graph structure. For the case where the model activation weight is positive, directly use the model weight as the graph connecting edge weight; for the case where the model activation weight is negative, calculate the norm of the weight as the graph connecting edge weight. For the structure of the convolutional layer, first expand the convolutional layer, and then simplify the convolutional graph structure to the case of the fully connected layer. As Figure 3 (a) shows, define the nodes and edges of the graph, and represent the convolutional operation as matrix multiplication. Specifically, for the input image X and the convolutional kernel The nodes of the l-th layer are defined as the output of the mapped features of the input of each filter f i Connect each of these nodes to the corresponding nodes of the output neurons in the next layer, and these connecting edges are weighted by the activation values of the neurons (i.e., the input image at a specific stride multiplied by the filter value at that position in the image). The undirected graph constructed through this process is the graph structure corresponding to the model structure, and the size of the connecting edges of the graph corresponds to the norm value of the model weight size; the graph nodes of the fully connected layer correspond one-to-one with the neurons of the model, and the graph nodes of the convolutional layer correspond to the convolutional kernels of the model.
[0050] Aiming at the problem that the complex network model structure is too complex, this method further optimizes the internal information of the graph structure by selecting appropriate neural pathways. While simplifying the graph structure features, it ensures that the graph structure features corresponding to different categories and poisoned or not poisoned models still have a high degree of diversity. Specifically, take the image recognition model based on the fully connected network as an example. Input the normal samples into the clean model and the poisoned model trained in step 1) in sequence, and record the output of each layer as:
[0051] Output l =Layerl (Output l-1 )
[0052] Among them, Layer l represents the layer function of the l-th layer, and Output l is the output matrix returned by inputting the parameters of the previous layer into the layer function of the l-th layer. Among them, denote the output of the sample at the l-th layer of the poisoned model as the output of the sample at the l-th layer of the clean model as Using 0.5 as the threshold to divide normal activation values and abnormal activation values, that is, taking the absolute value of the maximum output weight of each layer as the benchmark, and using a value of 0.5 times the benchmark value as the threshold to divide the neurons of each layer. Neurons above this threshold are activated neurons, and those below the threshold are unactivated neurons. The path between activated neurons is used as the activation path, and all weights and connection information of the activation path are saved; the path between inhibitory neurons is used as the inhibitory path, and all weights and connection relationships related to the inhibitory path are deleted; the path between activated neurons and inhibitory neurons is used as the semi-activated path, and all semi-activated paths connected to a certain inhibitory neuron in the output layer are calculated, and only the path with the largest weight norm is retained as the effective activation path. The undirected graph constructed between the finally retained activated neurons and the effective activation paths is the optimized model graph structure.
[0053] Step 3: Generate the graph features of the model graph;
[0054] (3.1) Calculate multiple basic metrics of each model graph respectively;
[0055] This method mainly calculates four metrics: graph clustering coefficient, average node degree, average path length, and modularity. Among them, the graph clustering coefficient C is used to measure the degree of aggregation of nodes in the graph. For a node i on the graph, e i represents the number of edges between the neighbors of node i, and k i represents the number of neighbor nodes of node i. Then the calculation method of graph clustering is:
[0056]
[0057] The node degree refers to the number of connections of a certain node in the graph with other nodes. For the directed graph constructed by this method, the in-degree and out-degree of the node need to be calculated simultaneously. The average degree of the graph refers to the average value of the degrees of all nodes in the graph. For a graph with n nodes, the degree of each node is d k , then the calculation process of the average degree is:
[0058]
[0059] The average path length of a graph refers to the average of the distances of the edges connecting any two nodes. For a weighted directed graph, the average path length of the graph is the weighted average of the directed distances between any nodes. The average path length of the graph is also called the characteristic path length or average distance of the graph, that is:
[0060]
[0061] where N is the number of nodes in the graph.
[0062] Graph modularity is a clustering algorithm that clusters a class of nodes with the same characteristics together to form clusters, and unsupervised divides the nodes in the graph into clusters. The clustering characteristics of the nodes are strongly correlated with the classification effects of the corresponding categories.
[0063] The relevant indicators of the optimized model graph are calculated by the above method, and the intermediate process in the poisoning training process is considered to enrich the graph index feature vector. Specifically, during the training process, the model structure in the intermediate process is selected every 5 training rounds to calculate the graph index. And the intermediate index vectors calculated for all training rounds of each model are concatenated into a graph feature matrix, and finally a graph feature matrix is constructed for each model as the input feature of the poisoning vulnerability tester.
[0064] (3.2) Concatenate multiple basic indicators of each model graph into a vector as the graph feature of the model graph;
[0065] Specifically, during the training process, the model structure in the intermediate process is selected every 5 training rounds to calculate the graph index. And the intermediate index vectors calculated for all training rounds of each model are concatenated into a graph feature matrix, and finally a graph feature matrix is constructed for each model as the input feature of the poisoning vulnerability tester.
[0066] Step 4: Combine the graph features of all models and their labels indicating whether they are backdoor models to form a training set for training a binary classification model. The trained binary classification model is called a backdoor model detector;
[0067] The graph feature matrix calculated in Step 3 is used as the input feature of the vulnerability evaluation device. Specifically, first construct a basic fully connected network as the basic model of the vulnerability tester. A common fully connected classifier maps each output to a different binary classification, corresponding to a robust model and a vulnerable model respectively, and the output layer consists of a shared softmax function.
[0068] The trained evaluator can be used to evaluate the vulnerability of the model. After the model activated by the samples is transformed into a graph structure, the graph metric feature vector can be quickly calculated to complete the evaluation. Compared with the traditional vulnerability testing methods, the diversity of the graph structure can be utilized to construct a large number of model graphs with only a small number of normal samples and test the poisoned model by using the feature matrix of the activated model graph. Since the optimized graph structure can better represent the structural features of the corresponding model, the poisoned vulnerability testing of the deep model can be realized without relying on specific poisoned samples, effectively improving the applicable range of the poisoned testing.
[0069] Step Five: For any model to be detected composed of a deep neural network, repeat Step Two and Step Three to obtain the graph features corresponding to the model to be detected, and input the graph features into the backdoor model detector. The backdoor model detector outputs the category of the model to be detected as a backdoor model or a normal model.
[0070] Construct model graphs for the normal model trained in Step One and the backdoor model trained with backdoor samples containing triggers by the methods in Step Two and Step Three in sequence, calculate the graph features at the intermediate moments of the training process for each model, and input them into the detector trained in Step Four. The detector uses the graph signature features for binary classification and divides each feature into a normal model graph and a backdoor model graph. Count the classification results of all the model graphs in the training process. The model corresponding to the backdoor model graph with more than half of the discriminants is detected as a suspicious backdoor model.
[0071] Those of ordinary skill in the art can understand that the above are only preferred examples of the invention and are not used to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, those skilled in the art can still modify the technical solutions described in the foregoing examples or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, etc. made within the spirit and principle of the invention shall be included within the protection scope of the invention.
Claims
1. A backdoor model detection method guided by graph feature vectors, characterized in that, Step 1: Construct a model set that mixes backdoor models and clean models; (1.1) Generate various types of backdoor triggers; (1.2) Add the backdoor triggers to normal training samples to form various types of backdoor samples, and train a deep neural network for classification with different poisoning ratios. After training, a backdoor model is obtained; (1.3) Mix different types of backdoor models and clean models to form a model set; Step 2: Construct a model graph based on classification categories; (2.1) For each category in the training set, select samples and input them into each model in the model set. Take the output value of the model node as one of the attributes of the node. Define the node with the neuron output value greater than or equal to the threshold as an activated node, and define the node with the node output value less than the threshold as a non-activated node; (2.2) Take the output value on the activated node as one of the attributes of the activated node, and set the attributes of the non-activated nodes to 0 to achieve the embedding of activation information; (2.3) For the activated nodes, extract the weights between them and all the activated nodes in the upper layer, and use them as the weights on the input connection edges; For the non-activated nodes, extract the maximum value of the weights between them and all the nodes in the upper layer as the weight value of the input connection edge to achieve the sparsification of the connection edges; Establish a corresponding graph structure according to the existing model structure; Step 3: Generate graph features of the model graph; (3.1) Calculate the four basic metrics of the global average degree, modularity, graph density, and average path length of each model graph respectively; (3.2) Concatenate the multiple basic metrics of each model graph into a vector as the graph feature of the model graph; Step 4: Use the graph features of all model graphs and their labels indicating whether they are backdoor models to form a training set for training a binary classification model. The trained binary classification model is called a backdoor model detector; Step 5: For any to-be-detected model composed of a deep neural network, repeat Step 2 and Step 3 to obtain the graph feature corresponding to the to-be-detected model, and input the graph feature into the backdoor model detector. The backdoor model detector outputs the category of the to-be-detected model as a backdoor model or a normal model.
2. The backdoor model detection method guided by graph feature vectors according to claim 1, characterized in that, in step (2.2), after the activation information is embedded, according to the category to which each input sample belongs, assign the same category attribute as the input sample to the activated nodes, and assign an additional attribute of a category different from all the samples in the training set to the non-activated nodes.
3. The backdoor model detection method guided by graph feature vectors according to claim 1, characterized in that, the threshold in step (2.1) is 0.
5.
4. The backdoor model detection method guided by graph feature vectors according to claim 1, characterized in that, the backdoor models and clean models in step (1.3) are models obtained by saving each time the training samples are trained for one round.
5. The backdoor model detection method guided by graph feature vectors according to claim 1, characterized in that, In the fifth step, after the backdoor model detector outputs the category, the detection results of the models trained from the same model structure and dataset are statistically analyzed. According to the voting mechanism, the backdoor model or clean model that appears the most in the detection results is selected as the final detection result.