Non-backdoor model node injection attack method in graph neural network
Through the node injection attack method without backdoor model, the general trigger generator model and multiple graph neural network test models are used to solve the problem that existing backdoor attack methods rely on backdoor models and complex algorithms, and achieve high concealment and practical operability attack effects.
Patent Information
- Application Number
- CN202510193448.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The existing backdoor attack methods rely on the construction of backdoor models and complex node selection algorithms, with high authority requirements and poor flexibility, making it difficult to effectively implement in actual scenarios.
A node injection attack method without backdoor model is proposed. By obtaining the original data set of the target graph neural network, pre-training the mediation model, building a trigger generator model with a specific structure, generating and binding the trigger nodes, forming a poisoned data set, and training it through the general trigger generator model, a graph neural network test model with multiple different structures is obtained for prediction to evaluate the attack performance.
It realizes an attack method that does not require relying on backdoor models and complex algorithms, improves the concealment and practical operability of the attack, and can effectively attack target nodes while maintaining the prediction accuracy of non-target nodes.
Smart Images

Figure CN120124049A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a method for attacking nodes of a backdoor-free model in a graph neural network. Background Art
[0002] Graph neural networks (GNNs) update node representations by aggregating neighbor node information and are a type of deep learning model for processing graph data. GNNs have been applied to tasks such as node classification, link prediction, and graph embedding, and their algorithms include classical graph convolutional networks (GCNs), graph attention networks (GATs), and GraphSAGE, etc. However, as GNNs are deployed in sensitive scenarios (such as financial fraud detection, medical diagnosis, etc.), their security issues have become increasingly prominent.
[0003] As a new type of security threat, backdoor attacks implant hidden triggers during the model training phase, causing the model to produce outputs specified by the attacker for specific inputs after deployment. Existing backdoor attack methods usually require the attacker to have full control over the dataset and the model training process. For example, trigger nodes are embedded in the training data and node features are modified, and the trained backdoor model is used to replace the original model for prediction to achieve the attack. Although these methods perform well in terms of attack success rate, they have high requirements for the attacker's permissions and are difficult to implement in actual scenarios. In addition, existing methods rely on complex node selection algorithms, which further limit the flexibility of the attack.
[0004] Therefore, how to design an attack method that does not rely on a backdoor model or complex algorithms, while taking into account both concealment and attack effect, is an important new direction. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a method for attacking nodes of a backdoor-free model in a graph neural network to solve the problems of relying on the construction of a backdoor model and complex node selection algorithms in the existing backdoor attack methods.
[0006] To achieve the above object, the present invention provides a method for attacking nodes of a backdoor-free model in a graph neural network, including:
[0007] Obtain the original dataset of the target graph neural network and pre-train an intermediary model based on the original dataset;
[0008] Construct a specific structure trigger generator model and generate trigger nodes, bind the trigger nodes to target nodes to obtain a poisoned dataset;
[0009] Train the specific structure trigger generator model based on the poisoned dataset and the intermediary model to obtain a general trigger generator model;
[0010] Build and train several graph neural network test models with different structures; use a general trigger generator to inject trigger nodes into multiple test data sets to obtain poisoned test data sets; predict the poisoned test data sets respectively through several graph neural network test models with different structures, and evaluate the attack performance according to the prediction results.
[0011] Optionally, the process of pre-training the mediation model based on the original data set includes:
[0012] Construct a graph structure based on the original data set and perform preprocessing, perform symmetric normalization on the adjacency relation matrix in the original data set to obtain an optimized graph structure; train the mediation model through the optimized graph structure, where a graph sampling aggregation neural network is used as the mediation model.
[0013] Optionally, the original data set includes a node feature matrix, an adjacency relation matrix, and node label information.
[0014] Optionally, the process of obtaining the poisoned data set includes:
[0015] Generate trigger nodes through a specific structure trigger generator model; construct a ring topology structure, based on the ring topology structure, connect the trigger nodes and the target nodes through edges, and update the generated trigger nodes and their connection relationships into the original graph structure to form a poisoned data set; where the specific structure trigger generator model includes a multi-layer perceptron and a residual connection.
[0016] Optionally, the process of obtaining the general trigger generator model includes:
[0017] Perform forward propagation through the mediation model using the poisoned data set to verify the effectiveness of the trigger generator model; construct a comprehensive loss function, perform iterative optimization in combination with the comprehensive loss function, update the parameters of the trigger generator model, and obtain the general trigger generator model.
[0018] Optionally, during the training process, the DropEdge mechanism is used to randomly remove the connections between the trigger nodes and the target nodes, and update the poisoned data set.
[0019] Optionally, the comprehensive loss function is obtained by weighting and adding the cross-entropy loss function and the cosine similarity loss function.
[0020] Optionally, several graph neural network test models with different structures include a graph convolutional network, a graph attention network, and a graph sampling aggregation network.
[0021] Optionally, it further includes processing the poisoned test data set to set up a no-defense scenario and a defense scenario. In the defense scenario, the poisoned test data set is operated through two defense mechanisms; the attack performance is evaluated by the ratio of the target nodes being misclassified as the attacker-specified class to the correct classification rate of the non-target nodes in the poisoned test data set.
[0022] Compared with the prior art, the present invention has the following advantages and technical effects:
[0023] The present invention proposes a training mechanism for a general trigger generation model. The triggers generated by this model are directly bound to the nodes of the original data set, enabling the graph neural network model trained based on the original data set to predict the attacker-specified label for the target nodes and maintaining the prediction accuracy of the non-target nodes. The node injection attack method for graph neural networks provided by the present invention solves the problems of relying on the construction of a backdoor model and a complex node selection algorithm in the existing backdoor attack methods, and improves the concealment and practical operability of the attack method. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0025] Figure 1 is the flowchart of the method according to the embodiment of the present invention;
[0026] Figure 2 is the implementation scheme diagram of the graph neural network backdoor attack according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.
[0028] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0029] Embodiment 1
[0030] As Figure 1-2 shown, this embodiment provides a node injection attack method without a backdoor model in a graph neural network, including:
[0031] Obtain the original data set of the target graph neural network, and pre-train an intermediary model based on the original data set;
[0032] Furthermore, the process of pre-training the mediation model based on the original dataset includes:
[0033] Construct a graph structure based on the original dataset and perform preprocessing. Symmetrically normalize the adjacency relation matrix in the original dataset to obtain an optimized graph structure. Train the mediation model through the optimized graph structure, where a graph sampling aggregation neural network is used as the mediation model.
[0034] Specifically, as Figure 2 shown in the training process, obtain the original dataset of the target graph neural network. These datasets usually include a node feature matrix, an adjacency relation matrix, and node label information:
[0035] Node feature matrix X: The feature representation of each node, usually sparse or high-dimensional data;
[0036] Adjacency matrix A: The connection relationship between nodes in the graph, representing the structure of the graph;
[0037] Label set Y: The true class labels of some nodes, used to supervise model training;
[0038] Construct a graph structure:
[0039] Use the above data to construct a graph structure G=(V, E), which is convenient to be input into the model for training, where V is the set of nodes and E is the set of edges, representing the relationship between nodes.
[0040] Preferably, the graph structure can be preprocessed, including removing isolated nodes or low-weight edges to simplify the graph structure, so as to reduce the training budget and enhance the information density; use a method to symmetrically normalize the adjacency matrix A to improve the stability of training.
[0041] Pre-training of the mediation model:
[0042] GraphSAGE (Graph Sample and Aggregate) is a graph neural network (GNN) framework based on neighbor sampling and feature aggregation. It gets its name from the inductive learning of node embeddings and mainly solves the problem of how to efficiently scale to large-scale graphs and generalize to new nodes. GraphSAGE is composed of "Graph" and "SAGE (Sample and Aggregate)". In addition, GraphSAGE also avoids information redundancy through neighbor sampling and introduces a flexible aggregation function to enhance the model's ability to express node feature patterns, thus performing well on large-scale graphs.
[0043] What makes GraphSAGE unique is that it addresses the problem of traditional GNNs relying on the entire graph structure through random neighbor sampling and a multi-layer aggregation mechanism, thereby improving the scalability and adaptability of the model. This enables GraphSAGE to effectively reduce the risk of overfitting and also perform excellently on dynamic graphs and partially observed graph data.
[0044] In the embodiments of the present invention, a graph sampling aggregation neural network (GraphSage) is used as an intermediary model for training. Neighbor nodes are sampled layer by layer, and the embedding representation of the target node is generated by aggregating the domain features. The specific formula is as follows:
[0045]
[0046] Among them, Sample(N(i)) represents a sampled subset of the neighbor nodes of node i, W (k) and are learnable weight matrices, representing the weights for learning features from neighbor nodes and self-features respectively. σ represents a non-linear activation function, and Aggregate is a function that combines the features of the sampled neighbor nodes. In the examples of the present invention, the mean aggregation function is used, which aggregates the features of neighbor nodes by taking the average of the feature vectors of neighbor nodes.
[0047] Construct a specific structure trigger generator model and generate trigger nodes, bind the trigger nodes to the target nodes to obtain a poisoned dataset;
[0048] Based on the poisoned dataset and the intermediary model, train the specific structure trigger generator model to obtain a general trigger generator model;
[0049] Furthermore, the process of obtaining the poisoned dataset includes:
[0050] Generate trigger nodes through a specific structure trigger generator model; construct a ring topology structure. Based on the ring topology structure, connect the trigger nodes and the target nodes through edges, and update the generated trigger nodes and their connection relationships into the original graph structure to form a poisoned dataset; wherein, the specific structure trigger generator model includes a multi-layer perceptron and a residual connection.
[0051] During the training process, the DropEdge mechanism is used to randomly remove the connections between the trigger nodes and the target nodes to update the poisoned dataset.
[0052] Furthermore, the process of obtaining the general trigger generator model includes:
[0053] Forward propagate through the mediation model using the poisoned dataset to verify the effectiveness of the trigger generator model; construct a comprehensive loss function, perform iterative optimization in combination with the comprehensive loss function, update the parameters of the trigger generator model, and obtain a general trigger generator model.
[0054] The comprehensive loss function is obtained by weighting and adding the cross-entropy loss function and the cosine similarity loss function.
[0055] Specifically, use a multi-layer perceptron (MLP) and a residual structure to generate the feature representation of the trigger node. The formula is: h tr = MLP(h t ) + h t , where h t is the feature representation of the target node, and h tr is the feature representation of the trigger node;
[0056] Construct a circular topological structure, connect the trigger node and the target node through an edge to enhance the attack effect of the trigger and reduce the budget required for the attack;
[0057] Adopt the DropEdge mechanism to randomly remove the connection between the trigger node and the target node to improve the concealment and generalization ability of the injected data.
[0058] During the trigger generation process, introduce the cosine similarity loss function to optimize the feature similarity between the trigger node and the target node, enhance the effectiveness of the adversarial defense mechanism, and then weight and add the cross-entropy loss function and the cosine similarity loss function to obtain the final loss function for model training until convergence.
[0059] Exemplarily, the trigger generator model is used to generate trigger nodes similar to the target node features for each attack group. This model takes the features of the target node as input, passes through a multi-layer perceptron (MLP) and residual connections, and generates the features of the trigger node. Let the feature vector of a certain attacked node in the k-th group be x i , and the generated feature representation of the trigger node is
[0060] The calculation process of the trigger generator model is expressed as:
[0061]
[0062] Among them, MLP(·) represents a multi-layer perceptron, and the residual connection directly adds the input feature to the output feature to ensure the stability and effectiveness of generating trigger features.
[0063] The connection structure between the trigger and the target node constitutes:
[0064] In this embodiment, for each attacked node V a , corresponding trigger nodes are generated according to the number of attacked nodes. Suppose there are |V a | attacked nodes, and k×|V a | trigger nodes are generated, where k is the proportional parameter for generating trigger nodes. A complete graph structure is formed between the generated trigger nodes and the attacked nodes. Specifically, the connection between the trigger nodes and the attacked nodes is unidirectional, that is, the trigger nodes are only connected to the attacked nodes.
[0065] In addition to the connection between the trigger nodes and the attacked nodes, this embodiment also forms a ring connection structure between the trigger nodes in each group to enhance the information propagation ability between the trigger nodes. The ring connection means that each trigger node t i is only connected to the next trigger node t i+1 , until the last trigger node is connected to the first trigger node to form a closed ring. This design helps to share information between the trigger nodes, increase the stealth and robustness of the attack, while reducing the number of trigger nodes and saving the attack budget.
[0066] The DropEdge mechanism is applied to the connection edges between the trigger and the target nodes:
[0067] In the embodiment of the present invention, in order to prevent the model from overfitting to a specific edge structure, the DropEdge technique is used to randomly delete the edges between the trigger nodes and the attacked nodes. During the training process, a proportion p of the edges are randomly deleted in each round to enhance the robustness of the model and prevent the model from relying on specific edges.
[0068] While ensuring the attack success rate by using a two-way optimized cross-entropy loss function, a relatively small impact is produced on non-target nodes. The specific formula is as follows:
[0069]
[0070] Among them, θ g represents the parameters of the trigger generation model, and λ is a hyperparameter used to balance the trade-off between the classification accuracy of clean nodes and the attack success rate of target nodes. V L and V P represent the clean node set and the attacked node set respectively. The symbol l(·) represents the cross-entropy loss function, and f m (·) represents the prediction result of the intermediate model. The function represents the connection operation between the trigger node and the attacked node, where is used to ensure the accurate classification of clean nodes, while is used to force the target node to be misclassified as the specified target category.
[0071] Introduce the cosine similarity loss function to ensure the stealthiness of the trigger and the performance of the adversarial defense mechanism. The specific formula is as follows:
[0072]
[0073] where x ti and x i represent the feature vectors of the trigger node and the attacked node i respectively, is the set of all target nodes. This loss aims to align the features of the generated trigger node as closely as possible with those of the attacked node.
[0074] Construct a comprehensive loss function by combining the bidirectional optimization cross-entropy loss function and the cosine similarity loss function as follows:
[0075]
[0076] where λ attack and λ cos are hyperparameters used to adjust the weights of different loss terms. Update the trigger generator model by gradient using this comprehensive loss function, and finally obtain the optimized trigger generator model, which is the core for subsequent implementation of backdoor attacks on graph neural networks.
[0077] Construct and train graph neural network test models with several different structures; use a general trigger generator to inject trigger nodes into multiple test datasets to obtain poisoned test datasets; predict the poisoned test datasets respectively through several graph neural network test models with different structures, and evaluate the attack performance according to the prediction results.
[0078] Furthermore, the graph neural network test models include graph convolutional networks, graph attention networks, and graph sampling aggregation networks.
[0079] Furthermore, it also includes processing the poisoned test datasets in a no-defense scenario and a defense scenario. In the defense scenario, operate on the poisoned test datasets through two defense mechanisms; evaluate the attack performance by the ratio of target nodes being misclassified as the attacker-specified class and the correct classification rate of non-target nodes in the poisoned test datasets.
[0080] Specifically, as shown in the attack process in Figure 2 , use the original dataset consistent with the training process to provide a data basis for subsequent verification of the attack effect. At the same time, to comprehensively test the generality of the backdoor attack method and ensure its applicability to graph neural network models with multiple architectures, three commonly used models are selected as test objects.
[0081] These benchmark models include Graph Convolutional Network (GCN), Graph Attention Network (GAT), and Graph Sampling Aggregation Network (GraphSAGE), which represent different feature aggregation mechanisms and application scenarios respectively. This selection provides a diverse test platform for subsequent injection attacks, ensuring the sufficiency and comprehensiveness of the verification process.
[0082] Randomly select target nodes, use the trigger generator model obtained from the above training, generate corresponding triggers according to the target node features, and bind them to the target nodes through the same connection structure as during training to obtain the injected dataset.
[0083] Process the injected dataset under the scenarios of no defense and defense, which are described as follows:
[0084] No defense scenario: Do not perform any operations on the injected dataset.
[0085] Defense scenario: Include two defense mechanisms, Prune and advanced Prune+LD, to operate on the injected dataset. Prune removes the parts that may carry backdoors by pruning redundant neurons in the neural network; Prune+LD is its enhanced version, which combines pruning and loss detection mechanisms. By analyzing the loss changes of the network when processing clean data and trigger data, it further detects and removes backdoors.
[0086] Let three different benchmark models predict the injected dataset respectively, and evaluate the effectiveness of the attack through two metrics, ASR and CA:
[0087] ASR is defined as the ratio of target nodes being misclassified as the categories specified by the attacker, and the calculation formula is:
[0088]
[0089] CA is defined as the correct classification rate of non-target nodes in the poisoned test dataset, and the calculation formula is:
[0090]
[0091] For the specific test results, please refer to Table 1. In the case of no defense mechanism, the attack success rates of this implementation method are all above 97%, demonstrating the reliability in terms of attack efficiency. When applying the Prune and Prune+LD defense strategies, the attack success rates of this implementation method still generally remain above 90%, except for a relatively poor performance in the citeceer dataset, showing that this implementation method has a certain ability to resist defense strategies.
[0092] Table 1
[0093]
[0094] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A backdoor-free model node injection attack method in a graph neural network, characterized in that: The following steps are involved: Obtain the original data set of the target graph neural network and pre-train the intermediate model based on the original data set; Constructing a specific structure trigger generator model and generating a trigger node, binding the trigger node to a target node, and obtaining a poisoned data set; Training a specific structure trigger generator model based on the poisoning data set and the intermediary model to obtain a general trigger generator model; Construct and train several graph neural network test models with different structures; use the universal trigger generator to inject trigger nodes into multiple test data sets to obtain poisoned test data sets; use several graph neural network test models with different structures to predict the poisoned test data sets respectively, and evaluate the attack performance based on the prediction results.
2. The backdoor-free model node injection attack method in a graph neural network according to claim 1 is characterized in that: The process of pre-training the intermediary model based on the original dataset includes: A graph structure is constructed based on the original data set and preprocessed, and the adjacency relationship matrix in the original data set is symmetric normalized to obtain an optimized graph structure; the intermediary model is trained by optimizing the graph structure, wherein a graph sampling aggregation neural network is used as the intermediary model.
3. The backdoor-free model node injection attack method in a graph neural network according to claim 2 is characterized in that: The original data set includes a node feature matrix, an adjacency relationship matrix and node label information.
4. The backdoor-free model node injection attack method in a graph neural network according to claim 1 is characterized in that: The process of obtaining the poisoned dataset includes: Generate trigger nodes through a specific structure trigger generator model; construct a ring topology structure, connect the trigger nodes with the target nodes through edges based on the ring topology structure, update the generated trigger nodes and their connection relationships to the original graph structure, and form a poisoned data set; wherein the specific structure trigger generator model includes a multi-layer perceptron and a residual connection.
5. The backdoor-free model node injection attack method in a graph neural network according to claim 1 is characterized in that: The process of obtaining a generic trigger generator model includes: The poisoned dataset is used to perform forward propagation through the intermediary model to verify the effectiveness of the trigger generator model. A comprehensive loss function is constructed, and iterative optimization is performed in combination with the comprehensive loss function to update the parameters of the trigger generator model and obtain a general trigger generator model.
6. The backdoor-free model node injection attack method in a graph neural network according to claim 5 is characterized in that: During the training process, the DropEdge mechanism is used to randomly remove the connection between the trigger node and the target node to update the poisoned dataset.
7. The backdoor-free model node injection attack method in a graph neural network according to claim 1 is characterized in that: The comprehensive loss function is obtained by weighting and adding the cross entropy loss function and the cosine similarity loss function.
8. The backdoor-free model node injection attack method in a graph neural network according to claim 1, characterized in that: Several graph neural network test models with different structures include graph convolutional networks, graph attention networks, and graph sampling aggregation networks.
9. The backdoor-free model node injection attack method in a graph neural network according to claim 1, characterized in that: It also includes setting up defenseless and defense scenarios for the poisoned test dataset, operating the poisoned test dataset through two defense mechanisms in the defense scenario; and evaluating the attack performance by the ratio of target nodes misclassified as the attacker-specified category and the correct classification rate of non-target nodes in the poisoned test dataset.
Citation Information
Patent Citations
Dual model replacement backdoor attack method and system based on federated learning
CN117151170A
Image classification backdoor attack method, device and equipment based on image contour
CN118379534A
System and Method for Video Backdoor Attack
US20220027462A1
Cited By
Deviation correction method for classification fairness of graph neural network
CN121705893A