Graph classification model backdoor attack method for enhancing concealment
Through the degree-centric calculation of edges and the dynamic sample screening mechanism driven by forgetting event, the backdoor attack of the graph classification model is optimized, and the attack effect is achieved with low pollution rate, high concealment and stable, solving the problems of obvious trigger structure and high randomness of sample selection in the existing technology.
Patent Information
- Application Number
- CN202510846107.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-24
AI Technical Summary
The backdoor attack method of the existing graph classification model has problems such as obvious trigger structure, poor concealment and high sample selection randomness, which leads to unstable attack efficiency and easy detection.
The graph structure perturbation trigger injection is used based on edge-based calculation, and combined with the forgetting event-driven dynamic sample selection mechanism, the poisoned sample set is optimized by adjusting the replacement ratio to generate a backdoor attack method with low pollution rate and high concealment.
It improves the attack success rate, reduces the risk of being detected, enhances the concealment of the attack, and maintains the classification accuracy of the clean samples.
Smart Images

Figure CN120356016A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a backdoor attack method for a graph classification model with enhanced concealment, belonging to the technical field of artificial intelligence security. Background Art
[0002] In graph classification tasks, graph neural network (GNN) models achieve graph-level class prediction by learning the structural information and node attributes of the entire graph. As a deep learning model commonly used to process graph-structured data, it is widely applied in scenarios such as social network analysis, protein classification, chemical molecule recognition, etc.
[0003] With the wide deployment of graph neural networks in various practical scenarios, their security issues have gradually attracted attention. Among them, the backdoor attack is an attack method that injects malicious behaviors into the dataset during the model training stage, which will bring great security threats to aspects such as social network analysis, fraud detection, drug toxicity prediction, and malware detection. The attacker carefully tampered with some training samples (such as adding specific structural perturbations and modifying labels). On the premise of not significantly affecting the normal prediction performance of the model, the model can correctly predict normal samples (without backdoor implantation), while when encountering test samples injected with specific trigger patterns, it outputs the classification results expected by the attacker, which makes the security risk of such attacks extremely high.
[0004] However, the existing backdoor attack methods in graph classification tasks still have the following problems: (1) The trigger structure is obvious: Most methods use fixed-pattern subgraphs (such as specific graph structures or graph markers) as triggers and insert them into the target graph, which are easily discovered by anomaly detection or interpretability tools, and the concealment is poor.
[0005] (2) The randomness of sample selection is high: The screening of poisoned samples is usually based on random selection, ignoring the impact of differences between different samples on the attack effect, resulting in a relatively high contamination ratio and unstable attack efficiency. At the same time, the relatively high contamination ratio also makes the attack easily detectable.
[0006] Therefore, there is an urgent need for a graph classification backdoor attack method with stronger concealment and lower contamination rate, which can achieve a stable attack effect at a low contamination rate through more refined structural perturbation design and an effective sample screening mechanism, improve the attack success rate and reduce the risk of being detected. Summary of the Invention
[0007] The present invention aims to creatively propose a backdoor attack method for a graph classification model to address the problems and deficiencies existing in the prior art.
[0008] The present invention first injects a graph structure perturbation trigger, and performs an edge removal operation according to the probability calculated based on the degree centrality of the edges. The smaller the degree centrality of an edge, the higher the probability of its removal. Then, a forgetting event-driven dynamic sample selection mechanism is used to screen and replace the poisoned sample set. In each round, the poisoned samples are sorted according to the number of forgetting events occurring during the training process, and the number of samples to be retained is determined according to the replacement ratio. According to this number, the poisoned samples are retained from high to low according to the sorting result, and new samples are sampled from the candidate set to replace the remaining non-retained samples. The replacement ratio is dynamically adjusted according to the change trend of the attack success rate to improve the sample screening efficiency. Finally, the training set corresponding to the round with the highest attack success rate is selected as the final backdoor training set, and a backdoor model is trained. When performing a backdoor attack, the trigger is injected into the sample to be predicted, and the sample injected with the backdoor trigger is predicted into the desired prediction result by using the trained backdoor model.
[0009] The present invention effectively improves the attack success rate, reduces the contamination ratio, and enhances the concealment of the attack by constructing a highly concealed edge perturbation trigger and combining a forgetting event-driven dynamic poisoned sample screening mechanism.
[0010] To achieve the above object, the present invention adopts the following technical solutions.
[0011] A method for backdoor attack on a graph classification model with enhanced concealment, comprising the following steps: Step 1: Injection of a graph structure perturbation trigger. It includes the following steps: Step 1.1: Obtain the training set of the graph classification task , where each graph sample is represented as , represents the set of nodes in the graph, represents the set of edges in the graph.
[0012] Step 1.2: For each graph sample , calculate the degree centrality of each node in the graph: , where represents the degree of each node, is the number of nodes in the graph.
[0013] Step 1.3: For each edge , calculate the degree centrality of the edge, , where and are the two end nodes of the edge respectively, represents the degree centrality of the -th node, represents the degree centrality of the The degree centrality of the nodes.
[0014] Step 1.4: Take the logarithm of the degree centrality of the edge and normalize it to get the removal probability , ,in, Indicates the value of the degree centrality of the edge after logarithmic change; represents the overall removal probability of the control edge; They represent the maximum and minimum values of the degree centrality of all edges after logarithmic transformation; It represents the calculated removal probability of each edge. The smaller the degree centrality, the greater the removal probability of the edge.
[0015] Step 1.5: All graph data in ,according to Delete the edges to inject triggers and generate a graph with structural perturbations .
[0016] Step 1.6: Generate the graph With the target label Pairing to form candidate poisoning samples All candidate poisoning samples constitute the candidate poisoning sample set .
[0017] Step 2: Based on the candidate poisoning sample set constructed in step 1 Perform forgetting event-driven dynamic poisoning sample screening.
[0018] Specifically, the screening strategy includes the following steps: Step 2.1: From Random sampling Sample, get the initial poisoned sample set .Will and Corresponding The remaining samples that have not been sampled together constitute the initialized backdoor training set ; Set the initialization replacement ratio .in, is the contamination rate.
[0019] Step 2.2: Summarize the poisoned sample set Rounds of screening, each round training the neural network model on the backdoor training set After the training, Each sample in , count the number of forgetting events during training ,in, For the The prediction results of the model for samples during round training ; is the prediction result of the model for samples in the round; ; is an indicator function that has a value of 1 when the condition holds and 0 otherwise. The more times the forgetting event occurs in a poisoned sample, the greater its impact on the backdoor attack.
[0020] Step 2.3: Calculate the attack success rate in each round, and obtain the change amount compared with the previous round . Record the for each round and the poisoned sample set in this round.
[0021] Step 2.4: Dynamically adjust the replacement ratio according to the value of : If If If and roll back the sample set to the previous round; where represents the sample replacement ratio in the round; represents the sample replacement ratio in the round; and are hyperparameters that determine the adjustment amplitude; represents the set adjustment discrimination threshold.
[0022] Step 2.5: According to the adjusted replacement ratio , in each round, remove the poisoned samples with the lowest number of forgetting events in , and then randomly supplement samples from to obtain the new sample set for the next round of training, where represents the number of samples in represents the poisoned sample set in the round; represents the poisoned sample set in the round.
[0023] Step 3: According to the and recorded in Step 2 for each round, select the dataset corresponding to the best value for training to obtain a backdoor model, and use the backdoor model for backdoor attacks.
[0024] When performing a backdoor attack, injecting a trigger into the sample to be predicted through Step 1 enables the trained backdoor model to predict the sample with the injected trigger as the desired prediction result, which is the attack target label. .
[0025] Beneficial effects The method of the present invention realizes a backdoor attack method for graph neural networks with a low pollution rate, high concealment, and strong attack effect by combining graph structure perturbation based on edge degree centrality and a forgetting event-driven dynamic sample screening mechanism. Compared with traditional subgraph insertion and random sample injection strategies, the present invention can improve the attack success rate while maintaining the classification accuracy of clean samples, reduce the risk of being detected, thereby increasing concealment, and can be widely adapted to different types of graph classification models and application scenarios. Description of the drawings
[0026] Figure 1 It is a flowchart of the core steps of the method of the present invention. Detailed implementation manners
[0027] The method of the present invention will be further described in detail below with reference to the drawings.
[0028] As Figure 1 shown, a backdoor attack method for a graph classification model with enhanced concealment includes the following steps: Step 1: Injecting a trigger for graph structure perturbation. It includes the following steps: Step 1.1: Obtaining the training set of the graph classification task , where each graph sample is represented as , represents the set of nodes in the graph, represents the set of edges in the graph.
[0029] Step 1.2: For each graph sample , calculating the degree centrality of each node in the graph. , where represents the degree of each node, is the number of nodes in the graph.
[0030] Step 1.3: For each edge , calculating the degree centrality of the edge, , where and are the two end nodes of the edge respectively, represents the degree centrality of the th node, represents the Degree centrality of nodes.
[0031] Step 1.4: Take the logarithm of the degree centrality of the edges and normalize it to obtain the removal probability , , where represents the value of the degree centrality of the edge after logarithmic transformation; represents the overall removal probability of the control edge; represent the maximum and minimum values of the degree centrality of all edges after logarithmic transformation, respectively; represents the calculated removal probability of each edge. The smaller the degree centrality of the edge, the greater the removal probability.
[0032] Step 1.5: For all the graph data in , according to delete the edges to inject triggers and generate the graph after structural perturbation .
[0033] Step 1.6: Pair the generated graph with the attack target label to form candidate poisoned samples . All candidate poisoned samples form the candidate poisoned sample set .
[0034] Step 2: Perform forgetting event-driven dynamic poisoned sample screening according to the candidate poisoned sample set constructed in Step 1.
[0035] The screening strategy includes the following steps: Step 2.1: Randomly sample samples from to obtain the initialized poisoned sample set . Combine and the remaining samples in that correspond to and have not been sampled to jointly form the initialized backdoor training set ; set the initialized replacement ratio .
[0036] Step 2.2: Perform a total of rounds of screening on the poisoned sample set. In each round, train the neural network model for rounds on the backdoor training set. After the training is completed, for each sample in , count the number of forgetting events that occur during the training process , where is the prediction result of the model for the sample in the is the prediction result of the round model for the sample ; is an indicator function that is 1 when the condition holds and 0 otherwise; the more times the forgetting event occurs in the poisoned sample, the greater the impact on the backdoor attack.
[0037] Step 2.3: Calculate the attack success rate in each round, and obtain the change amount compared with the previous round , record the of each round and the poisoned sample set of this round
[0038] Step 2.4: Dynamically adjust the replacement ratio according to the value of : If If If and roll back the sample set to the previous round; where represents the sample replacement ratio at the round; represents the sample replacement ratio at the round; and are hyperparameters that determine the adjustment amplitude; represents the set adjustment discrimination threshold; Step 2.5: According to the adjusted replacement ratio , in each round, remove the poisoned samples with the lowest number of forgetting events in , and then randomly supplement samples from to obtain the new sample set for the next round of training, where represents the number of samples in represents the poisoned sample set at the round; represents the poisoned sample set at the round; Step 3: According to the and recorded in Step 2 for each round, select the dataset corresponding to the best value for training to obtain a backdoor model, and use the backdoor model for backdoor attacks.
[0039] Specifically, it includes the following steps: Step 3.1: Select the Optimal As the final poisoned sample set .
[0040] Step 3.2: Combine with the remaining corresponding samples in it as the final backdoor training set, and train a backdoor model.
[0041] When performing a backdoor attack, just inject the trigger into the sample to be predicted through Step 1, and the trained backdoor model can predict the sample with the injected trigger into the desired prediction result.
[0042] Embodiment In this embodiment, two publicly available graph classification datasets, PROTEINS and NCI1, are selected. The dataset is divided as follows: 75% for training, 5% for validation, and 20% for testing. The model used for experimental sample screening is the GCN model, and the models to be attacked are the GCN model and the GIN model. The model is trained using the Adam optimizer, with a learning rate of 0.01 and a weight decay coefficient of . The maximum number of training epochs is set to 100, and an early stopping mechanism (patience = 10) is adopted to prevent overfitting. The initial replacement ratio is set to 0.5, the number of screening rounds is set to 40, and is set to 1, is set to 0.05, is set to 0.8, and the sample contamination rate is set to 10%.
[0043] This embodiment compares the proposed backdoor attack method with the prior art, and uses the attack success rate as the evaluation metric. It represents the proportion of samples that originally belonged to non-target classes in the test set and were successfully misclassified as target classes after inserting the backdoor trigger. Denote the method of the present invention as Deg, and the experimental results are shown in Table 1:
[0044] It can be seen from the results in Table 1 that the method of the present invention has a better attack effect than other methods under the same contamination rate, especially in the PROTEINS dataset, which is significantly better than other methods. This proves that the model has learned the backdoor trigger pattern described in the present invention. At the same time, the method of the present invention considers the contribution of samples to the backdoor attack, performs sample screening, and achieves better results.
Claims
1. A backdoor attack method for graph classification models to enhance concealment, characterized in that It includes the following steps: First, inject the graph structure perturbation trigger, and perform edge removal operations according to the probability calculated based on the degree centrality of the edges. The smaller the degree centrality of the edge, the higher the probability of removal; Then, use the forgetting event-driven dynamic sample selection mechanism to screen and replace the poisoned sample set. In each round, sort the poisoned samples according to the number of forgetting events that occur during the training process, determine the number of samples to be retained according to the replacement ratio, retain the poisoned samples from high to low according to the sorting result according to this number, and sample new samples from the candidate set to replace the remaining un-retained samples; Dynamically adjust the replacement ratio according to the change trend of the attack success rate to improve the sample screening efficiency; Finally, select the training set corresponding to the round with the highest attack success rate as the final backdoor training set, and train to generate a backdoor model; When performing a backdoor attack, inject the trigger into the sample to be predicted, and use the trained backdoor model to predict the sample injected with the backdoor trigger into the desired prediction result.
2. The backdoor attack method for a graph classification model with enhanced concealment according to claim 1, wherein It includes the following steps: Step 1: Inject the graph structure perturbation trigger; Step 1.1: Obtain the training set for graph classification tasks , where each graph sample is represented as , represents the set of nodes in the graph, represents the set of edges in the graph; Step 1.2: For each graph sample , calculate the degree centrality of each node in the graph ; Step 1.3: For each edge , calculate the degree centrality of the edge ; Step 1.4: Take the logarithm of the degree centrality of the edges and normalize it to obtain the removal probability ; Step 1.5: For all the graph data in , according to delete the edges to inject triggers, and generate a graph with structural perturbations ; Step 1.6: Pair the generated graph with the attack target label to form a candidate poisoned sample ; All candidate poisoned samples form a candidate poisoned sample set ; Step 2: According to the candidate poisoned sample set Perform dynamic poisoned sample screening driven by forgetting events; The screening strategy includes the following steps: Step 2.1: Randomly sample samples from to obtain an initial poisoned sample set ; Combine the remaining samples in and that correspond to the unsampled ones to form a backdoor training set ; Set an initial replacement ratio ; where is the contamination rate. Step 2.2: Total the poisoned sample set Perform round screening. In each round, train a neural network model on the backdoor training set rounds. After the training ends, for each sample in , count the number of forgetting events that occur during the training process ; Step 2.3: Calculate the attack success rate for each round , and obtain the change compared with the previous round . Record the for each round and the poisoned sample set of this round ; Step 2.4: According to value, dynamically adjust the replacement ratio : Step 2.5: According to the adjusted replacement ratio , eliminate the poisoning samples with the lowest number of forgetting events in each round, and then randomly supplement samples from to obtain the new sample set for the -th round for the next round of training; Among them, represents the number of samples in represents the poisoned sample set at the -th round; represents the poisoned sample set at the -th round; Step 3: According to the and recorded in each round of Step 2, select the dataset corresponding to the optimal value for training to obtain a backdoor model, and use the backdoor model to conduct a backdoor attack; When performing a backdoor attack, inject the trigger into the sample to be predicted through Step 1, and let the trained backdoor model predict the sample injected with the trigger into the desired prediction result.
3. The backdoor attack method for a graph classification model to enhance concealment according to claim 2, wherein, In Step 1.2, calculate the degree centrality of each node in the graph The method for is as follows: , where represents the degree of each node, is the number of nodes in the graph.
4. The backdoor attack method for a graph classification model with enhanced concealment according to claim 2, characterized in that, In step 1.3, calculate the degree centrality of the edge The method is as follows: , where and are the two end nodes of the edge respectively, represents the degree centrality of the th node, represents the degree centrality of the th node.
5. The backdoor attack method for a graph classification model with enhanced concealment according to claim 2, characterized in that, In step 1.4, the method for calculating the removal probability is as follows: , where represents the value of the degree centrality of the edge after logarithmic transformation; represents the overall removal probability of the control edge; respectively represent the maximum and minimum values of the degree centrality of all edges after logarithmic transformation, represents the calculated removal probability of each edge, and the smaller the degree centrality of the edge, the greater the removal probability.
6. The backdoor attack method for a graph classification model to enhance concealment according to claim 2, characterized in that, In step 2.2, calculate the number of forgetting events , where is the prediction result of the model for the sample in the round of training; is the prediction result of the model for the sample in the round of the model; is an indicator function that has a value of 1 when the condition is met and 0 otherwise. The more poisoning samples with more forgetting events occur, the greater the impact on the backdoor attack.
7. The backdoor attack method for a graph classification model enhancing concealment according to claim 2, characterized in that, In step 2.4, according to dynamically adjust as follows: If If If and roll back the sample set to the previous round; Among them, represents the sample replacement ratio at the th round; represents the sample replacement ratio at the th round; and are hyperparameters that determine the adjustment amplitude; represents the set adjustment discrimination threshold.
8. The backdoor attack method for a graph classification model to enhance concealment according to claim 2, characterized in that, Step 3 includes the following steps: Step 3.1: Select all the optimal as the final poisoned sample set ; Step 3.2: Combine the remaining corresponding samples in and as the final backdoor training set, and train a backdoor model.
Citation Information
Patent Citations
Backdoor attack method based on class activation features in forgetting event
CN117579363A
Backdoor attack detection method and device based on forgetting learning and expected transition probability
CN118379607A
Back door defense method and system based on reverse forgetting
CN120030542A
Post-Training Detection and Identification of Backdoor-Poisoning Attacks
US20210256125A1