Rumor detector-oriented backdoor attack generation method, system and device
By building an adaptive discrete trigger generator, the trigger nodes are injected into the key nodes in the rumor propagation graph, the problem that the existing technology is easily detected when attacking the rumor propagation tree is solved, and a hidden and transferable backdoor attack is realized, which improves the security of the rumor detector.
Patent Information
- Application Number
- CN202510094225.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-21
Smart Images

Figure CN119989335A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of rumor detection and backdoor attack, and in particular to a backdoor attack generation method, system and device for a rumor detector. Background Art
[0002] With the rapid development of the Internet, social media has become an important channel for users to obtain information and communicate. A large amount of inaccurate and unverified information, especially malicious information such as rumors, spreads widely and rapidly, causing significant harm to society. Therefore, it is crucial to detect rumors on social platforms.
[0003] Deep learning plays an important role in rumor detection. It can automatically and efficiently learn feature vectors containing deep semantic information from rumor text and images. For example, methods based on recurrent neural networks (RNNs) can effectively capture the temporal relationship between posts. Methods based on convolutional neural networks (CNNs) can learn local spatial feature representations of rumors. However, these methods mainly focus on the text information of rumors and ignore the structural information inherent in the spread of rumors. This structural information is crucial because it reflects the relationship between posts, reveals the collective wisdom and stance of user comments, and is a key feature to distinguish rumors from non-rumors. Therefore, in order to enhance detection capabilities, recent studies have begun to incorporate propagation structures into rumor detection models by utilizing graph neural networks (GNNs) and integrating other advanced technologies. The propagation structure of rumors provides valuable clues to how the stance in the original post spreads and develops over time.
[0004] Although propagation-based methods have performed well in rumor detection tasks, malicious users and criminals may exploit the vulnerabilities of graph neural networks (GNNs) to evade or interfere with rumor detection results, raising concerns about the security of rumor detectors. Although recent research has focused on adversarial attacks, such as using a reinforcement learning framework to explore the vulnerabilities of GNN-based rumor detectors under structural adversarial attacks, another type of attack, backdoor attacks, has been largely overlooked. In a backdoor attack, the target model is poisoned by injecting triggers into a small portion of the data in the training set and modifying the labels. After training, the model runs normally on clean samples, but misclassifies samples containing triggers as the labels expected by the attacker. This allows the attacker to covertly change the detection results during system operation. Such attacks can bypass initial security reviews and testing, posing a significant threat to rumor detection models.
[0005] Several studies have explored backdoor attacks on graph neural networks (GNNs). They usually form fixed-size subgraph triggers by perturbing the graph topology, which performs well in graph classification tasks but is not suitable for rumor detection tasks because the propagation tree sizes vary greatly. More importantly, rumor propagation trees are significantly different from graphs, and there are the following flaws when trying to attack them: (i) Any perturbation to their structure is easily detectable, such as Figure 1 As shown in Figure 2, adding an edge may destroy the tree structure, and deleting an edge may destroy the link. In addition, on real social platforms, attackers cannot tamper with the social records posted by other users. Even if users delete posts, social platforms may retain historical information. (ii) Selecting appropriate content as a trigger for feature injection is also a major challenge. The content of the injected post must be related to the topic of the post it replies to and retain the attack attributes to evade detection and effectively trigger the backdoor. This dual requirement of context relevance and effectiveness poses a huge challenge to implementing backdoor attacks on rumor detectors. Summary of the invention
[0006] In view of the shortcomings of existing attack technologies that are easily detected when attacking the rumor propagation tree and are not suitable for rumor detection tasks, the present invention proposes a backdoor attack generation method, system and device for rumor detectors. By constructing an adaptive discrete trigger generator, trigger nodes are injected into key nodes to create a covert and transferable attack, thereby solving the problems existing in the prior art.
[0007] A backdoor attack generation method for a rumor detector includes the following steps:
[0008] Obtain a rumor propagation graph; the rumor propagation graph includes a graph node set consisting of source posts and their comments and an edge set consisting of the relationship between each comment or between the source post and the comment;
[0009] The importance scores of the node structure and features of the rumor propagation graph are generated by using the degree centrality and the similarity between the node and the root node, respectively. The degree centrality is used to measure the degree of the node as the measure of the node centrality. The root node represents the source post, the node represents the comment of the source post, and the degree of the node represents the number of edges connected to each node. After sorting the importance scores, the M nodes with the highest values are selected as the attachment nodes, and each attachment node is attached with a trigger node. The comments newly added outside the original propagation graph are used as the trigger nodes. An adaptive feature generator is constructed. The adaptive feature generator uses the node features of the attachment nodes as input to generate an aggressive adaptive trigger node. At the same time, similarity constraints are added to the node features so that the trigger node can adaptively approach the node features of the attachment nodes. The adaptive feature generator is used to adaptively generate corresponding node features for each trigger node, and the trigger node with node features is injected into the attachment node to generate a backdoor graph.
[0010] The generated backdoor graph is used to attack the propagation-based rumor detector to test the security of the rumor detector.
[0011] Furthermore, the degree centrality and the similarity between the node and the root node are used to generate the importance score S(v) of the node structure and feature of the propagation graph, respectively, which is expressed as:
[0012]
[0013] Among them, sim(·) represents the cosine similarity function, is the node centrality value of node v, β is the fusion factor used to control the ratio of the two indicators, x root and x v Represent the characteristics of the root node and the attached node respectively.
[0014] Furthermore, the constraints of the adaptive feature generator are expressed as:
[0015]
[0016] Among them, x u 、x v They represent the features of the generated trigger node u and the attached node v, T represents the cosine similarity threshold, and E g Represents the edge set containing the edges connecting the trigger node and the attachment node.
[0017] Furthermore, the adaptive feature generator is used to adaptively generate corresponding node features for each trigger node, which is expressed as:
[0018] x u =σ(W1·x υ +b1)W2+b2
[0019] where W1, W2, b1, and b2 are parameters that the feature generator can learn, and σ(·) is the activation function.
[0020] Furthermore, the backdoor graph G gt It is expressed as:
[0021]
[0022] Where G represents a clean graph and m(·) represents the trigger g t A hybrid function that is injected into a given graph to generate a trigger embedding graph, where M represents the number of attached nodes and A Ggt represents the adjacency matrix of the trigger embedding graph, A G Adjacency matrix representing the clean embedded graph.
[0023] The present invention also proposes a backdoor attack generation system for rumor detectors, comprising:
[0024] An acquisition module is used to acquire a rumor propagation graph; the rumor propagation graph includes a graph node set consisting of source posts and their comments and an edge set consisting of the relationship between each comment or between the source post and the comment;
[0025] A generation module is used to generate the importance scores of the node structure and features of the rumor propagation graph respectively by using the degree centrality and the similarity between the node and the root node, wherein the degree centrality is used to measure the degree of the node as the measure of the node centrality, the root node represents the source post, the node represents the comment of the source post, and the degree of the node represents the number of edges connected to each node; after sorting the importance scores, the highest M nodes are selected as the attachment nodes, and each attachment node is attached with a trigger node; wherein the newly added comments outside the original propagation graph are used as the trigger nodes; an adaptive feature generator is constructed, wherein the adaptive feature generator uses the node features of the attachment nodes as input to generate an aggressive adaptive trigger node, and adds similarity constraints to the node features so that the trigger node can adaptively approach the node features of the attachment nodes; the adaptive feature generator is used to adaptively generate corresponding node features for each trigger node, and the trigger node with the node features is injected into the attachment node to generate a backdoor graph;
[0026] The attack module is used to attack the propagation-based rumor detector through the generated backdoor graph to test the security of the rumor detector.
[0027] The present invention also proposes a computer device for generating a backdoor attack for a rumor detector, comprising: a memory, a processor, and a computer program stored in the memory, wherein the processor implements the steps of the method for generating a backdoor attack for a rumor detector when executing the computer program.
[0028] The present invention also proposes a readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, they are used to execute the steps of the backdoor attack generation method for the rumor detector.
[0029] The present invention provides a backdoor attack generation method for rumor detectors, which has the following beneficial effects:
[0030] The present invention proposes a backdoor attack framework for a propagation-based rumor detection method, which aims to launch a targeted attack on a specific rumor while maintaining the overall performance of the rumor detector unaffected. By screening out attached nodes and their triggering nodes, an adaptive feature generator is proposed to mislead a model to capture the hidden association between the trigger and the target class. The adaptive feature generator can use the node features of the attached nodes as input to generate an aggressive adaptive triggering node, thereby creating a hidden and transferable attack. In this way, newly added comments outside the original propagation graph can be used as trigger nodes to simulate the controlled behavior of users in rumor propagation to attack the propagation-based rumor detector, thereby solving the problem that the existing attack technology is easily detected when attacking the rumor propagation tree and is not suitable for the rumor detection task. By attacking the rumor detector based on propagation with the backdoor attack generated by the present invention, the test effect of the rumor detector is improved, and the security of rumor detection is further enhanced. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Schematic diagram of the difference between backdoor attack and propagation tree in the background technology; (a) shows an example diagram of injecting backdoor trigger into the graph; (b) shows an example diagram of injecting backdoor trigger into the rumor propagation tree;
[0032] Figure 2 It is a schematic diagram of the overall framework of the injection-based backdoor attack (IBAttack) in an embodiment of the present invention;
[0033] Figure 3 This is a schematic diagram of average similarity distribution in an embodiment of the present invention;
[0034] Figure 4 Schematic diagram of the effect of trigger size on attack effect in an embodiment of the present invention;
[0035] Figure 5 Schematic diagram of the effect of trigger size on attack concealment in an embodiment of the present invention;
[0036] Figure 6 Schematic diagram of the effect of the poisoning rate on Twitter 16 in an embodiment of the present invention;
[0037] Figure 7 Schematic diagram of hyperparameter sensitivity analysis of T and α in an embodiment of the present invention;
[0038] Figure 8 Schematic diagram of the hyperparameter sensitivity analysis of β in an embodiment of the present invention. DETAILED DESCRIPTION
[0039] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0040] This paper proposes an injection-based backdoor attack (IBAttack), an adaptive trigger generation framework for propagation-based rumor detectors. To adapt to the rumor propagation structure, the present invention adopts a discrete trigger structure and generates adversarial node features related to the original post context. These trigger points are then attached to key representative nodes to generate a powerful, covert and transferable backdoor attack.
[0041] Rumor detection based on propagation: The rumor detection task can be defined as a graph-level classification task. Specifically, the rumor detection dataset is represented as C = {c1, c2, …, c i , …, c n}, where c i is the i-th event, n is the number of events, and each event c i =(G i ,y i ) including its propagation structure G i =(V i , E i ), where V i and E i Represent the graph node set (source post and its comments) and edge set (relationship between replies or between source post and reply), as well as its true label y i ∈{N, R} (i.e., non-rumor or rumor) or a fine-grained label y i ∈{N, F, T, U} (i.e., non-rumor, false rumor, true rumor, unconfirmed rumor). Given a dataset, the goal of the rumor detection task is to learn a classifier:
[0042] f:G→Y (1)
[0043] Map each graph G to Y = {y1, ..., y n}. The parameters of f are optimized by gradient descent using the loss function L train (e.g. cross entropy) are learned on labeled datasets.
[0044] Backdoor attack on rumor detector: The backdoor attack injects malicious functions into the target model as hidden neuron Trojans. When the trigger exists, the hidden Trojan will be activated and mislead the model to get the expected output.
[0045] Attacker’s goal: The attacker uses the backdoor attack to influence the final trained rumor detector. There are two goals: (i) The backdoor rumor detection model should predict the expected label on the propagation graph with triggers. (ii) The backdoor attack should not affect the accuracy of the rumor detection model on the benign propagation graph, making the attack concealed. Therefore, the attack goal can be defined as:
[0046]
[0047] Among them, G represents the clean graph, y t is the target attack category, m(·) represents the hybrid function that injects triggers into a given graph to generate a trigger embedding graph, and f θ and They represent benign rumor detector and backdoor rumor detector respectively.
[0048] Figure 1 Schematic diagram of the difference between backdoor attack and propagation tree; (a) shows an example graph of injecting backdoor trigger into the graph; (b) shows an example graph of injecting backdoor trigger into the rumor propagation tree. The subgraph composed of red nodes and edges represents the trigger. Due to the complex graph structure, the graph with embedded trigger is still roughly similar to the original graph in structure, allowing the backdoor attack to evade detection.
[0049] The overall framework of the attack of the present invention is as follows Figure 2 As shown. Unlike the existing backdoor attack methods for graph classification tasks, the triggers in the attack of the present invention are more flexible and are manifested as multiple discrete new nodes. Specifically, the present invention selects the most vulnerable and important nodes as attachment nodes, and injects new nodes into them to create a trigger embedding graph, and then trains a feature generator to adaptively generate features for these injected nodes. In this way, malicious comments (new comments outside the original propagation graph) can be used as trigger nodes to simulate the controlled behavior of users in rumor propagation, thereby attacking the rumor detector based on propagation. The method specifically includes:
[0050] Adaptive trigger generation;
[0051] In order to keep the original graph intact, the present invention adopts a trigger strategy consisting of discrete injection nodes, that is, connecting new nodes with a certain budget to the propagation graph to form a further evolved rumor propagation tree. Based on such a trigger structure, the objectives of the present invention are decomposed into attachment node selection, trigger feature generation and trigger injection.
[0052] (1) Attached Nodes Selection: The present invention designs an attached node selection scheme to decide which nodes the injected node should be connected to. Since the graph neural network (GNN) model aggregates information by learning the features of nodes and their neighbors, it is crucial for key nodes to transmit information in the graph structure. Cui and Jia (2024) analyzed the structural characteristics of the rumor propagation tree (RPT) and summarized three principles for assigning importance scores to nodes to improve rumor detection performance. Similarly, connecting the trigger node to important nodes can achieve better attack performance than randomly injecting nodes. IBAttack recommends the use of two node importance metrics, including degree centrality and the similarity between the node and the root node.
[0053] Degree centrality uses the degree of a node as a measure of node centrality. Many studies have shown that nodes with lower degrees are less robust to attacks than nodes with higher degrees. Intuitively, posts that are heavily discussed and replied to are more difficult to be influenced.
[0054] The similarity between a node and the root node; in the rumor propagation tree, the root node, as the source post, usually contains more important and rich information. Intuitively, nodes with lower similarity to the root node are more likely to affect the entire propagation graph.
[0055] The present invention uses the above two metrics to reflect the importance of node structure and features respectively. In different situations, different strategies or combinations of the two metrics can be used to obtain the importance score:
[0056]
[0057] Among them, sim(·) represents the cosine similarity function, is the node centrality value of node v, and β is the fusion factor used to control the ratio of the two indicators. root and x v Represent the characteristics of the root node and the node respectively. For the candidate graph, the present invention selects the first m nodes as the attachment nodes according to the importance score, and each attachment node is attached with a trigger node. The budget and injection process will be described in the trigger injection module.
[0058] (2) Triggering feature generation: Graphs in the real world, such as social networks, usually exhibit homogeneity, that is, nodes with similar features are connected by edges. Rumor propagation graphs are no exception. The average similarity distribution between nodes in Twitter15 and Twitter16 is shown in Figure 2. Figure 3As shown in Figure 2, most graphs show extremely high homogeneity because all nodes in the propagation graph represent discussions related to the source post. Therefore, the reply content is more likely to embed the same topic information as the original post. Based on this phenomenon, the present invention adds similarity constraints to the node features so that the trigger node can adaptively approach the features of the attached node. Let E g represents the edge set containing the edges connecting the trigger node and the attachment node. The constraints on the generated adaptive trigger can be written as:
[0059]
[0060] Among them, x u 、x v denote the features of the generated trigger nodes and attachment nodes, respectively, and T is a relatively high cosine similarity threshold that can be adjusted according to the dataset.
[0061] To generate an adaptive trigger node similar to the attached node, the adaptive feature generator takes the node feature of the attached node as input. Specifically, unlike the adversarial attack that learns and generates adversarial perturbations for each graph, the present invention designs and trains a feature generator to mislead the model to capture the hidden association between the trigger and the target class. The way to generate the injection node feature is expressed as:
[0062] x u =σ(W1·x υ +b1)W2+b2, (5)
[0063] where W1, W2, b1, and b2 are learnable parameters of the feature generator and σ(·) is the activation function.
[0064] (3) Trigger injection: For each graph G to which trigger points are to be injected, the present invention obtains M attachment nodes according to the attachment node selection. Subsequently, the trained feature generator is used to adaptively generate corresponding node features for each trigger node. Then, the mixing function m(·) uses the M trigger nodes with generated features as g t Inject its attached nodes while keeping the original edges and nodes unchanged. Formally, the backdoor graph can be expressed as:
[0065]
[0066] Where G represents a clean graph and m(·) represents the trigger g t A hybrid function that is injected into a given graph to generate a trigger embedding graph, where M represents the number of attached nodes and A Ggt represents the adjacency matrix of the trigger embedding graph, A G Represents the adjacency matrix of the clean graph.
[0067] To ensure that the backdoor attack is not detected, the present invention randomly samples a certain number of candidate graphs to obtain the backdoor graphs; these graphs are placed in the training data set, and they participate in the target model training phase to inject hidden Trojans (i.e., the trigger mode of the present invention). Once the backdoor model training is completed, in the testing phase, a trigger can be inserted on any graph to activate the Trojan. It is worth noting that due to the large difference in the size of the rumor propagation tree (RPT), the present invention flexibly controls the trigger node size (i.e., the injection budget) for different samples, and adjusts it according to the number of nodes in each graph. The ratio of the trigger node size to the graph size may be different at different stages.
[0068] Feature loss function: Considering the similarity constraint in formula (4), the key idea is to ensure that the attached nodes are connected to the triggering nodes with high cosine similarity to improve the concealment of the triggering nodes and the reliability of the edges. The present invention defines a differentiable similarity loss to help optimize the feature generator:
[0069]
[0070] Where T represents the similarity threshold. This loss is applied between all attached nodes and their triggering nodes to ensure that the generated triggering nodes meet the similarity constraint. During the training phase, the present invention randomly selects attached nodes to reduce the risk of overfitting.
[0071] For a trained model, the candidate graphs before trigger injection tend to be classified as their true labels. Under the constraints of structure, content, and budget, the attacker needs the model to associate the trigger pattern with the target label as closely as possible. Therefore, to ensure the effectiveness of generating trigger nodes, the present invention optimizes the adaptive trigger feature generator to attack the surrogate model. θ *. Specifically, the attack loss can be defined as:
[0072]
[0073] Among them, f θ *(·) y and f θ *(·) yt Represent the true label y and the target label y respectively t The log-probability output.
[0074] Optimization: The present invention divides the dataset C into two parts: target class y t Figure C[y t ] and other graphs in the class C[\y t ],C[\y t ] is used to train the feature generator. L atkThe trigger point is forced to be embedded in the graph C[y t ] have similar embeddings. s Ensure similarity constraints. Therefore, the attack target of the present invention can be expressed as the following optimization problem:
[0075]
[0076] Among them, α is a trade-off hyperparameter used to balance the two terms of the loss function; It is the loss function for training the rumor detector (the proxy model in the attack of this invention).
[0077] Experimental proof:
[0078] The present invention conducts experiments on three real-world rumor datasets, including Twitter16, Twitter15 and Pheme; detailed statistics are shown in Table 1. Non-rumor or False Rumor is selected as the target category because attackers usually hope that rumors can bypass detection.
[0079] Table 1 Dataset statistics
[0080]
[0081] Surrogate Models and Rumor Detectors:
[0082] GCN (graph convolutional neural network), GAT (graph attention network), Graph-SAGE (a graph neural network model for graph node embedding learning) are three classic GNN models, which are often used in graph classification tasks. We attack them and use them as proxy models to attack specialized rumor detection models to test the effectiveness of the attacks; BiGCN is a GCN-based model that uses rumor propagation and diffusion to capture the global structure of the rumor tree; EBGCN uses the Bayesian method to capture robust structural features and adaptively reconsider the reliability of potential relationships. Their performance on clean datasets is shown in Table 2:
[0083] Table 2 Accuracy of the model on the clean dataset
[0084]
[0085] Two metrics are used to evaluate the effectiveness and stealth of the attack. The first metric is the attack success rate (ASR), which measures the success rate of the backdoor rumor detector in predicting the trigger embedding graph as a specified category:
[0086]
[0087] The second indicator is Clean Accuracy Drop (CAD), which is used to measure the difference in prediction classification accuracy between the clean rumor detector and the backdoor rumor detector on the clean test dataset.
[0088] Since there have been no previous attempts to attack propagation-based detectors, we select five existing graph classification backdoor attack methods to verify the performance of IBAtack, and compare them with its heuristic variants IBAtack / F and IBAtack / S as benchmarks: (1) ER-B generates universal triggers through the Erd"os-R'en yi (ER) model; (2) MIA selects the most important nodes from the GNN interpreter and replaces their connections with those of the triggers in the graph; (3) GTA updates the trigger generator using a two-layer optimization algorithm to generate an adaptive subgraph as the trigger of the graph; (4) TRAP performs perturbation actions on the graph structure through a gradient-based score matrix to generate perturbation triggers for the graph; (5) Motif uses the distribution differences of graphs in the dataset to search for trigger structures. IBAtack / F removes the feature generation module and randomly selects features from other nodes in the graph as the injected trigger node features; IBAtack / S removes the node selection module to randomly select the trigger injection location.
[0089] The dataset splitting rule is applied to all methods. According to the setting of BiGCN, each dataset is divided into two parts: 80% as training dataset and 20% as test dataset. 5% of the non-target classes in the training set are randomly selected as backdoor graphs. The trigger size is set to 10% and 20% of the number of graph nodes during training and testing, respectively. The epoch and learning rate in the training phase are set to 100 and 0.001, respectively. All baselines use the optimal parameters.
[0090] Experimental analysis:
[0091] Effectiveness and concealment of attacking rumor propagation trees: IBAtack is compared with baseline methods in terms of effectiveness and concealment on three real-world datasets. For fairness, GCN is used as the attacked model. The results are shown in Table 3:
[0092] Table 3 Attack results compared with baseline methods (ASR (%) | CAD (%))
[0093]
[0094] According to Table 3, we can get the following: (1) All the baseline methods of backdoor attacks, except GTA, fail on the rumor propagation tree even if the tree restriction is removed. These methods operate on the structure as a trigger. However, since the nodes of the propagation graph are the comment information about the source post, connecting two nodes that are not directly related will not have a great impact on the final graph representation. Although GTA has achieved relatively high performance due to the modification of the trigger features, its trigger structure in the form of a subgraph is still very easy to detect and limits its performance. In contrast, the method of the present invention obtains the best results due to the simultaneous consideration of the structure and feature characteristics of the propagation tree; (2) Compared with IBAtack / F and IBAtack / S, the method of the present invention achieves better results, which proves the effectiveness of the additional node selection module and feature generation module in the framework of the present invention. When the node features are randomly selected, the attack fails like other baseline methods. This shows that the injected attributes will greatly affect the effectiveness of the trigger because the propagation tree restricts the perturbation of the structure. In addition, the location of the selected trigger also improves the performance of the model to a certain extent.
[0095] Performance on different rumor detectors. Table 4 shows the results of testing the method of the present invention on different rumor detection models. The method of the present invention achieves extremely high ASR and extremely low CAD on all models. The present invention achieves almost 100% on Pheme, probably because it is relatively easy to attack as a binary classification task. The performance of attacking GAT on Twitter15 and Twitter16 is slightly worse, because the attack of the present method always adds edge nodes to the graph, and these nodes are often assigned smaller weights by the attention mechanism of GAT, making them less efficient than other models. It is worth mentioning that the present method also achieves excellent results on the robust model EBGCN. This is due to the similarity constraint, which reduces the suspicion of malicious triggering nodes, thereby increasing the confidence of the edges, which proves the robustness of the attack of the method of the present invention.
[0096] Table 4 Attack results on different rumor detection models
[0097]
[0098] Table 5 Attack results of proxy models using different architectures (ASR (%) | CAD (%))
[0099]
[0100] In addition, Table 5 shows the results of attacking rumor detectors with different architectures using a simple proxy model.
[0101] The attack method of the present invention performs well in a black-box setting and shows transferability across different architectures. Using different models as proxy models does not significantly degrade the attack performance, making the attack applicable to all propagation-based rumor detectors. And although the performance is slightly worse when attacking GAT, using GAT as a proxy model does not affect the effectiveness of the attack.
[0102] Experiments are conducted to explore the impact of different node importance metrics. Table 6 shows the ASR and CAD when attacking each model using other node importance metrics on Twitter16. Closeness centrality is calculated as the inverse of the sum of the shortest path lengths between a node and all other nodes in the graph. PageRank centrality is widely used in web page ranking, based on the principle that the importance of a web page on the Internet depends on the number and quality of its inbound links. Among all metrics, node similarity to the root node and degree centrality achieved better results on different datasets. Therefore, these two metrics are selected as the metrics for importance scoring in IBAtack.
[0103] Table 6 Attack results of different node importance measures (ASR (%) | CAD (%))
[0104]
[0105] The present invention conducts experiments on the number of trigger nodes (ie, trigger sizes) with different budgets in training and testing to explore the attack performance of IBAtack. Figure 4 The ASR of attacking BiGCN on two datasets is shown. As the ratio of trigger nodes to graph nodes in the test set increases, the ASR gradually increases, while the trend of the ratio of trigger nodes in the training set is the opposite, and this trend applies to all datasets. Regarding CAD, only the changes in the ratio of trigger nodes in the training set are reported, because it is only affected during the training phase, as Figure 5 As shown, although no obvious pattern emerges, it can be observed that generally larger values correspond to a decrease in CAD. Except when the trigger node ratio on GAT exceeds 0.1, the overall CAD is less than 2%. Therefore, 0.1 is selected in the training phase and 0.2 is selected in the testing phase to balance effectiveness and concealment.
[0106] The present invention also carried out experiments to explore the influence of poisoning rate, and the results are as follows Figure 6 As shown. Figure 6 As shown in (a), as the poisoning rate increases, the ASR also increases; however, the increment of ASR decreases as the poisoning rate increases, which indicates that the model of the present invention does not rely on a higher poisoning rate to succeed; Figure 6As shown in (b), as the poisoning rate increases, there is an overall upward trend, but it remains below 2% in most cases. This trend applies to all datasets.
[0107] The present invention further studies the impact of hyperparameters α and T on the performance of IBAttack. α controls the weight of similarity loss when training the feature generator, and T controls the threshold of similarity constraint. α and T are set to {0, 100, 500, 800, 1000} and {0.8, 0.85, 0.9, 0.95, 1} respectively. The attack success rate (ASR) and clean accuracy degradation (CAD) of attacking BiGCN on Twitter16 are shown in Figure 1. Figure 7 As shown. Figure 7 (a) It can be observed that (i) in general, as α increases, the ASR first increases and then decreases to stabilize. At the beginning, when the similarity loss is minimal, the generator cannot capture the similarity relationship between the trigger node and the attached node, which makes the characteristics of the trigger more irregular, thereby weakening the performance of the attack. However, when the similarity loss exceeds a certain level, the similarity actually reaches the set threshold. From this point on, larger α values will not significantly affect the training of the trigger. (ii) As T increases, the ASR gradually decreases. This is because a larger threshold imposes stricter constraints on the attack budget of the present invention. In the extreme case (ie, T = 1), this means injecting nodes that are exactly the same as the attached post, thereby degenerating the attack of the present invention into non-malicious. As Figure 7 As shown in (b), whether it is α or T, CAD fluctuates within a certain range, but in most cases remains below 2%. In practice, T can be set according to the average similarity of the data set. In addition, the present invention also studies the impact of the hyperparameter β on the performance of IBAttack. In the additional node selection module, β is a fusion factor used to control the ratio of the importance metrics of two nodes. β is set to {0, 0.25, 0.5, 0.75, 1}, Figure 8 (a) Reports the ASR attack on each rumor detection model on Twitter16. Figure 8 (b) The CAD of each rumor detection model attacking Twitter16 is reported. For BiGCN, degree centrality is more effective, while for GCN and EBGCN, similarity to the root node is a better metric. For GAT and GraphSAGE, the combination of these two methods produces better results. Except for GAT, the CAD fluctuations of all models are relatively small, and the present invention selects the optimal β value for different models.
[0108] The present invention proposes an injection-based backdoor attack framework, which is applied to a propagation-based rumor detector. Specifically, a node selection module is used to select key nodes as additional nodes to make full use of the attack budget, and the features of the triggering nodes are adaptively generated by using attack loss and similarity loss to ensure effectiveness and concealment.
[0109] Based on the same inventive concept, the present invention also proposes a backdoor attack generation system for rumor detectors, comprising:
[0110] The acquisition module is used to obtain a rumor propagation graph; the rumor propagation graph includes a graph node set consisting of source posts and their comments and an edge set consisting of the relationship between each comment or between a source post and a comment.
[0111] A generation module is used to generate the importance scores of the node structure and features of the rumor propagation graph by using degree centrality and the similarity between the node and the root node, wherein degree centrality is used to measure the degree of the node as the measure of the node centrality, the root node represents the source post, the node represents the comments of the source post, and the degree of the node represents the number of edges connected to each node; after sorting the importance scores, the M nodes with the highest value are selected as the attachment nodes, and each attachment node is attached with a trigger node; wherein the comments newly added to the propagation graph are used as the trigger nodes; an adaptive feature generator is constructed, which takes the node features of the attachment nodes as input to generate aggressive adaptive trigger nodes, and adds similarity constraints to the node features so that the trigger nodes can adaptively approach the node features of the attachment nodes; the adaptive feature generator is used to adaptively generate corresponding node features for each trigger node, and the trigger nodes with node features are injected into the attachment nodes to generate a backdoor graph.
[0112] The attack module is used to attack the propagation-based rumor detector through the generated backdoor graph to test the security of the rumor detector.
[0113] The present invention also proposes a computer device for generating a backdoor attack for a rumor detector, comprising: a memory, a processor, and a computer program stored in the memory. When the processor executes the computer program, the steps of the method for generating a backdoor attack for the rumor detector are implemented.
[0114] The present invention also proposes a readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, they are used to execute the steps of a backdoor attack generation method for a rumor detector.
[0115] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A backdoor attack generation method for rumor detectors, characterized in that: The following steps are involved: Obtain a rumor propagation graph; the rumor propagation graph includes a graph node set consisting of source posts and their comments and an edge set consisting of the relationship between each comment or between the source post and the comment; The importance scores of the node structure and features of the rumor propagation graph are generated respectively by using the degree centrality and the similarity between the node and the root node, wherein the degree centrality is used to take the degree of the node as a measure of the node centrality, the root node represents the source post, the node represents the comment of the source post, and the degree of the node represents the number of edges connected to each node; After sorting the importance scores, the top M nodes are selected as attachment nodes, and each attachment node is attached with a trigger node; the newly added comments outside the original propagation graph are used as trigger nodes; construct An adaptive feature generator, wherein the adaptive feature generator takes the node features of the attached node as input, generates an aggressive adaptive trigger node, and adds similarity constraints to the node features so that the trigger node can adaptively approach the node features of the attached node; Use an adaptive feature generator to adaptively generate corresponding node features for each trigger node, inject the trigger node with node features into the attached node, and generate a backdoor graph; The generated backdoor graph is used to attack the propagation-based rumor detector to test the security of the rumor detector.
2. A method for generating a backdoor attack for a rumor detector according to claim 1, characterized in that: The degree centrality and the similarity between the node and the root node are used to generate the importance score S(v) of the node structure and feature of the rumor propagation graph, respectively, which is expressed as: Among them, sim(·) represents the cosine similarity function, is the node centrality value of the node, β is the fusion factor used to control the ratio of the two indicators, x root and x v Represent the characteristics of the root node and the attached node respectively.
3. A method for generating a backdoor attack for a rumor detector according to claim 1, characterized in that: The constraints of the adaptive feature generator are expressed as: Among them, x u 、x v They represent the features of the generated trigger node u and the attached node v, T represents the cosine similarity threshold, and E g Represents the edge set containing the edges connecting the trigger node and the attachment node.
4. A method for generating a backdoor attack for a rumor detector according to claim 3, characterized in that: The adaptive feature generator is used to adaptively generate corresponding node features for each trigger node, which is expressed as: x u =σ(W1·x v +b1)W2+b2 where W1, W2, b1, and b2 are parameters that the feature generator can learn, and σ(·) is the activation function.
5. A method for generating a backdoor attack for a rumor detector according to claim 1, characterized in that: The backdoor diagram G gt It is expressed as: Where G represents a clean graph and m(·) represents the trigger g t A hybrid function that is injected into a given graph to generate a trigger embedding graph, where M represents the number of attached nodes and A Ggt represents the adjacency matrix of the trigger embedding graph, A G Adjacency matrix representing the clean embedded graph.
6. A backdoor attack generation system for rumor detectors, characterized in that: include: An acquisition module is used to acquire a rumor propagation graph; the rumor propagation graph includes a graph node set consisting of source posts and their comments and an edge set consisting of the relationship between each comment or between the source post and the comment; A generation module, used to generate the importance scores of the node structure and features of the rumor propagation graph respectively by using the degree centrality and the similarity between the node and the root node, wherein the degree centrality is used to take the degree of the node as the measure of the node centrality, the root node represents the source post, the node represents the comment of the source post, and the degree of the node represents the number of edges connected to each node; After sorting the importance scores, the top M nodes are selected as attachment nodes, and each attachment node is attached with a trigger node; the newly added comments outside the original propagation graph are used as trigger nodes; construct An adaptive feature generator, wherein the adaptive feature generator takes the node features of the attached node as input, generates an aggressive adaptive trigger node, and adds similarity constraints to the node features so that the trigger node can adaptively approach the node features of the attached node; Use an adaptive feature generator to adaptively generate corresponding node features for each trigger node, inject the trigger node with node features into the attached node, and generate a backdoor graph; The attack module is used to attack the propagation-based rumor detector through the generated backdoor graph to test the security of the rumor detector.
7. A computer device for generating backdoor attacks for rumor detectors, characterized in that: include: A memory, a processor, and a computer program stored in the memory, wherein when the processor executes the computer program, the steps of the backdoor attack generation method for a rumor detector as described in any one of claims 1 to 5 are implemented.
8. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, they are used to execute the steps of the backdoor attack generation method for a rumor detector as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Rumor propagation model construction method with anti-attack mechanism
CN115051923A
Rumor detection framework based on adaptive data enhancement and antagonism training
CN116795989A
Rumor detection method and system based on causal feature discovery
CN118708933A
Serverless mutual authentication
WO2023183925A1
AU2021102006A4
Cited By
Back door detection method of text map diffusion model based on attention transfer
CN121167722A
A backdoor detection method of a text-to-image diffusion model based on attention shift
CN121167722B