Causal influence discovery-based disguised topological element attack method

Through the causal influence discovery and hierarchical filtering mechanism, combined with cross-entropy loss and confidence minimization, the attack method of graph neural network is optimized, and the problems of resource waste and inefficiency in the existing technology are solved, achieving more efficient graph structure perturbation and model damage.

CN120263530APending Publication Date: 2025-07-04SOUTHEAST UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510619680.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing attack methods are unreasonable in the graph neural network, and the failure to accurately select key topological elements, resulting in waste of resources and inefficient attacks, and it is difficult to effectively perturb the graph structure within a limited budget to reduce model performance.

Method used

A hierarchical filtering mechanism of causal influence discovery is adopted, combining cross-entropy loss and confidence minimization, key nodes are evaluated through causal effects, and graph structural perturbations are optimized to maximize attack effect.

Benefits of technology

It realizes accurate selection of vulnerable nodes and edges within a limited budget, improves the efficiency and stability of attacks, enhances the destructiveness of graph neural networks, and improves the success rate and resource utilization efficiency of attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263530A_ABST
    Figure CN120263530A_ABST
Patent Text Reader

Abstract

The invention discloses a causal influence discovery-based disguised topological element attack method. The method aims at solving the problems that an existing attack method does not distinguish a disturbance total graph, neglects attack budget and node differences, does not select key topological elements, is low in resource utilization rate, is limited in success rate and the like. According to the method, graph structure elements susceptible to attack are discovered and perceived by means of causal influence, fragile nodes are identified by using a layered filtering mechanism, and accurate allocation of resources is realized; integrating causal effect evaluation to determine key nodes; and carrying out optimal element operation by adopting a joint loss optimization disturbance method so as to maximize the attack effect in a prediction. Experiments show that CTEA is superior to an existing method in efficiency and effect, and the CTEA shows adaptability, robustness, accuracy and resource efficiency on multiple GNN architectures and data sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a graph neural network adversarial attack technology, belonging to the cross - technical field of computer network security and artificial intelligence, and particularly relates to a method for attacking disguised topological elements based on causal influence discovery. Background Art

[0002] In recent years, research on adversarial attacks against graph - structured datasets has emerged in an endless stream. [1] . These methods can be divided into poisoning attacks [5] and evasion attacks [6] . Among them, poisoning attacks occur in the training stage, and evasion attacks occur in the testing stage. The purpose of evasion attacks is to disrupt the input of the GNN during the inference stage, thereby reducing the performance of the victim model. Therefore, the more the classification accuracy of the victim model decreases, the more successful the attack is. The most direct method is random sampling, randomly perturbing the graph structure or node features of a clean graph to generate adversarial samples. However, its attack performance is usually not as good as that of more strategically designed perturbations.

[0003] In addition, according to the knowledge of the attacker, attacks can be divided into white - box attacks, grey - box attacks, and black - box attacks. In grey - box evasion attacks, the attacker can only obtain the training labels and the adjacency matrix, and cannot obtain any knowledge of the target model. Most existing GNN attacks belong to white - box or grey - box attacks. Usually, in this case, the attacker first uses a surrogate model, such as Simplified Graph Convolution (SGC), and uses white - box attacks to generate perturbations. For example, a target - free evasion attack for graph classification and node classification is designed using reinforcement learning techniques. Zugner [1] proposed a targeted evasion attack called Nettack for a two - layer GCN and demonstrated state - of - the - art attack performance. They constructed a surrogate linear model for the GCN by omitting the ReLU activation function. In addition, they also defined a graph structure that can maintain perturbations and minimized the difference in node degree distribution between the graphs before and after the attack. To improve the concealment of the attack, Ma et al. [7] proposed ReWatt, which re - defines the action space of reinforcement learning. The perturbations generated by ReWatt are still imperceptible because the number of edges in the adversarial sample is the same as that in the clean sample. In addition to these attack strategies, if the opponent knows the specific operations existing in the victim model, they may also target specific components of the GNN. For example, some classic methods add hierarchical pooling operations to the model, and these operations will select key topological elements to reduce the number of nodes and determine the structure of the coarse - grained graph [8] .

[0004] Although these existing methods are very excellent in terms of attack performance, most of them lack concealment or rely on specific model architectures, resulting in insufficient generality in the impact on different types of GNNs. In addition, these methods often lack pertinence when selecting perturbed edges, easily wasting valuable attack budgets. In contrast, CTEA can more accurately locate vulnerable nodes and edges through hierarchical node filtering and causal effect analysis. It also shows strong operability on multiple GNN models and can maintain excellent attack effects, so it is more efficient and general.

[0005] [1] Z¨ugner, D., Akbarnejad, A., G¨unnemann, S., 2018. Adversarial attacks on neural networks for graph data, in: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Association for Computing Machinery, New York, NY, USA. p. 2847–2856.

[0006] [2] Dai, H., Li, H., Tian, T., Huang, X., Wang, L., Zhu, J., Song, L., 2018. Adversarial attack on graph structured data, in: International conference on machine learning, PMLR. pp. 1115–1124.

[0007] [3] Tao, H., Cao, J., Chen, L., Sun, H., Shi, Y., Zhu, X., 2024a. Black-box attacks on dynamic graphs via adversarial topology perturbations. Neural Networks 171, 308–319.

[0008] [4] Wang, H., Liu, T., Sheng, Z., Li, H., 2024b. Explanatory subgraph attacks against graph neural networks. Neural Networks 172, 106097.

[0009] [5] Wang, B., Gong, N. Z., 2019. Attacking graph-based classification via manipulating the graph structure, in: Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security, pp. 2023–2040.

[0010] [6] Wang, B., Lin, M., Zhou, T., Zhou, P., Li, A., Pang, M., Li, H., Chen, Y., 2024a. Efficient, direct, and restricted black-box graph evasion attacks to any-layer graph neural networks via influence function, in: Proceedings of the 17th ACM International Conference on Web Search and Data Mining, pp. 693–701.

[0011] [7] Ma, Y., Wang, S., Derr, T., Wu, L., Tang, J., 2021. Graph adversarial attack via rewiring, in: Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pp. 1161–1169.

[0012] [8] Ying, Z., You, J., Morris, C., Ren, X., Hamilton, W., Leskovec, J., 2018. Hierarchical graph representation learning with differentiable pooling. Advances in neural information processing systems 31.

[0013] [9]Zhang, Z., Bu, J., Ester, M., Zhang, J., Yao, C., Yu, Z., Wang, C., 2019. Hierarchical graph pooling with structure learning. arXiv preprint arXiv:1911.05954. Summary of the Invention

[0014] Technical Problem:

[0015] Unreasonable allocation of attack resources: Existing attack methods often perturb the entire graph indiscriminately, without considering the limited attack budget and the differences in node vulnerabilities, resulting in low resource utilization efficiency. When selecting attack positions, a global optimization method is adopted, which increases the computational complexity due to the large scale of the graph and causes resource waste. This application aims to achieve more accurate element selection.

[0016] Difficulty in accurately selecting key topological elements: Existing methods rely on local gradients, heuristic rules, or scoring functions to select perturbation positions, making it difficult to comprehensively capture the actual impact of nodes and edges on the model's classification performance, unable to effectively balance the perturbation budget and the attack effect, and prone to over - allocating resources to unimportant positions, affecting the attack efficiency and stability. This application aims to achieve a more stable and efficient attack effect.

[0017] Technical Solution:

[0018] The present invention discovers vulnerable graph structure elements through causal influence perception, identifies vulnerable nodes using a hierarchical filtering mechanism to achieve precise resource allocation; integrates causal effect evaluation to determine key nodes; and adopts a joint - loss - optimized perturbation method to perform optimal element operations to maximize the attack effect within the budget.

[0019] The specific method is as follows:

[0020] A camouflaged topological element attack method based on causal influence discovery, comprising the following steps:

[0021] Step 1, select multiple graph structure data sets. Determine the target GNN model for evaluation and select a suitable proxy model;

[0022] Step 2, according to the trained GNN model, extract node prediction logits. Use the proxy model to calculate the node prediction probability distribution, and perform preliminary screening through the classification confidence margin CCM, retain the nodes with high classification confidence, concentrate the attack targets on the nodes that the model considers to be stably classified, and avoid wasting the attack budget on error - prone nodes. Refine the screening of the nodes to be attacked based on the node degree and the confidence boundary;

[0023] Step 3: Calculate the causal effect of nodes on the model prediction by means of counterfactual analysis in causal inference. By intervening in the attention map, calculate the output difference of the model under different attention values to obtain the direct causal effect of attention, and thus evaluate the impact of nodes on the final prediction. Optimize the causal effect metric, use the identity matrix as the counterfactual attention tensor, promote the attention model to allocate attention more accurately, and screen out the final set of attack candidate nodes;

[0024] Step 4: Construct an attack loss function by combining maximizing cross-entropy loss and minimizing confidence. Cross-entropy loss measures the difference between the model's predicted distribution and the true label. Maximizing this loss can make the model produce incorrect predictions on the perturbed graph; confidence represents the degree of certainty of the model's prediction. Minimizing confidence can interfere with the prediction stability of the model. Combine the two, and adjust the hyperparameter λ to balance the weights to enhance the attack effect;

[0025] Step 5: Based on the attack loss function, calculate the gradient with respect to the graph adjacency matrix. Select the edge with the largest absolute value of the gradient for modification. If the gradient is positive and the edge does not exist, add it; if the gradient is negative and the edge exists, delete it. Update the adjacency matrix after each modification, recalculate the cross-entropy loss and confidence, and continue to iterate until the attack budget is reached to generate adversarial samples.

[0026] Furthermore, the specific implementation method of Step 2 is as follows.

[0027] When attacking a GNN, due to the limitation of the attack budget, it is necessary to preferentially perturb the nodes that have the greatest impact on the model to minimize the overall classification performance as much as possible. However, the attacker cannot directly obtain the parameters or true labels of the target model, so a surrogate model is used to calculate the logarithm to estimate the classification confidence, so as to screen out potential attack targets. Specifically, the surrogate model calculates the predicted probability distribution of node v. Use the classification confidence measure (CCM) to measure the stability of node classification:

[0028]

[0029] where and are the highest and second-highest logits values output by the surrogate model for point v. CCM reflects the certainty of the model's classification decision. The larger the CCM value, the more stable the classification, and the smaller the CCM value, the more vulnerable the classification result of the node is to interference. According to this indicator, conduct a preliminary screening and only retain those nodes with high classification confidence:

[0030] F ccm (V) = {v|CCM(v) > threshold ccm , v ∈ V U (Equation 2)

[0031] Among them, threshold ccm is the classification confidence threshold, and V U is the set of nodes in the initial dataset. This filtering step ensures that the attack targets are concentrated on the nodes that the model considers to be stably classified, making the attack more destructive, while avoiding wasting the attack budget on error-prone nodes.

[0032] The specific implementation method in Step 3 is as follows.

[0033] Research shows that low-degree nodes are more vulnerable to structural perturbations because their feature propagation ability drops faster when their neighborhood information decreases. In contrast, high-degree nodes are more resistant to attacks because they have more neighbors and the classification results are less dependent on local topological changes. Based on this observation, the filtered node set F ccm (V) is further screened:

[0034] F deg (V) = {v∣Degree(v) < threshold deg , v ∈ F ccm (V)} (Equation 3)

[0035] Among them, threshold deg is the degree threshold used to control whether the selected nodes are sparse enough. Degree(v) is the degree of node v. Through this screening step, it can be ensured that the attack is preferentially concentrated on the nodes with a weaker topological structure, so that the removal of edges has a greater impact on the classification performance of the model and improves the overall destructive power of the attack. Further, confidence boundary filtering is introduced to screen out the nodes that are most vulnerable to perturbations. Since the classification confidence boundary calculated by the surrogate model logits can effectively measure the stability of classification, a threshold mar is set for further screening:

[0036] F mar (V) = {v∣CCM(v) < threshold mar , v ∈ F deg (V)} (Equation 4)

[0037] Among them, threshold mar is a threshold set empirically to distinguish between stably classified nodes and nodes vulnerable to interference. This filter ensures that the attack targets are ultimately concentrated on those nodes with low classification credibility and fragile topological structures, thus maximizing the attack effect.

[0038] Furthermore, the specific implementation method of Step 4 is as follows.

[0039] In the present invention, the performance trade-off between the primary task and the secondary task is alleviated by explicitly maximizing the causal effect of attention on the primary task. Suppose a simple case of performing a node classification task using a typical L-layer GAT. Every l layers, the node representations of the previous layer are used as inputs, where Z is the node feature, R is the real number field, n is the number of nodes in the graph, and d is the dimension of the node feature. Then, the factual attention map A l is used for feature aggregation and update to obtain the factual output feature Z l = f(Z l-1 , X l ), where X is the attention map. Similarly, when disturbing the attention map of the l-th layer (for example, assigning dummy values using the do(·) operation), the counterfactual output feature is obtained. Further, a learnable matrix (c represents the number of classes) is used to obtain the factual predicted label of the node and the counterfactual predicted label Note that the causal effect of the l-th layer is Therefore, the causal effect can be used as a supervision signal to explicitly indicate the attention acquisition process. For GAT, the new objective can be formulated as follows:

[0040]

[0041] where y is the true label, λ l is the balanced training coefficient, is the entropy loss, represents the original objective, such as the standard classification loss, represents the result obtained after the causal effect of the attention mechanism of the l-th layer acts on the model prediction and is used to calculate the cross-entropy loss. Note that equation (5) is the general format for shaping the causal effect of evaluating node influence, and an additional loss is calculated at each GAT layer to directly supervise the attention. In practical applications, it is found that simply selecting one or two layers is sufficient to bring satisfactory performance improvement by CTEA.

[0042] Furthermore, the specific implementation manner of step 5 is as follows

[0043] First, the present invention is optimized based on the cross-entropy loss with the aim of maximizing the classification error of the model. The cross-entropy loss measures the difference between the model prediction distribution and the true label. The larger the cross-entropy loss, the greater the error of the model in the classification task. For node v, its calculation formula is

[0044]

[0045] Where \(V\) represents the set of all nodes in the graph, \(K\) is the number of categories in the classification task, and \(y\) vk is the true label of node \(v\) belonging to category \(k\), and \(\hat{y}\) is the predicted probability of the surrogate model for node \(v\) belonging to category \(k\). By maximizing the cross-entropy loss, the target model can be forced to make incorrect predictions on the perturbed graph, thereby affecting its classification accuracy.

[0046] However, simply maximizing the cross-entropy loss is not sufficient to ensure the stability of the attack. Especially when the target model has a high confidence in the predictions of certain nodes or edges, this may cause the model to respond slowly to perturbations. To improve the effectiveness of the attack, confidence minimization (CM) is introduced as an auxiliary optimization objective. Confidence is the degree of trust the model has in its predictions, usually represented as the maximum probability of the predicted class by the model. In graph classification or node classification tasks, the higher the confidence of the model, the more certain its prediction is, while the lower the confidence, the more unstable the model's prediction is and the more vulnerable it is to perturbations. Therefore, the goal is to minimize the confidence of the model in the classification results, which will further disrupt the prediction stability of the model and make it more vulnerable to adversarial perturbations. The objective function of confidence minimization is defined as

[0047]

[0048] Introducing this objective is to reduce the classification confidence of the target model, so that the model has a lower prediction confidence in adversarial samples, thereby increasing the success rate of the attack. To optimize the cross-entropy loss and confidence minimization simultaneously, these two objectives are combined to construct the final attack loss function

[0049]

[0050] where \(\lambda\) is a hyperparameter used to balance the weights of the cross-entropy loss and confidence minimization. By adjusting \(\lambda\), the influence degrees of the cross-entropy loss and confidence minimization in the attack process can be controlled. When \(\lambda\) is large, the attack will pay more attention to confidence minimization, making the model's predictions more unstable. When \(\lambda\) is small, the attack mainly relies on maximizing the cross-entropy loss, thereby optimizing the error rate of the node classification results. During the attack process, the most influential edges are selected for perturbation by calculating the gradient of the loss function with respect to the graph adjacency matrix \(A\). The adjacency matrix \(A\) describes the connectivity between nodes in the graph, and the attacker disrupts the graph structure by modifying the presence or absence of these edges. To perform the attack efficiently, for each element \(A_{ij}\) in the adjacency matrix ij the gradient of the attack loss function is calculated:

[0051]

[0052] According to the magnitude of the gradient, select the edge with the largest absolute value of the gradient for modification. Specifically, if A ij and the gradient is positive, then add the edge (i, j); if A ij = 1 and the gradient is negative, then delete the edge (i, j). This method can ensure that each modification will have a greater impact on the target loss function, thereby effectively disrupting the classification results of the graphical model. The gradient of the adjacency matrix is calculated in each iteration, and important edges are selected for perturbation until the attack budget is reached. As the iteration progresses, the structure of the graph gradually changes, and finally a perturbed adversarial sample is generated and used for the evaluation of the classification task of the target model.

[0053] Beneficial effects:

[0054] The present invention adopts a hierarchical filtering mechanism, combines multiple factors such as node degree and classification confidence to screen attack targets, accurately locates nodes that are vulnerable and have a great impact on model performance, avoids waste of resources, and improves attack efficiency. By evaluating the causal effect to determine key nodes and edges, the attack is more targeted. Compared with existing methods that rely on local gradients, heuristic rules or scoring functions, it can capture the interaction between elements more accurately and effectively utilize perturbation resources. In practical applications, CTEA helps to discover potential security risks of GNN models, promotes the improvement and optimization of models, and improves their security and robustness in complex environments. Description of the drawings

[0055] Figure 1 is the flow chart of the present invention.

[0056] Figure 2 is the schematic diagram of the implementation process of the present invention. Detailed implementation manners

[0057] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.

[0058] Embodiment 1:

[0059] The purpose of the present invention is to conduct a targeted attack on a well-trained graph model. To improve the attack success rate, a camouflaged topological element attack is introduced, and the attack budget is concentrated on nodes that are vulnerable and have a great influence.

[0060] The overall framework of CTEA is as shown in Figure 1 and 2As shown, it mainly consists of three key parts: (a) hierarchical filtering for vulnerable nodes, (b) influence assessment through causal effects, and (c) objective-driven topology perturbation based on joint optimization. To effectively identify local attack targets, a hierarchical filtering strategy is adopted, aiming to precisely locate vulnerable elements. This strategy includes three steps: (1) classification confidence screening based on the surrogate model, (2) topological filtering based on node degree, and (3) target refinement screening based on the confidence boundary. Then, an effective causal effect evaluation strategy for node influence is introduced to select the nodes that have the relatively greatest impact on the model for the classification task. Finally, using the hierarchical filtering strategy and the causal effect evaluation of node influence, the third module systematically adds fake edges or deletes existing edges between vulnerable nodes and high-influence nodes according to the attack budget, thereby constructing an adversarial graph.

[0061] 1. Hierarchical filtering of vulnerable nodes

[0062] When attacking a GNN, due to the limitation of the attack budget, it is necessary to preferentially perturb the nodes that have the greatest impact on the model to minimize the overall classification performance. However, the attacker cannot directly obtain the parameters or true labels of the target model, so a surrogate model is used to calculate the logits to estimate the classification confidence, thereby screening out potential attack targets. Specifically, the surrogate model calculates the predicted probability distribution of node v. The classification confidence measure (CCM) is used to measure the stability of node classification:

[0063]

[0064] where and represent the highest and second-highest logits values output by the surrogate model for node v. CCM reflects the certainty of the model's classification decision. The larger the CCM value, the more stable the classification, and the smaller the CCM value, the more vulnerable the classification result of the node is to interference. According to this metric, a preliminary screening is carried out, and only those nodes with high classification confidence are retained:

[0065] F ccm (V) = {v|CCM(v) > threshold ccm , v ∈ V U (Equation 2)

[0066] where threshold ccm is the classification confidence threshold, and V U is the set of nodes in the initial dataset. This filtering step ensures that the attack targets are concentrated on the nodes that the model considers to have stable classifications, making the attack more destructive and avoiding wasting the attack budget on error-prone nodes.

[0067] Research shows that low-degree nodes are more vulnerable to structural perturbations because their feature propagation ability degrades faster when their neighborhood information decreases. In contrast, high-degree nodes are more resistant to attacks because they have more neighbors and the classification results are less dependent on local topological changes. Based on this observation, the filtered node set F was further screened. ccm (V):

[0068] F deg (V) = {v | Degree(v) < threshold deg , v ∈ F ccm (V)} (Equation 3)

[0069] where threshold deg is the degree threshold used to control whether the selected nodes are sparse enough, and Degree(v) is the degree of node v. Through this screening step, it can be ensured that the attack is concentrated on nodes with a weaker topological structure first, so that the removal of edges has a greater impact on the classification performance of the model and improves the overall destructive power of the attack. Confidence boundary filtering was further introduced to screen out the nodes most vulnerable to perturbations. Since the classification confidence boundary calculated by the surrogate model logits can effectively measure the stability of classification, a threshold mar was set for further screening:

[0070] F mar (V) = {v | CCM(v) < threshold mar , v ∈ F deg (V)} (Equation 4)

[0071] where threshold mar is a threshold set empirically to distinguish between nodes with stable classification and nodes vulnerable to interference. This filter ensures that the attack targets are finally concentrated on those nodes with low classification credibility and fragile topological structures, thus maximizing the attack effect.

[0072] 2. Impact assessment through causal effects

[0073] First, introduce how to derive the causal effect of attention from the GAT-based structural causal model. In particular, this requires the use of counterfactual analysis widely adopted in causal inference. Subsequently, based on the obtained causal effect, a scheme for combining it with the attention-based GAT training is introduced in detail. Since the direct causal effect of attention on the final prediction (i.e., the link X→Y) is calculated and used as a measure of the quality of attention, to a certain extent, it reflects the node effect that affects the final prediction. Therefore, the direct causal effect can be used as a measure to evaluate the impact of nodes on the final prediction, so as to find the nodes that have an impact on the final prediction.

[0074] As mentioned above, since it is not feasible to directly calculate the impact of nodes on the final prediction. Previous work mainly measures the influence of nodes through indicators such as degree centrality, median centrality, and proximity centrality, but these indicators can only reflect the importance of nodes to the graph structure itself and have little impact on the final prediction. Fortunately, recently emerging causal inference techniques provide an effective tool to evaluate the influence of nodes by analyzing the causal relationships between model variables. In this way, the causal effect of attention can be directly used to measure the influence. Given that the obtained causal effect mainly depends on the model itself, this makes it a more accurate and unbiased method for evaluating the actual impact effect.

[0075] To calculate the causal effect of node influence, the counterfactual analysis method widely used in causal inference is introduced. The most crucial point of counterfactual causality is as follows: Under a specific data background (i.e., node features Z), if the treatment (i.e., attention graph X) is not observed, what will the result (i.e., model prediction Y) be? To answer such a question, the values of several variables must be changed to observe their changes. This operation is called intervention in causal inference and can be expressed as do(·). During the do(·) operation, a counterfactual value must be selected to replace the original factual value of the intervened variable. Once a variable is intervened, all its input connections in the microcontroller will be cut off, and its value will be given separately, while other unaffected variables will maintain their original values. For example, in the example, do(X = x * ) means asking the attention X to take a non-factual value x * (e.g., reverse / random attention), so that the link Z→X is cut off, and X will no longer be affected by its causal parent variable Z at all. The mathematical formula for the intervention operation is as follows: This indicates that after the do(·) operation changes the attention value to x * , the output value of the model will also become

[0076] In summary, when a fictional value is assigned to the attention map so that all adjacent nodes of each self-node share the same attention weight, the feature clustering in the graph attention model degenerates into unweighted averaging. In this case, according to the theory of causal inference, the direct causal effect DCE of attention on the model prediction can be obtained by calculating the difference between the model result Y z,x and . Thus, the impact of a node on the final prediction (NI) can also be evaluated through the quality metric of attention, i.e.:

[0077] In the present invention, by explicitly maximizing the causal effect of attention on the main task, the performance trade-off between the main task and the secondary task is alleviated. Assume a simple case of performing a node classification task using a typical L-layer GAT. Every l layers, the node representations of the previous layer are used as input, where Z is the node feature, R is the real number field, n is the number of nodes in the graph, and d is the dimension of the node feature. Then, the factual attention map A l is used for feature aggregation and update to obtain the factual output feature Z l = f(Z l-1 , X l ), where X is the attention map. Similarly, when disturbing the attention map of the l-th layer (e.g., assigning dummy values using the do(·) operation), the counterfactual output feature is further obtained using the learnable matrix (c represents the number of classes) to obtain the factual predicted label of the node and the counterfactual predicted label Note that the causal effect of the l-th layer is Therefore, the causal effect can be used as a supervision signal to explicitly indicate the attention acquisition process. For GAT, the new objective can be formulated as follows:

[0078]

[0079] where y is the true label, λ l is the balance training coefficient, is the entropy loss, represents the original objective, such as the standard classification loss, represents the result obtained after the causal effect of the attention mechanism of the l-th layer acts on the model prediction and is used to calculate the cross-entropy loss. Note that Equation (5) is the general format for shaping the causal effect to evaluate the node impact, and an additional loss is calculated at each GAT layer to directly supervise the attention. In practical applications, it is found that simply selecting one or two layers is sufficient to bring satisfactory performance improvement by CTEA.

[0080] 3. Joint Optimization Based Target-Driven Topology Perturbation

[0081] The present invention proposes an iterative optimization attack method based on the combination of maximizing cross-entropy loss and minimizing confidence. This method is applicable to the gray-box attack scenario, that is, the attacker cannot directly obtain the parameters of the target model, but obtains the model prediction information by training a surrogate model and uses this information for perturbation optimization. The core objective of the attack is to force the target model to output incorrect classification results by modifying the structure of the graph and perturbing its edges, thereby minimizing the classification accuracy of the model to the greatest extent and increasing the randomness of the model prediction. First, the present invention is optimized based on cross-entropy loss to maximize the classification error of the model. Cross-entropy loss measures the difference between the model prediction distribution and the true label. The larger the cross-entropy loss, the greater the error of the model in the classification task. For node v, its calculation formula is

[0082]

[0083] where V represents the set of all nodes in the graph, K is the number of categories in the classification task, y vk is the true label of node v belonging to category k, is the predicted probability of node v belonging to category k by the surrogate model. By maximizing the cross-entropy loss, the target model can be forced to make incorrect predictions on the perturbed graph, thereby affecting its classification accuracy.

[0084] However, simply maximizing the cross-entropy loss is not sufficient to ensure the stability of the attack. Especially when the target model has a high confidence in the prediction of some nodes or edges, this may lead to a slow response of the model to perturbations. To improve the effectiveness of the attack, confidence minimization CM is introduced as an auxiliary optimization objective. Confidence is the degree of trust of the model in its prediction, usually expressed as the maximum probability of the model prediction category. In graph classification or node classification tasks, the higher the confidence of the model, the more certain its prediction is, while the lower the confidence, the more unstable the model's prediction is and it is more vulnerable to perturbations. Therefore, the goal is to minimize the confidence of the model in the classification result, which will further interfere with the prediction stability of the model and make it more vulnerable to adversarial perturbations. The objective function of confidence minimization is defined as

[0085]

[0086] Introducing this objective is to reduce the classification confidence of the target model, reduce the prediction confidence of the model in adversarial samples, and thus improve the success rate of the attack. To optimize the cross-entropy loss and confidence minimization simultaneously, these two objectives are combined to construct the final attack loss function:

[0087]

[0088] where λ is a hyperparameter used to balance the weights of cross - entropy loss and confidence minimization. By adjusting λ, the influence degrees of cross - entropy loss and confidence minimization during the attack can be controlled. When λ is large, the attack pays more attention to confidence minimization, making the model's predictions more unstable. While when λ is small, the attack mainly relies on maximizing the cross - entropy loss, thereby optimizing the error rate of the node classification results. During the attack, by calculating the loss function and the gradient with respect to the graph adjacency matrix A to select the most influential edges for perturbation. The adjacency matrix A describes the connectivity between nodes in the graph, and the attacker disrupts the graph structure by modifying the existence or non - existence of these edges. To perform the attack efficiently, for each element A in the adjacency matrix ij calculate the gradient of the loss function :

[0089]

[0090] According to the magnitude of the gradient, select the edge with the largest absolute value of the gradient for modification. Specifically, if A ij and the gradient is positive, add the edge (i, j); if A ij = 1 and the gradient is negative, delete the edge (i, j). This method ensures that each modification will have a greater impact on the target loss function, thus effectively disrupting the classification results of the graphical model. The gradient of the adjacency matrix is calculated in each iteration, and important edges are selected for perturbation until the attack budget is reached. As the iteration progresses, the structure of the graph gradually changes, and finally a perturbed adversarial sample is generated and used for the evaluation of the classification task of the target model.

[0091] The specific embodiments described in this article are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains can make various modifications or supplements to the described specific embodiments or use similar ways to substitute, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. A method for attacking camouflaged topological elements based on causal influence discovery, characterized in that The steps include: Step 1: Select multiple graph structure datasets; determine the target GNN model for evaluation and select a suitable proxy model; Step 2: Extract node prediction logits based on the trained GNN model; use the proxy model to calculate the node prediction probability distribution, perform preliminary screening through the classification confidence margin CCM, retain nodes with high classification confidence, and focus the attack targets on nodes that the model believes are stable in classification, avoiding wasting attack budget on error-prone nodes; refine and screen the nodes to be attacked based on node degree and confidence margin; Step 3: With the help of counterfactual analysis in causal inference, calculate the causal effect of the node on the model prediction; by intervening in the attention graph, calculate the output difference of the model under different attention values, and obtain the direct causal effect of attention, so as to evaluate the impact of the node on the final prediction; optimize the causal effect measurement, use the unit matrix as the counterfactual attention tensor, promote the attention model to allocate attention more accurately, and screen out the final set of candidate attack nodes; Step 4: Construct the attack loss function by combining the method of maximizing the cross entropy loss and minimizing the confidence. The cross entropy loss measures the difference between the model's predicted distribution and the true label. Maximizing this loss can cause the model to make incorrect predictions on the perturbation graph. The confidence indicates the degree of certainty of the model's prediction. Minimizing the confidence can interfere with the model's prediction stability. Combining the two, the attack effect can be enhanced by adjusting the hyperparameter λ to balance the weights; Step 5: Based on the attack loss function, calculate the gradient of the graph adjacency matrix; select the edge with the largest absolute value for modification according to the gradient size. If the gradient is positive and the edge does not exist, add it; if the gradient is negative and the edge exists, delete it; after each modification, update the adjacency matrix, recalculate the cross entropy loss and confidence, and continue to iterate until the attack budget is reached to generate adversarial samples.

2. The method for attacking camouflaged topological elements based on causal influence discovery according to claim 1, wherein: The specific implementation of step 2 is as follows: The proxy model calculates the predicted probability distribution of node v and uses the classification confidence CCM to measure the stability of node classification: Among them, and represent the highest and the second highest logits values output by the surrogate model of node v; CCM reflects the certainty of the model's classification decision. The larger the CCM value, the more stable the classification, and the smaller the CCM value, the more easily the classification result of the node is affected by interference. According to this indicator, a preliminary screening is carried out to retain only those nodes with high classification confidence: F ccm (V) = {v | CCM(v) > threshold ccm , v ∈ V U (Equation 2) Among them, threshold ccm is the classification confidence threshold, and V U is the set of nodes in the initial dataset. This filtering step can ensure that the attack targets are concentrated on the nodes that the model considers to be stably classified, making the attack more destructive, while avoiding wasting the attack budget on error-prone nodes.

3. The method for attacking disguised topological elements based on causal influence discovery according to claim 2, wherein: The specific implementation of step 3 is as follows: Further screen the filtered node set F ccm (V): F deg (V) = {v | Degree(v) < threshold deg , v ∈ F ccm (V)} (Equation 3) Among them, threshold deg is the degree threshold, which is used to control whether the selected nodes are sparse enough, and Degree(v) is the degree corresponding to node v; through this screening step, it is ensured that the attack is mainly concentrated on the nodes with weaker topological structures, so that the removal of edges has a greater impact on the classification performance of the model and improves the overall destructive power of the attack; further introduce confidence boundary filtering to screen out the nodes most vulnerable to perturbations; since the classification confidence boundary calculated by the surrogate model logits can effectively measure the stability of classification, a threshold mar is set for further screening: F mar (V) = {v | CCM(v) < threshold mar , v ∈ F deg (V)} (Equation 4) Among them, threshold mar is a threshold set based on experience, used to distinguish nodes with stable classification from nodes vulnerable to interference; it can ensure that the attack targets are finally concentrated on those nodes with low classification credibility and fragile topological structures, thus maximizing the attack effect.

4. The method for attacking camouflaged topological elements based on causal influence discovery according to claim 1, wherein: The specific implementation of step 4 is as follows: Mitigate the performance trade-off between the main task and the secondary task by explicitly maximizing the causal effect of attention on the main task; use a typical L-layer GAT to perform the node classification task. Every l layers, use the node representations of the previous layer as input, where Z is the node feature, R is the real number field, n is the number of nodes in the graph, and d is the dimension of the node feature; then, use the factual attention graph A l to perform feature aggregation and update to obtain the factual output feature Z l = f(Z l-1 , X l ), where X is the attention graph; Similarly, when disturbing the attention map of the l-th layer, the counterfactual output features can be obtained Further, a learnable matrix where c represents the number of classes, is used to obtain the factual prediction label of the node as well as the counterfactual prediction label The causal effect of the l-th layer is Therefore, the causal effect can be used as a supervision signal to explicitly indicate the attention acquisition process; for GAT, the new objective is formulated as follows: where y is the true label, and λ l is the balanced training coefficient, is the entropy loss, represents the original target, represents the result obtained after the causal effect of the l-th layer attention mechanism acts on the model prediction and is used to calculate the cross-entropy loss; Equation (5) is the general format for shaping the causal effect of the evaluation node's influence, and additional losses will be calculated at each GAT layer to directly supervise the attention.

5. The method for attacking disguised topological elements based on causal influence discovery according to claim 1, wherein: The specific implementation method in step 5 is as follows: First, optimize based on the cross-entropy loss to maximize the classification error of the model; the cross-entropy loss measures the difference between the predicted distribution of the model and the true labels; the larger the cross-entropy loss, the greater the error of the model in the classification task; for node v, its calculation formula is Among them, V represents the set of all nodes in the graph, K is the number of categories of the classification task, and y vk is the true label of node v belonging to category k, is the predicted probability of the surrogate model for node v belonging to category k; by maximizing the cross-entropy loss, the target model is forced to make incorrect predictions on the perturbed graph, thereby affecting its classification accuracy; However, simply maximizing the cross-entropy loss is not sufficient to ensure the stability of the attack. Especially when the target model has a high confidence in the predictions of certain nodes or edges, it will cause the model to respond slowly to perturbations. To improve the effectiveness of the attack, confidence minimization (CM) is introduced as an auxiliary optimization objective. Confidence is the degree of trust the model has in its predictions, expressed as the maximum probability of the model's predicted class. In graph classification or node classification tasks, the higher the confidence of the model, the more certain its prediction is, while the lower the confidence, the more unstable the model's prediction is and the more vulnerable it is to perturbations. Therefore, the goal is to minimize the confidence of the model in the classification results, which will further disrupt the prediction stability of the model and make it more vulnerable to adversarial perturbations. The objective function of confidence minimization is defined as Introduce the objective function is to reduce the classification confidence of the target model, so that the prediction confidence of the model for adversarial samples is reduced, thereby improving the success rate of the attack; in order to optimize the cross-entropy loss and confidence minimization simultaneously, these two objectives are combined to construct the final attack loss function where λ is a hyperparameter used to balance the weights of cross - entropy loss and confidence minimization; by adjusting λ, the influence degrees of cross - entropy loss and confidence minimization during the attack can be controlled; when λ is large, the attack will pay more attention to confidence minimization, making the model's predictions more unstable; while when λ is small, the attack mainly relies on maximizing the cross - entropy loss, thereby optimizing the error rate of node classification results; during the attack, by calculating the loss function and the gradient of the graph adjacency matrix A to select the most influential edges for perturbation; the adjacency matrix A describes the connectivity between nodes in the graph, and the attacker disrupts the graph structure by modifying the presence or absence of these edges; to efficiently perform the attack, for each element A in the adjacency matrix ij calculate the gradient of the attack loss function : According to the magnitude of the gradient, select the edge with the largest absolute value of the gradient for modification; specifically, if A ij and the gradient is positive, add the edge (i, j); if A ij = 1 and the gradient is negative, delete the edge (i, j); it can ensure that each modification will have a greater impact on the target loss function, thus effectively disturbing the classification result of the graphical model; the gradient of the adjacency matrix is calculated in each iteration, and important edges are selected for perturbation until the attack budget is reached; as the iteration progresses, the structure of the graph gradually changes, and finally a perturbed adversarial sample is generated and used for the evaluation of the classification task of the target model.

Citation Information

Cited By

  • Graph neural network recommendation system anti-fact interpretation method based on generative model

    CN122198163A