Visual analysis method for robustness of graph neural network
Through the robust visual analysis method of graph neural network, the robustness problem of graph neural network in the face of graph structure perturbation attacks is solved, and effective analysis and improvement means are provided to help experts improve the robustness and interpretability of the model.
Patent Information
- Application Number
- CN202510126889.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively solve the robustness of graph neural networks in the face of poisoning attacks and avoiding attacks, especially in the case of graph structure perturbation.
A robust visual analysis method for graph neural networks is proposed. Through steps such as data extraction, overall visualization, simulation of attacks, identification of perturbation nodes and focus node visual analysis, it helps experts understand model performance, diagnose problems during training and make adjustments.
This method provides comprehensive visual analysis methods to help experts comprehensively and effectively evaluate the robustness of graph neural networks in different attack scenarios, identify weak links of the model, and design more effective defense methods.
Smart Images

Figure CN120068927A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of robust analysis of graph neural networks, and specifically refers to a visual analysis method for the robustness of graph neural networks. Background Art
[0002] In the real world, many data are represented in the form of graphs, such as social networks, chemical molecules, and financial data. With the development of graph neural networks (GNNs), they have become increasingly popular due to their unique message passing mechanism. This mechanism updates node representations by aggregating neighbor information, enabling node representations to capture node features, neighbor information, and local graph structures. Based on this, graph neural networks have been widely applied to tasks such as node classification, graph classification, and link prediction. In addition, to improve model performance, researchers have proposed a variety of advanced GNN operations, such as graph convolutional networks (GCNs) and graph attention networks (GATs).
[0003] However, due to the nature of graph neural networks, they are more vulnerable to malicious attacks. According to the stage at which the attack occurs, these attacks are mainly divided into two categories: poisoning attacks and evasion attacks. Poisoning attacks occur during the training phase of the GNN model, where attackers modify the training graph to generate a model with impaired performance. Evasion attacks, on the other hand, target the inference phase, where attackers aim to interfere with the graph structure during inference, resulting in incorrect behavior of a well-trained GNN model.
[0004] Currently, most attacks are achieved by injecting perturbations into graph data. This perturbation method usually manifests as modifications to the graph structure, such as adding or deleting edges, or by changing node attributes. Empirical studies have shown that perturbing edges is more effective than perturbing node attributes, so graph structure perturbation has become the main attack method currently. Another common perturbation method is to inject malicious nodes into graph data, which interfere with classification by connecting to normal benign nodes. This method of injecting malicious nodes is often used when attackers cannot modify existing edges or nodes.
[0005] To defend against the above types of attacks, researchers have conducted a large amount of research and proposed a variety of defense methods. However, due to the highly discrete nature of the graph structure, the concept of perturbing graph samples is fundamentally different from that of ordinary samples (such as images) in Euclidean space, which makes the study of the robustness of GNNs full of challenges. At the same time, compared with the fields of images and texts, the research on the interpretability of graph models is relatively less, but this direction is crucial for understanding deep graph neural networks.
[0006] To develop more robust solutions, it is necessary to deeply understand the prediction process of adversarial examples and identify the root causes of incorrect predictions. By visually explaining the mechanism by which adversarial examples lead to misclassification and analyzing the reasons for errors, experts can identify the weaknesses of the model and thus design more effective attack or defense methods. Summary of the Invention
[0007] To solve the above problems, help experts understand the performance of the graph neural network model, diagnose problems in the training process, and guide experts to adjust and improve the graph neural network model, the present invention proposes a visual analysis method for the robustness of graph neural networks, including the following steps:
[0008] S1. Extraction of visual analysis data for existing graph neural networks: Extract data from the trained graph neural network model and its dataset. By collecting the training results of the graph neural network model on the dataset, obtain the basic data for subsequent analysis.
[0009] S2. Overall visualization: Use the basic data extracted in step S1 to generate visual charts for the data in the dataset and the model results in the overall view, helping users intuitively understand the overall structure of the dataset and the model classification results from an overall perspective.
[0010] S3. Simulated attack: Select an attack method to apply an attack to the dataset imported in step S1, simulating the situation where the dataset is maliciously attacked, for subsequent analysis and evaluation of the robustness of the graph neural network model imported in step S1.
[0011] S4. Extraction of visual analysis data for the graph neural network after being attacked: Use the dataset after being attacked as input, re-infer and judge through the graph neural network model, and extract the output results of the network model after being attacked to obtain the basic data for subsequent analysis.
[0012] S5. Identification of nodes vulnerable to perturbation: Based on the basic data extracted in step S4, identify the nodes vulnerable to attack perturbation and highlight these nodes.
[0013] S6. Visual analysis of focus nodes: Users can select the highlighted nodes vulnerable to perturbation in step S5 as focus nodes for further visual analysis.
[0014] S6.1. Based on the data extracted in step S4, generate the k-hop subgraph of the focus node and the visualization chart of the neighborhood features of the focus node.
[0015] S6.2. Through an interpretive method, extract the subgraph according to the focus node and obtain the influence scores of each node and edge in the subgraph on the focus node.
[0016] Preferably, the data to be extracted in step S1 includes: the true labels of the nodes in the training results, the predicted labels of the nodes, the prediction accuracy rates of the nodes, the topological structures of the nodes, the feature vectors of the nodes, etc., aiming to provide a basis for the visualization and performance analysis of the graph neural network.
[0017] Preferably, in step S2, the force-directed layout algorithm is used to display the topological structure of the data set. The force-directed layout algorithm will automatically adjust the positions of the nodes according to the connection relationships between the nodes in the data set, presenting the natural distribution and relationship density of the network, which helps users observe the aggregation areas of the nodes and the degree of their mutual connections.
[0018] Preferably, in step S2, the true labels and predicted labels of each node in the data set are encoded by colors, so as to intuitively display the label information of the nodes, which helps users quickly understand the overall structure of the data set and the classification results of the model.
[0019] Preferably, the optional attack methods for the data set in step S3 include edge deletion attack, edge addition attack and edge perturbation attack. It aims to simulate the impact of different types of graph structure change attacks on the model performance in the node classification task.
[0020] Preferably, the data to be extracted in step S4 includes: the predicted labels of the nodes after being attacked, the prediction accuracy rates of the nodes after being attacked, the topological structures of the nodes after being attacked, the feature vectors of the nodes after being attacked, etc., aiming to provide a basis for the performance change analysis of the graph neural network after being attacked.
[0021] Preferably, in step S5, the vulnerability degree of the nodes is evaluated by the classification margin (CM). The classification margin of any node v is expressed as:
[0022]
[0023] where G is the graph data, including the adjacency matrix and feature matrix of the graph data, f represents the graph neural network model as the classifier, and y * is the true label of node v in graph data G, is the predicted label of node v by the model. p(·) represents the probability distribution, represents the set of all classification labels excluding the true label. The smaller the classification margin (CM) value, the stronger the robustness of node v.
[0024] Preferably, in step S6.1, in order to evaluate the robustness of the graph neural network model at the overall level of the data set, the robustness score (RS) is defined:
[0025] RS τ (f) = AR τ>0 (f) - AR τ=0(f)
[0026] Among them, the attack risk (AR):
[0027]
[0028] Among them, τ represents the attack budget, that is, the intensity of the attack. G′ is the graph data after being attacked, and d(G′, G) represents the difference distance between the graph data before and after the attack.
[0029] Preferably, the objective of step S6.2 can be defined as finding the target subgraph by learning the mask matrix M to maximize the mutual information:
[0030]
[0031] For the graph data G = (A, X), where A is the adjacency matrix of the graph data and X is the feature matrix of the graph data, and H(·) is the conditional entropy.
[0032] Preferably, based on the adjacency subgraph generated in step S6.1, the nodes and edges in the subgraph G S will be highlighted and encoded according to the importance weights obtained from the mask matrix M. The width of the edges will be adjusted according to their weights, and the edges with larger weights will appear thicker.
[0033] Compared with the prior art, the present invention has the following advantages:
[0034] 1. The present invention provides comprehensive visualization analysis means, including aspects such as the topological structure of the graph neural network, node classification results, and model prediction effects, to help experts intuitively understand the behavior and performance of the model. Through interactive exploration and multi-level data display, the present invention enables experts to analyze the characteristics and problems of the network model more comprehensively and deeply.
[0035] 2. Compared with the common robustness evaluation methods in the prior art, the present invention provides more types of attack methods and more detailed evaluation means. Combining with the neural network interpretability method, it helps experts comprehensively and effectively evaluate the robustness of the graph neural network in the face of different attack scenarios and understand the basis for the model to make predictions.
[0036] 3. The present invention innovatively combines the multi-dimensional visualization analysis of the focus node and the local structure, supports users to perform visualization analysis on specific nodes, deeply understand the impact of attacks on the model performance, and conduct detailed analysis of the model through different dimensions. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is the flowchart of the present invention;
[0038] Figure 2This is a visualization diagram of the classification performance of the dataset and model of the present invention;
[0039] Figure 3 This is a visualization diagram of the labels and prediction distributions of the adjacent subgraphs of the focus nodes of the present invention;
[0040] Figure 4 This is a visual bar chart of the consistency of the focus nodes before and after being attacked and a visual feature matrix diagram of similar nodes of the present invention. Detailed implementation manners
[0041] The present invention will be described in detail below in conjunction with the accompanying drawings and specific examples.
[0042] A graph neural network robustness visual analysis method provided by the present invention, as Figure 1 shown, includes the following steps:
[0043] S1. Extraction of visual analysis data of the existing graph neural network.
[0044] In this step, first, data is extracted from the trained graph neural network model and its dataset. By collecting the training results of the model on the dataset, basic data for subsequent analysis is generated. The basic data is used to construct a visual analysis module for the graph neural network. The extracted basic data includes the true labels of the nodes, the predicted labels of the nodes, the prediction accuracy of the nodes, the topological structure of the nodes, the feature vectors of the nodes, etc. These data provide a basis for the subsequent visualization and performance analysis of the graph neural network.
[0045] The specific extraction steps are as follows:
[0046] 1. Import of the model and dataset: The user imports the trained graph neural network model and the corresponding dataset, ensuring that the input files meet the data format requirements of the graph neural network, such as network model files, node feature files, edge feature files, etc. The uploaded model and dataset will be loaded to the background server for preprocessing.
[0047] 2. Extraction of basic data: Convert the inference performance data of the model into visual basic data. By parsing information such as the accuracy rate and error rate of the output layer of the model, the classification results and prediction performances of the nodes are statistically analyzed, and the distribution of various node categories and the overall accuracy are calculated.
[0048] 3. Data storage and arrangement: After the data extraction is completed, the basic data is saved as a visual dataset. This dataset contains the feature matrices and connection relationship matrices of each node and edge, and records information such as the prediction accuracy rate, true label, and predicted label, so as to be efficiently used in the subsequent visualization steps.
[0049] S2. Overall visualization.
[0050] In this step, the dataset and model performance imported in the previous step will be initially visualized to help users establish a basic understanding of the dataset structure and model classification performance. Figure 2 As shown, this process utilizes the force-directed layout algorithm to display the topological structure of the data set, the correct classification of the nodes, and the predicted classification of the model, allowing users to intuitively view the topological distribution of each node in the network, the correct classification, and the initial classification performance of the model.
[0051] This visualization is divided into the following sub-steps:
[0052] 1. Force-directed layout algorithm display data set: The force-directed layout algorithm is used to visualize the node and edge structure of the graph neural network, making the overall topological structure of the network clearer. The force-directed layout algorithm automatically adjusts the node position according to the connection relationship between the nodes, so that nodes with dense connections are clustered together, while nodes with fewer edges or no connections are pushed away, thus presenting the natural distribution and relationship density of the network. This display helps users observe the clustering areas of nodes and the degree of their interconnection.
[0053] 2. Visualization of node classification results: In the overall topology of the data set, this method uses color and shape coding to display the correct classification label of each node. For each node, it will be color-coded according to its category, so that users can quickly identify the node clustering area of the same category. Through this color distribution method, users can observe the overall category distribution of the data set, the density of nodes in each category, and other information.
[0054] 3. Visualization of model prediction classification: Visualize the nodes based on the model's prediction results, compare the predicted labels of the nodes with the true labels, and use different colors to mark the nodes with correct and incorrect predictions. This method can intuitively show the classification accuracy of the model in the initial situation and help users identify areas where the model performs poorly.
[0055] S3. Simulate attack. In this step, the present invention provides a series of built-in attack algorithms for users to select and execute attacks on the imported graph neural network model. The main purpose of the simulated attack is to detect the performance of the model under malicious interference, especially focusing on the attack interference on the graph data structure in the node classification task. After selecting the attack method, the attack process will be executed according to the set parameters. By attacking the model, the user can evaluate its robustness in an interference environment.
[0056] This method supports the following attack methods:
[0057] 1. Edge deletion attack: This attack algorithm changes the neighbor relationships of nodes by deleting a certain number of edges from the graph structure, misleading the model in node classification. Edge deletion attacks usually target nodes with high connectivity, isolating them or weakening their association with specific category nodes. Users can set the proportion or number of edges to be deleted to control the intensity of the attack.
[0058] 2. Edge addition attack: Contrary to the edge deletion attack, the edge addition attack changes the connection relationships between nodes by adding new edges to the graph. This attack method may introduce false connections, causing the model to generate incorrect feature propagation paths, thus affecting the node classification effect. This method supports randomly adding edges or adding edges between specific nodes to test the model's sensitivity to noisy connections.
[0059] 3. Edge perturbation attack: The edge perturbation attack combines the operations of edge deletion and edge addition. By randomly adding or deleting edges, it simulates the uncertainty of the graph structure in practical applications. Users can set the perturbation proportion or specific nodes, and randomly add or delete edges during the attack execution to detect the model's classification performance in a mixed interference environment.
[0060] S4. Visual analysis data extraction of the graph neural network after being attacked.
[0061] In this step, the dataset after being attacked will be used as input and re-inferred through the graph neural network model. At this time, since the network model has been attacked specifically, it is necessary to extract the output results of the model in the attacked state to analyze the impact of the attack on the performance of the graph neural network and provide basic data for subsequent analysis.
[0062] The specific extraction steps are as follows:
[0063] 1. Inference of the attacked model: Input the dataset processed by the attack into the graph neural network model for re-inference. At this time, the model will make predictions based on the attacked data. Ensure that the input data format meets the requirements of the graph neural network and that the model can be loaded correctly.
[0064] 2. Basic data extraction: Convert the inference results of the model after being attacked into the basic data required for visual analysis. The data to be extracted includes: the predicted labels of the nodes after being attacked, the prediction accuracy of the nodes after being attacked, the topological structure of the nodes after being attacked, the feature vectors of the nodes after being attacked, etc. By extracting these data, it provides support for subsequent visualization and performance analysis to help analyze the impact of the attack on the model's performance.
[0065] S5. Identify nodes vulnerable to perturbation.
[0066] In this step, based on the basic data extracted in the previous step, this method can identify the nodes vulnerable to attack perturbations and highlight these nodes. Through the highlighting, these relatively vulnerable nodes can be presented more prominently, attracting the user's attention. In this way, the user can quickly identify which nodes show low robustness when facing attacks, and thus conduct further analysis and exploration.
[0067] The degree of vulnerability of a node is calculated by the classification margin (CM) of the node, that is, the probability difference of misclassification after the node is perturbed. The classification margin of any node v is expressed as:
[0068]
[0069] where G is the graph data, including the adjacency matrix and feature matrix of the graph data, f represents the graph neural network model as a classifier, and y * is the true label of node v in graph data G, is the predicted label of the model for node v. p(·) represents the probability distribution, represents the set of all classification labels excluding the true label. The smaller the classification margin (CM) value, the stronger the robustness of node v.
[0070] S6. Visual analysis of focus nodes.
[0071] In step S5, this method has identified and highlighted the nodes vulnerable to perturbations, and the user can select these nodes. The selected nodes will be used as focus nodes to display the basic information of these nodes, such as the original features, true labels, predicted labels, classification changes before and after attacks, etc. In addition, the focus nodes will be further visualized in multiple dimensions to help the user explore the impact of attacks on the model performance for nodes with different features, reveal the weak links of the model, and thus provide valuable guidance for subsequent optimization and protection measures.
[0072] S6.1. Visual analysis of the neighborhood of focus nodes
[0073] Based on the basic data extracted in step S4, multiple specific metrics for visualization in this step can be obtained.
[0074] To intuitively understand the performance of the model, the overall accuracy (acc) of the model will be calculated based on the true label of the node and the predicted label of the model:
[0075]
[0076] where y i is the true label of node i, is the predicted label of the node, I(·) is the indicator function, and N is the total number of nodes.
[0077] To quantitatively evaluate the robustness of the graph neural network model at the overall dataset level, a robustness score (RS) needs to be obtained. Combining the classification margin (CM) obtained in step S5, first define the model adversarial attack risk (AR):
[0078]
[0079] where τ represents the attack budget, that is, the intensity of the attack. G′ is the graph data after being attacked, and d(G′, G) represents the difference distance between the graph data before and after the attack. The adversarial attack risk (AR) measures the average possibility of misjudgment of the model under a certain attack budget τ. The greater the probability, the more likely the model is to make mistakes when facing attacks, and the worse the overall robustness.
[0080] The robustness score (RS) is then obtained from the adversarial attack risk (AR):
[0081] RS τ (f) = AR τ>0 (f) - AR τ=0 (f)
[0082] The robustness score (RS) measures the relative performance of the model before and after the attack, that is, the relative robustness of the model.
[0083] In the visual analysis of the focus node neighborhood, the main visual displays include the following:[[]]
[0084] 1. Focus node k - adjacent sub - graph display: In this step, the k - adjacent sub - graph of the focus node is dynamically generated according to the k value selected by the user, as Figure 3 shown. The sub - graph intuitively shows the relationship between the focus node and its neighbor nodes. Users can view the structural relationship between the focus node and its neighbor nodes, deeply analyze the changes in the local structure before and after the attack, and understand the impact of the attack on the model.
[0085] 2. Pie - chart analysis: The pie - chart shows data such as the node attribute characteristics and topological structure characteristics of the nodes that the user is interested in through multi - layer structure comparison, as Figure 3 shown in the upper - right corner. It helps users intuitively compare the label characteristics of the focus node and its neighborhood nodes, revealing the distribution changes of the focus node and its neighborhood nodes before and after the attack. The pie - chart shows the label changes of the focus node and its neighborhood in a three - layer structure: the central solid circle represents the true label of the focus node; the left and right halves of the middle arc layer represent the predicted labels before and after the attack respectively; the left and right halves of the outer arc layer show the label distributions of the neighborhood nodes before and after the attack.
[0086] 3. Radar chart analysis: The radar chart visualizes the multi-dimensional index data of several different categories of the true label, original predicted label, and post-attack predicted label of the focus node, revealing the changing trend of classification performance and helping users evaluate the changes in the performance of the model for different types of nodes before and after the attack. The multi-dimensions of the radar chart include the overall classification accuracy, the impact of the model classification output, and the model robustness score, as Figure 3 shown in the upper right corner.
[0087] 4. Feature matrix chart analysis: The feature matrix chart shows the features of the focus node and its top 5 neighbor nodes with the shortest feature distance in the form of 6×n, as Figure 4 shown below. Among them, 6 represents 1 focus node and its top 5 nodes with the shortest feature distance, and n represents the length of the feature vector of the node. By visually comparing the feature matrices, users can evaluate the similarity and difference between the focus node and the neighborhood nodes in the feature space and analyze the impact of the attack on feature perturbation.
[0088] 5. Bar chart analysis: This chart makes a comparison in the form of a bar chart to quantitatively analyze the consistency of the focus node with its neighborhood structure. The performance before the model attack and the classification results after the attack are distinguished by different colors to show the degree of interference of the attack on the neighborhood label consistency, as Figure 4 shown above. Through this comparison, users can clearly see the impact of the current attack method on the overall classification accuracy of the model and further understand the destructive effect of the attack.
[0089] S6.2. Explanation of model prediction results:
[0090] In this step, through interpretive methods, subgraphs closely related to the prediction of the focus node in the graph data will be extracted, and the influence score of the connections in the subgraph on the prediction of the focus node by the model will be obtained to analyze the internal mechanism when the model makes predictions.
[0091] Mutual information MI measures the information contribution degree of the subgraph to the model prediction result. The larger the mutual information value, the stronger the correlation between the information contained in the subgraph and the model prediction result. For the graph neural network f and the graph data G, given a target node v∈G, the goal is to find the subgraph that is most important for the prediction result Y to maximize the mutual information MI, which can be specifically expressed as:
[0092]
[0093] Since the network f is the given trained GNN network, H(Y) is a determined value, so the above goal can be transformed into finding the minimum conditional entropy:
[0094]
[0095] where H(·) is the conditional entropy:
[0096]
[0097] Since the graph data G = (A, X), where A is the adjacency matrix of the graph data and X is the feature matrix of the graph data. Define a mask matrix M for the adjacency matrix A, and the goal can ultimately be optimized to find the target subgraph by learning the mask matrix M to maximize the mutual information:
[0098]
[0099] Through the mask matrix M, the subgraph G that is most relevant to the focus node prediction can be directly extracted S , and the importance weights of each connection in the subgraph can be obtained. These weights reflect the contribution degree of the corresponding connection to the focus node prediction result. Higher weights indicate that they have a greater impact on the prediction result and may be the most relied-upon parts when the model makes a prediction.
[0100] Based on the adjacency subgraph generated in step S6.1, the nodes and edges in the subgraph G S will be highlighted and encoded according to the importance weights obtained from the mask matrix M. The width of the edge will be adjusted according to its weight, and edges with larger weights will appear thicker.
[0101] This visualization method intuitively shows the relative contributions of each connection in the graph to the target node prediction, enabling users to clearly see which parts play a key role in the model decision-making and understand the decision-making basis of the model. In addition, for the exploration of multiple focus nodes, it can also allow users to discover the commonalities among the influencing factors of different focus nodes, enabling users to discover potential laws in the model behavior and graph structure, further optimizing the model, and improving the robustness of the model.
Claims
1. A graph neural network robustness visual analysis method, characterized in that: The following steps are involved: S1. Obtain the trained graph neural network model and its data set, collect the training results of the graph neural network model on the data set, perform data extraction, and obtain basic data; S2. Generate visualization charts based on the data in the dataset and the results of the graph neural network model using basic data; S3, select an attack method to attack the data set, simulating the situation where the data set is subjected to malicious attacks; S4, taking the attacked data set as input, re-inferring and judging through the graph neural network model, and extracting data from the output results of the attacked network model; S5. Based on the data extracted in step S4, identify nodes that are vulnerable to attack disturbances and mark them; S6. Select the marked susceptible nodes as focal nodes for visual analysis.
2. According to claim 1, a graph neural network robustness visual analysis method is characterized in that: The data extracted in step S1 include: the true value label of the node in the training result, the predicted label of the node, the prediction accuracy of the node, the topological structure of the node and the feature vector of the node.
3. A graph neural network robustness visual analysis method according to claim 2, characterized in that: The specific steps of generating a visualization chart in step S2 are as follows: S2.
1. Display the topological structure of the data set through the force-directed layout algorithm. The force-directed layout algorithm automatically adjusts the node positions according to the connection relationship between the nodes in the data set, presenting the natural distribution and relationship density of the network; S2.
2. Encode the true value label and predicted label of each node in the dataset by color to display the label information of the node.
4. According to claim 3, a graph neural network robustness visual analysis method is characterized in that: The attack methods described in step S3 include edge deletion attack, edge addition attack and edge disturbance attack.
5. A graph neural network robustness visual analysis method according to claim 4, characterized in that: The susceptibility of the node in step S5 is evaluated by the classification interval CM. The classification interval of any node v is expressed as: Where G is the graph data, including the adjacency matrix and feature matrix of the graph data, f represents the graph neural network model as a classifier, and y * is the true label of node v in the graph data G, is the model's predicted label for node v, p(·) represents the probability distribution, Represents the set of all classification labels that do not contain the true label. The smaller the classification interval CM value, the stronger the robustness of the node v.
6. A graph neural network robustness visual analysis method according to claim 5, characterized in that: The specific implementation process of step S6 is as follows: S6.1, based on the data extracted in step S4, generating a focal node k-neighbor subgraph and a focal node neighborhood feature visualization chart; S6.
2. Through an explanatory method, a subgraph is extracted based on the focal node, and the influence score of each node and edge in the subgraph on the focal node is obtained.
7. A graph neural network robustness visual analysis method according to claim 6, characterized in that: In step S6.1, the robustness score RS is defined to evaluate the robustness score of the graph neural network model at the overall level of the dataset: RS τ (f)=AR τ>0 (f)-AR τ=0 (f) Where τ represents the attack budget, that is, the intensity of the attack, AR is the attack risk, G′ is the graph data after the attack, and d(G′, G) represents the difference distance between the graph data before and after the attack.
8. The method for visual analysis of graph neural network robustness according to claim 7, characterized in that: In step S6.1, the k-neighbor subgraph of the focal node and the focal node neighborhood feature visualization chart are generated. The specific process is as follows: Focus node k-adjacent subgraph: dynamically generate a focus node k-adjacent subgraph based on the k value, display the relationship between the focus node and its neighboring nodes, and analyze the changes in the local structure before and after the attack; Pie chart: Through multi-layer structure comparison, it shows the node attribute characteristics and topological structure characteristic data of the node. The pie chart shows the label changes of the focal node and its neighborhood in a three-layer structure: the central solid circle represents the true value label of the focal node; the left and right halves of the middle arc layer represent the predicted labels before and after the attack; the left and right halves of the outer arc layer show the label distribution of the neighborhood nodes before and after the attack; Radar chart: By visualizing the multi-dimensional indicator data of several different categories of the true value label, original prediction label and post-attack prediction label of the focus node, the changing trend of the classification performance is revealed. The multi-dimensionality of the radar chart includes the overall classification accuracy, the impact of the model classification output, and the model robustness score; Feature matrix diagram: The feature matrix diagram displays the features of the focal node and its top 5 neighboring nodes with the shortest feature distance in a 6×n format. By visually comparing the feature matrices, the similarities and differences between the focal node and the neighboring nodes in the feature space are evaluated; Bar chart: Quantitatively analyzes the consistency between the focal node and its neighborhood structure, distinguishes the performance of the model before the attack and the classification results after the attack, and shows the degree of interference of the attack on the consistency of the neighborhood labels.
9. A graph neural network robustness visual analysis method according to claim 8, characterized in that: Step S6.2 is specifically implemented as follows: extracting a subgraph from the graph data that is closely related to the prediction of the focal node through an explanatory method, and obtaining the influence score of the connection in the subgraph on the model's prediction of the focal node, and analyzing the internal mechanism of the model when making predictions: Mutual information MI measures the information contribution of a subgraph to the model prediction results. The larger the mutual information value, the stronger the correlation between the information contained in the subgraph and the model prediction results. For a graph neural network f and graph data G, given a target node v∈G, the goal is to find the subgraph that is most important to the prediction result Y. Make the mutual information MI maximum, specifically expressed as: Since the network f is a given trained GNN network, H(Y) is a fixed value, and the goal is to find the minimum conditional entropy: where H(·) is the conditional entropy: Since the graph data G = (A, X), where A is the adjacency matrix of the graph data, and X is the feature matrix of the graph data; define a mask matrix M for the adjacency matrix A, and the ultimate optimization goal is to find the target subgraph by learning the mask matrix M so that the mutual information is maximized: Through the mask matrix M, the subgraph G with the strongest correlation with the focus node prediction is directly extracted S , and obtain the importance weight of each connection in the subgraph; Based on the adjacency subgraph generated in step S6.1, subgraph G S The nodes and edges in will be highlighted and encoded according to their importance weights obtained in the mask matrix M, and the width of the edges will be adjusted according to their weights.
Citation Information
Cited By
Method and system for safety and robustness simulation analysis of industrial network topology structure
CN120358152A
A method and system for simulating security robustness of an industrial network topology
CN120358152B