A Defense Model for Graph Adversarial Attacks and Its Construction Method
By constructing message generation components based on node characteristics, message aggregation components of attention mechanism and multi-aggregator, as well as message update components of gate mechanism and adaptive residual mechanism, the problem of insufficient robustness in the face of adversarial samples is solved, and the effective and low-time consumption graph-anti-aggregation attack defense effect is achieved.
Patent Information
- Application Number
- CN202310468686.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-04-27
AI Technical Summary
Existing graph neural network models are susceptible to perturbation when facing adversarial samples, resulting in insufficient robustness, especially in security or financial applications.
A defense model for graph-oriented attacks is proposed. By constructing a message generation component based on node characteristics, a message aggregation component based on attention mechanism and multi-aggregator, and a message update component based on gating mechanism and adaptive residual mechanism, we can reduce the impact of perturbation edges and improve the robustness of the model.
This model significantly improves robustness when facing adversarial samples, can effectively reduce time consumption, and is adapted to other GNN models that meet the messaging mechanism, providing a defense model construction method that can realize the robustness of industrial systems related to graph neural networks without understanding the attack details.
Smart Images

Figure CN116502087B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graph data, and particularly relates to a defense model for graph adversarial attacks and a construction method thereof. Background Art
[0002] Graph data is used to model a set of objects (nodes) and their relationships (edges). Graph data is very important in many research fields. For example, in the e-commerce field, graph-based learning systems can effectively utilize the interactions between users and products to make more accurate recommendations; in the chemical field, molecules can be expressed as graphs to discover new drugs by identifying their biological activities; in the biological field, new complexes can be generated through protein-protein interaction networks. Obviously, analyzing graph data has become a very important direction.
[0003] Graph Neural Networks (GNNs) are a type of deep learning model that satisfies the message passing mechanism. It aggregates the features of neighboring nodes in a recursive manner, can process inputs based on graph data, and then analyze the graph data. In the past few years, due to the good performance of graph neural networks, GNNs have been widely used in various graph data analysis tasks. However, GNNs are also vulnerable to adversarial samples, just like traditional deep learning models. By artificially designed elaborate perturbations, GNNs can be induced to make incorrect predictions for certain specific nodes with high confidence. [2] These adversarial attacks are usually concealed and only require small perturbations of the graph structure and / or node features (e.g., by adding or deleting edges). For example, Nettack constrains the perturbations by keeping the node degree distribution unchanged and co-occurrence. Such imperceptible adversarial perturbations pose a serious security threat to deep learning model application systems based on GNNs, especially in security or financial applications, where the risks are extremely high. For example, in the e-commerce field, injecting non-existent links between users and products can degrade the recommendation effect; in the social network and financial system fields, an adversary may disguise their illegal behavior by conducting multiple transactions with trustworthy agents.
[0004] To improve the robustness of the GNN model against adversarial samples, many researchers have made some efforts. GNNJacard calculates the feature similarity between node pairs through the Jaccard distance, and preprocesses the perturbed graph by deleting nodes with relatively low feature similarity. GNNSVD discovered the Nettack [3] attack as a high-rank attack and proposed a low-rank defense mechanism based on this discovery: using a low-rank approximation of the adjacency matrix to remove noise information through SVD decomposition. RGCN [7]The model uses a Gaussian distribution as the hidden layer to absorb the impact of attacks in the variance. ProGCN shows that the clean graph is low-rank, sparse, and the eigenvalues of interconnected nodes are more similar by analyzing the different attributes of the perturbed graph and the clean graph. This guides the design of a robust graph neural network, and at the same time learns the clean graph structure from the parameters of the perturbed graph GNN. GNNGUARD improves the robustness of GCN when facing perturbed graphs by assigning higher weights to nodes with similar features, smaller weights to suspicious neighbors, and pruning edges that may be perturbed. Moreover, it uses Layer-Wise GraphMemory to retain part of the memory of the graph structure pruned from the previous layer to maintain the stability of training. SimPGCN proposes that GCN will destroy the node similarity in the original feature space during the aggregation process, and the original feature information of nodes can mitigate the impact of adversarial attacks. Based on this, an aggregation strategy for preserving node feature similarity is designed, which can adaptively fuse the graph structure and the original node features. MedianGNN explains the reason for the robustness of GCN (Graph Convolutional Neural Network) by the decomposition point theory, which depends on the robustness of the aggregation function. It is proposed that the decomposition point of the median and pruning aggregator is higher than that of the traditional sum aggregator, so it is more robust. ElasticGNN proposes graph smoothing based on l1 regularization, which is fused with l2 regularization as a new message passing mechanism. L1 regularization can better maintain local smoothness and can resist the attack of adversarial samples to a certain extent.
[0005] The above defense schemes can be divided into preprocessing strategies (GNNJacard, GNNSVD), graph structure learning strategies (ProGCN, GNNGuard), and robust optimization strategies (RGCN, SimpGCN ElasticGNN). The schemes they designed are all considered too complex.
[0006] For example, the main problem of ProGCN in graph structure learning is that it also needs to learn the graph structure during the training process, perform operations such as edge deletion on the perturbed graph, which consumes a lot of time and the learned structure may not be similar to the original graph, resulting in secondary pollution of the graph structure. The problem of GNNGuard is that the data for graph structure learning needs to be transferred between the CPU and the GPU, with high communication costs. In addition, there is also the problem of secondary pollution of the graph structure in ProGNN.
[0007] SimPGCN in the robust optimization strategy considers the KNN feature map (KNN means selecting the K most similar neighbor nodes under a certain metric criterion to form a graph), self-supervised learning and other modules, etc. Then, the entire message passing layer is coupled together, and the objective function is also modified due to self-supervised learning, resulting in low adaptability; RGCN considers representing the entire message passing layer with a Gaussian distribution, but in order to force the learned representation to conform to the Gaussian distribution characteristics, the objective function is also modified, making the adaptability low. The problem of MedianGCN is to modify the aggregation function in the message passing. Although it has adaptability, the problem is that it learns the representation from the feature dimension, resulting in a long training time. The new message passing mechanism proposed by ElasticGNN cannot be adapted to traditional GNN models. Preprocessing operations such as GCNSVD and GCNJaccard solve the influence of the message passing stage before training, and process the graph structure by edge deletion and low-rank approximation. Although it is simple, there is also the problem that the graph structure after purification is not similar to the original graph, which easily leads to suboptimal performance. Summary of the Invention
[0008] The purpose of the present invention is: aiming at the above existing problems, the present invention provides a defense model for graph adversarial attacks and its construction method, which meets adaptability, low time consumption simplicity and high accuracy. The present invention provides a message generation component based on node features, a message aggregation component based on attention mechanism and multi-aggregator, and a message update component based on gating mechanism and adaptive residual mechanism. These components all transform the features of classification nodes through simple learnable parameters and linear functions, with simplicity and low time consumption; at the same time, these components can be adapted to other GNN models that meet the message passing mechanism, improving the robustness of these models against adversarial sample attacks.
[0009] The technical solution adopted by the present invention is as follows:
[0010] A method for constructing a defense model for graph adversarial attacks, using G=(V, E, X) to represent a graph, where V={v1, v2, ……, v n} is the set of nodes, represents the set of edges, X={x1, x2, ……, x n} represents the set of node features, x i ∈R F is an F-dimensional vector, representing the node feature of node v I ∈V; using A∈R N×N to represent an adjacency matrix, where the element A ij ∈{0, 1} represents whether there is an edge e i between node v j and node v i,j exists; using to represent the node v i 's neighbor set, including the following steps: constructing a message generation component, a message aggregation component, and a message update component, where
[0011] the message generation component performs weighted processing on the input graph data to reduce the impact of perturbed edges, generates and outputs the message msg ij ;
[0012] the message aggregation component aggregates the message msg ij through different aggregators, then splices the aggregation results of different aggregators, and uses the attention mechanism to assign different weights to different aggregators to further improve the model's defense performance, generates and outputs the aggregation result c i ;
[0013] the message update component includes a gating sub-component, a loop judgment sub-component, and an adaptive residual mechanism sub-component. The gating sub-component uses the feature filter f of the gating mechanism to process the aggregation result c i to obtain the node feature u after feature filtering i ;
[0014] the loop judgment sub-component judges whether the node feature u after feature filtering i needs to return to the message generation component for further processing based on the current layer number k of the node feature u after feature filtering i ;
[0015] the adaptive residual mechanism sub-component is used to measure the importance of the original node feature and the node feature u after feature filtering of the last layer, and perform importance assignment i .
[0016] The construction steps of the message generation component are as follows:
[0017] a) Input the graph data, and perform K-layer processing on the graph data in total. If the processing layer k = 1, the system receives the graph data, reduces the dimension of all node features X, and obtains H = {h1, h2,... h i …, h n}, h i ∈R (F′) , where i represents the node v i , and F' is the node feature dimension of the current layer; if the processing layer k ∈ {2,.. K - 1}, then receive the data of the message update component of the k - 1 layer, so that h i = u i ;
[0018] b) Calculate the message two-norm product of the central node v i and the neighbor node v j feature through equation (1):
[0019] norm(h i ,h j ) = ||h i ||2·||h j ||2 (1)
[0020] In equation (1), h i is the eigenvector of the central node v i , h j is the eigenvector of the neighbor node v j , j ∈ N(v i ) ∪ {v i} represents all neighbors of the central node, norm(h i ,h j ) represents calculating the product of the two-norms of h i and h j ;
[0021] c) Based on the learnable parameter α, score the message two-norm to obtain the weight w m of the message two-norm, as shown in equation (2):
[0022] w m = α·norm(h i ,h j ) (2)
[0023] In equation (2), α is a learnable parameter, α ∈ R (1) ;
[0024] d) Use the softmax function to normalize the weight w m of the message two-norm to obtain the message weight w j transmitted by the neighbor v i to the central node v j , as shown in equation (3):
[0025]
[0026] In equation (3), exp is the exponential function;
[0027] e) Use the message weight w i of the central node v j to assign different scores to different neighbor messages to obtain the message msg j transmitted by the neighbor v ij to the central node, as shown in equation (4):
[0028] msg ij = w j ·h j (4)
[0029] f) Output the message msg ij to the message aggregation component.
[0030] The construction steps of the message aggregation component are as follows:
[0031] a) Receive the message msg sent by the message generation component ij ;
[0032] b) Use the aggregator AGG to aggregate the message msg ij as shown in Equation (5):
[0033] agg im = AGG im (msg ij ) (5),
[0034] In Equation (5), j ∈ N(v i ) ∪ {v i}, agg im ∈ R (F′) represents the aggregation result of the m-th aggregator in the aggregator combination for node v i , F′ is the node feature dimension of the current layer, m ∈ {1,..M}, and M represents the number of aggregators;
[0035] c) Concatenate the aggregation results of different aggregators, as shown in Equation (6):
[0036] Q i = CAT(agg i1 ,…agg im ,…agg iM ) (6)
[0037] In Equation (6), Q i ∈ R (M,F′) , CAT represents the concatenation operation along dimension 0, m ∈ {1,..M}, and M represents the number of aggregators;
[0038] d) Learn the importance of different aggregators through the attention mechanism. Adopt the dot-product attention mechanism to avoid introducing additional parameters, as shown in Equation (7):
[0039]
[0040] In Equation (7), where, Q i = K i = V I ∈ R (M,F′) , represent the query value, key value, and value respectively, F′ represents the dimension of the node feature at this layer; s I ∈ R (M) represents the importance of each of the M aggregators;
[0041] e) Process the feature results transmitted by the attention mechanism using Equation (8):
[0042] c i = SUM(V i * s i ) (8)
[0043] In Equation (8), * represents the multiplication of two elements at corresponding positions of two matrices, V i * s i ∈ R (M,F′) represents weighting the aggregator; SUM() represents summing along dimension 0 to sum the results of the weighted aggregator, and the central node obtains the aggregated result c i ∈ R (F′) , where F' is the node feature dimension of the current layer;
[0044] f) Output the aggregated result c i to the message update component.
[0045] The types of the AGG im include a sum aggregator, a mean aggregator, a maximum aggregator, and a minimum aggregator. The aggregator combination can be selected according to different data sets.
[0046] Furthermore, the construction steps of the gating sub-component are as follows:
[0047] a) Receive the aggregated result c transmitted by the message aggregation component i ;
[0048] b) Implement the feature filter f using the sigmoid function, as shown in Equation (9):
[0049] f = sigmoid(w b @ add(c i , h i )) (9)
[0050] In Equation (9), w b ∈ R (F′) is a learnable parameter, F' is the node feature dimension of the current layer; add is an element-wise addition of vectors;
[0051] c) Filter the aggregated result c i using the feature filter f, as shown in Equation (10),
[0052] u i = f * c i (10)
[0053] In Equation (10), f is the feature filter, ci is the aggregation result processed by the node message aggregation component, where * represents the multiplication of two elements at the corresponding positions of two matrices, and u I ∈R (F′) represents the node features after feature filtering, and F′ is the node feature dimension of the current layer;
[0054] d) Output the node features u after feature filtering I to the loop judgment component.
[0055] Furthermore, the loop judgment sub-component receives the node features u after feature filtering transmitted by the gating sub-component I , and determines whether the current layer number k is the last layer. If so, it outputs the node features u after feature filtering I to the adaptive residual mechanism sub-component. If it is not the last layer, the current layer number k is incremented by 1, and the message generation component is returned for continued processing.
[0056] Furthermore, the adaptive residual mechanism sub-component uses the adaptive residual mechanism to perform importance assignment on the original node features and the node features u after feature filtering of the last layer I as shown in Equation (11):
[0057]
[0058] In Equation (11), a ∈ R (1) is a learnable parameter, is the initial feature of node v I , is the feature of the node after being processed by the message update component in the (K - 1)th layer.
[0059] Furthermore, the output result of the adaptive residual mechanism sub-component is sent to the classification layer for classification.
[0060] A graph adversarial attack defense model constructed according to the above defense model construction method for graph adversarial attacks.
[0061] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are as follows:
[0062] 1. The present invention designs robust defense components in the message generation, aggregation, and update stages to mitigate the impact of edge perturbations. In addition, in the update stage of the last layer, a residual mechanism adaptive function is designed to improve the expressive ability of the model.
[0063] 2. Each component of the present invention can be inserted into any GNN model that satisfies the message passing mechanism, thereby improving the robustness of the model against adversarial samples.
[0064] 3. Compared with other defense models, the components provided by the present invention transform the features of classification nodes through simple learnable parameters and linear functions, with shorter training time. Inserting the corresponding defense components into other traditional GNN models also consumes less time.
[0065] 4. The present invention provides a method for constructing a defense model for the robustness of graph neural network-related industrial systems without the need to understand the details of attacks against existing graph adversarial samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is the overall flowchart of the defense model of the present invention;
[0067] Figure 2 is the construction flowchart of the message generation component of the present invention;
[0068] Figure 3 is the construction flowchart of the message aggregation component of the present invention;
[0069] Figure 4 is the construction flowchart of the message update component of the present invention;
[0070] Figure 5 is the comparison graph of time consumption test of the present invention under the cora dataset;
[0071] Figure 6 is the comparison graph of time consumption test of the present invention under the cora_ml dataset;
[0072] Figure 7 is the comparison graph of time consumption test of the present invention under the pubmed dataset;
[0073] Figure 8 is the test graph of the adaptability of the message generation component of the present invention;
[0074] Figure 9 is the test graph of the adaptability of the message aggregation component of the present invention;
[0075] Figure 10 is the test graph of the adaptability of the message update component of the present invention;
[0076] Figure 11 is the node dimensionality reduction graph of the cora_ml dataset of the present invention under the Nettack attack;
[0077] Figure 12 is the node dimensionality reduction graph of the cora_ml dataset of GAT under the Nettack attack;
[0078] Figure 13It is the node dimensionality reduction diagram of the cora_ml dataset under the SGAttack attack in the present invention;
[0079] Figure 14 It is the node dimensionality reduction diagram of the cora_ml dataset under the SGAttack attack in GAT. Specific implementation manner
[0080] The present invention will be described in detail below with reference to the accompanying drawings.
[0081] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0082] Embodiment 1
[0083] A method for constructing a defense model for graph adversarial attacks, as Figure 1-4 shown, uses G=(V, E, X) to represent a graph, where V={v1, v2, ……, v n} is the set of nodes, represents the set of edges, and X={x1, x2, ……, x n} represents the set of node features. x i ∈R F is an F-dimensional vector representing the node feature of node v i ∈V; uses A∈R N×N to represent an adjacency matrix, where the element A ij ∈{0, 1} represents whether there is an edge e i between node v j and node v i,j ; uses to represent the neighbor set of node v i , and includes the following steps: constructing a message generation component, a message aggregation component, and a message update component, where,
[0084] The message generation component performs weighted processing on the input graph data to reduce the impact of perturbed edges, and generates and outputs a message msg ij ;
[0085] The message aggregation component aggregates the message msg ij through different aggregators, then splices the aggregation results of different aggregators, and uses the attention mechanism to assign different weights to different aggregators to further improve the model defense performance, and generates and outputs an aggregation result c i ;
[0086] The message update component includes a gating sub-component, a loop judgment sub-component, and an adaptive residual mechanism sub-component. The gating sub-component uses the feature filter f of the gating mechanism to process the aggregation result c i to obtain the node feature u after feature filtering i ;
[0087] The loop judgment sub-component determines whether the node feature u after feature filtering i needs to be returned to the message generation component for further processing based on the current layer number k of the node feature u after feature filtering i ;
[0088] The adaptive residual mechanism sub-component is used to measure the importance of the original node feature and the node feature u after feature filtering in the last layer, and perform importance assignment i .
[0089] The construction steps of the message generation component are as follows
[0090] a) Input graph data, and perform K-layer processing on the graph data in total. If the processing layer k = 1, the system receives the graph data, reduces the dimension of all node features X, and obtains H = {h1, h2,... h i …, h n}, h i ∈R (F′) , where i represents the node v i , and F′ is the node feature dimension of the current layer; if the processing layer k ∈ {2,..K - 1}, then accept the data of the message update component of the k - 1 layer, so that h i = u i ;
[0091] b) Calculate the message two-norm product of the features of the central node v i and the neighbor node v j through formula (1):
[0092] norm(h i , h j ) = ||h i ||2 · ||h j ||2 (1)
[0093] In formula (1), h i is the feature vector of the central node v i , h j is the feature vector of the neighbor node v j , j ∈ N(v i ) ∪ {v i} represents all neighbors of the central node, and norm(h i , h j ) represents calculating h i and hj Product of the second norm;
[0094] c) Score the second norm of the message based on the learnable parameter α to obtain the weight w of the second norm of the message m , as shown in Equation (2):
[0095] w m = α · norm(h i , h j ) (2)
[0096] In Equation (2), α is a learnable parameter, α ∈ R (1) ;
[0097] d) Use the softmax function to normalize the weight w of the second norm of the message m to obtain the message weight w j passed from neighbor v i to central node v j , as shown in Equation (3):
[0098]
[0099] In Equation (3), exp is the exponential function;
[0100] e) Use the message weight w i of central node v j to assign different scores to different neighbor messages to obtain the message msg j passed from neighbor v ij to central node, as shown in Equation (4):
[0101] msg ij = w j · h j (4)
[0102] f) Output the message msg ij to the message aggregation component.
[0103] The construction steps of the message aggregation component are as follows:
[0104] f) Receive the message msg passed from the message generation component ij ;
[0105] g) Use the aggregator AGG to aggregate the message msg ij , as shown in Equation (5):
[0106] agg im = AGG im (msg ij ) (5),
[0107] In formula (5), j ∈ N(v i ) ∪ {v i}, agg im ∈ R (F′) represents the aggregation result of the m-th aggregator in the aggregator combination for node v i . F' is the node feature dimension of the current layer, m ∈ {1,..M}, and M represents the number of aggregators;
[0108] h) Concatenate the aggregation results of different aggregators, as shown in formula (6):
[0109] Q i = CAT(agg i1 ,…agg im ,…agg iM ) (6)
[0110] In formula (6), Q i ∈ R (M,F′) , CAT represents the concatenation operation along dimension 0, m ∈ {1,..M}, and M represents the number of aggregators;
[0111] i) Learn the importance of different aggregators through the attention mechanism. The dot product attention mechanism is adopted to avoid introducing additional parameters, as shown in formula (7):
[0112]
[0113] In formula (7), where Q i = K i = V i ∈ R (M,F′) , represent the query value, key value, and value respectively, and F' represents the dimension of the node features at this layer; s i ∈ R (M) represents the importance of each of the M aggregators;
[0114] j) Process the feature results passed from the attention mechanism using formula (8):
[0115] c i = SUM(V i * s i ) (8)
[0116] In formula (8), * represents the multiplication of two elements at the corresponding positions of the two matrices. V i * s i ∈ R (M,F′) represents the weighting of the aggregators; SUM() represents the summation along dimension 0, summing the results of the weighted aggregators, and the central node obtains the aggregation result c i ∈ R (F′), F′ is the node feature dimension of the current layer;
[0117] f) Output the aggregation result c i to the message update component.
[0118] The AGG im includes sum aggregator, mean aggregator, max aggregator and min aggregator.
[0119] The construction steps of the gating sub-component are as follows:
[0120] a) Receive the aggregation result c passed from the message aggregation component i ;
[0121] b) Implement the feature filter f using the sigmoid function, as shown in Equation (9):
[0122] f = sigmoid(w b @add(c i , h i )) (9)
[0123] In Equation (9), w b ∈R (F′) is a learnable parameter, F′ is the node feature dimension of the current layer; add is an element-wise addition of vectors;
[0124] c) Filter the aggregation result c using the feature filter f i as shown in Equation (10),
[0125] u i = fc i (10)
[0126] In Equation (10), f is the feature filter, c i is the aggregation result obtained by processing the node message aggregation component, * represents the multiplication of two elements at the corresponding positions of two matrices, u i ∈R (F′) represents the node feature after feature filtering, and F′ is the node feature dimension of the current layer;
[0127] d) Output the node feature u after feature filtering i to the loop judgment component.
[0128] The loop judgment sub-component receives the node feature u after feature filtering passed from the gating sub-component i , and judges whether the current layer number k is the last layer. If so, output the node feature u after feature filtering i to the adaptive residual mechanism sub-component. If it is not the last layer, then increment the current layer number k by 1 and return to the message generation component for further processing.
[0129] The adaptive residual mechanism sub-component uses the adaptive residual mechanism to assign importance to the original node features and the node features u after the last layer feature filtering i as shown in Equation (11):
[0130] In Equation (11), a ∈ R (1) is a learnable parameter is the initial feature of node v i and is the feature of the node after being processed by the message update component in the (K - 1)-th layer
[0131] The output result of the adaptive residual mechanism sub-component is fed into the classification layer for classification
[0132] Example 2
[0133] A graph adversarial attack defense model constructed according to the method for constructing a defense model for graph adversarial attack described in Example 1. The obtained graph adversarial attack defense model is tested. As Figure 5-7 shown, the horizontal axis represents each attack; the vertical axis represents the time consumption; different bars represent different defense techniques; the name of the model in this example during testing is MODENET. Since the time consumption of ProGNN and GNNGuard is very large, they are not drawn in the test results; in the pubmed dataset, since the time consumption of MedianGNN is also very large, it is not drawn. It can be seen from the figure that the graph adversarial attack defense model provided in this example has the shortest time consumption compared to several other defense models in the cora_ml, cora, and pubmed datasets
[0134] Example 3
[0135] A graph adversarial attack defense model constructed according to the method for constructing a defense model for graph adversarial attack described in Example 1. The graph adversarial attack defense model provided in this example is adapted to other GNN models. In this example, GCN and GAT are selected. As Figure 8-10 shown, the line with asterisks is the adapted model; the line with plus signs is the original model. Among them Figure 8 is to replace the original message generation part of a certain model (GCN, GAT) with the message generation component of this example Figure 9 represents replacing the original aggregation part of a certain model (GCN, GAT) with the message aggregation component of this example Figure 10It means replacing the original update part of a certain model (GCN, GAT) with the update component of the present invention. It can be seen that after adapting the graph adversarial attack defense model of this embodiment to other GNN models, the defense performance of the original model is effectively improved.
[0136] Example 4
[0137] A graph adversarial attack defense model constructed according to the method for constructing a defense model against graph adversarial attacks described in Example 1. Under the Nettack attack, the nodes of the cora_ml dataset are respectively dimension-reduced by using the graph adversarial attack defense model provided in this embodiment and the GAT model, as Figure 11 shown, is the graph of the dimension-reduced nodes of the cora_ml dataset of the graph adversarial attack defense model provided in this embodiment under the Nettack attack, as Figure 12 shown, is the graph of the dimension-reduced nodes of the cora_ml dataset of GAT under the Nettack attack. It can be seen that the graph adversarial attack defense model provided in this embodiment has stronger discrimination ability and the obtained graph is clearer.
[0138] Under the SGAttack attack, the nodes of the cora_ml dataset are respectively dimension-reduced by using the graph adversarial attack defense model provided in this embodiment and the GAT model, as Figure 13 shown, is the graph of the dimension-reduced nodes of the cora_ml dataset of the graph adversarial attack defense model provided in this embodiment under the SGAttack attack, as Figure 14 shown, is the graph of the dimension-reduced nodes of the cora_ml dataset of GAT under the SGAttack attack. Similarly, it can be seen that the graph adversarial attack defense model provided in this embodiment has stronger discrimination ability and the obtained graph is clearer.
[0139] Example 5
[0140] A graph adversarial attack defense model constructed according to the method for constructing a defense model against graph adversarial attacks described in Example 1. Under the Nettack attack, a comparison is made between the graph adversarial attack defense model provided in this embodiment and the defense models of other solutions, and the results are shown in Table 1 below:
[0141] Table 1 Comparison of the defense performance of the present invention and other solutions under the Nettack attack
[0142]
[0143] Under the SGAttack attack, a comparison is made between the graph adversarial attack defense model provided in this embodiment and the defense models of other solutions, and the results are shown in Table 2 below:
[0144] Table 2 Comparison of the defense performance of the present invention and other solutions under the SGAttack attack
[0145]
[0146]
[0147] From the data in the above table, it can be seen that the defense performance of the graph adversarial attack defense model provided in this embodiment is better than that of the other solutions.
[0148] In this article, specific embodiments are used to elaborate on the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
[0149] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of this invention is usually placed when in use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0150] In the description of the present invention, it should also be noted that unless otherwise clearly specified and defined, the terms "set", "install", "connect", "couple" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
Claims
1. A method for constructing a defense model against graph adversarial attacks, using G=(V, E, X) to represent a graph, where V={v1, v2, ……, v n} is the set of nodes, represents the set of edges, X={x1, x2, ……, x n} represents the set of node features, and x i ∈ R F is an F-dimensional vector representing the node feature of node v i ∈ V; Use \(A\in\mathbb{R}\) N×N to represent an adjacency matrix, where the element \(A\) ij \(\in\{0, 1\}\) represents whether there is an edge \(e\) i between node \(v\) j and node \(v\) i,j ; Use to represent the set of neighbors of node \(v\) i , which is characterized by including the following steps: constructing a message generation component, a message aggregation component, and a message update component, where The message generation component performs weighted processing on the input graph data to reduce the impact of the perturbed edges and generates and outputs the message msg ij ; The message aggregation component aggregates the message msg through different aggregators ij Then, it splices the aggregation results of different aggregators, assigns different weights to different aggregators using the attention mechanism, further improves the model's defense performance, and generates and outputs the aggregation result c i ; The message update component includes a gating sub-component, a loop judgment sub-component, and an adaptive residual mechanism sub-component. The gating sub-component uses the feature filter f of the gating mechanism to process the aggregation result c i to obtain the node feature u after feature filtering i ; The loop judgment sub-component determines the node feature u after feature filtering based on the current layer number k of i i whether the node feature u after the feature filtering needs to be returned to the message generation component for further processing; i The adaptive residual mechanism sub-component is used to measure the importance of the original node features and the node features u after the last layer of feature filtering, and perform importance assignment. i of each, and perform importance assignment.
2. The method for constructing a defense model against graph adversarial attacks according to claim 1, wherein The construction steps of the message generation component are as follows: a) Input graph data and perform K layers of processing on the graph data. If the processing layer k = 1, the system receives the graph data and reduces the dimensionality of all node features X to obtain H = {h1, h2, … h i …, h n}, h i ∈R (F′) , where i represents node v i , and F′ is the node feature dimension of the current layer; if the processing layer k ∈ {2,..K - 1}, then receive the data of the message update component of the (k - 1)-th layer, such that h i = u i ; b) Calculate the central node v through formula (1) i and the neighbor node v j The product of the message two-norms of the features norm(h i ,h j )=||h i ||2·||h j ||2 (1) In Equation (1), h i is the eigenvector of the central node v i , h j is the eigenvector of the neighbor node v j , j ∈ N(v i ) ∪ {v i} represents all neighbors of the central node, norm(h i , h j ) represents calculating the product of the second norms of h i and h j ; c) Score the message two - norm based on the learnable parameter α to obtain the weight w of the message two - norm m , as shown in Equation (2): w m = α · norm(h i , h j ) (2) In formula (2), α is a learnable parameter, α ∈ R (1) ; d) Use the softmax function to normalize the weight w of the message two-norm m to obtain the neighbor v j and pass the message weight w i to the central node v j , as shown in Equation (3): In Equation (3), exp is the exponential function; e) Using the message weight w i of the central node v j , different scores are assigned to different neighbor messages to obtain the message msg j sent by neighbor v to the central node, as shown in Equation (4): ij msg ij = w j ·h j (4) f) Output message msg ij To the message aggregation component.
3. The method for constructing a defense model against graph adversarial attacks according to claim 1, characterized in that, The construction steps of the message aggregation component are as follows: a) Receive the message msg passed from the message generation component ij ; b) Aggregate the message msg using the aggregator AGG ij as shown in Equation (5): agg im = AGG im (msg ij ) (5), In formula (5), j ∈ N(v i ) ∪ {v i}, agg im ∈ R (F′) represents the aggregation result of the m-th aggregator in the aggregator combination for node v i , F ′ is the node feature dimension of the current layer, m ∈ {1,..M}, and M represents the number of aggregators; c) Concatenate the aggregation results of different aggregators, as shown in Equation (6): Q i = CAT(agg i1 ,…agg im ,…agg iM ) (6) In Equation (6), Q i ∈R (M,F′) , CAT represents the concatenation operation along dimension 0, m ∈ {1,..M}, and M represents the number of aggregators; d) Learn the importance of different aggregators through the attention mechanism. The dot product attention mechanism is adopted to avoid introducing additional parameters, as shown in Equation (7): In formula (7), where Q i = K i = V i ∈ R (M,F′) , respectively represent the query value, key value, and value, and F ′ represents the dimension of the node features at this layer; s i ∈ R (M) represents the importance of each of the M aggregators; e) Process the feature results transmitted by the attention mechanism using Equation (8): c i = SUM(V i * s i ) (8) In Equation (8), * represents the multiplication of two elements at the corresponding positions of two matrices, and V i *s i ∈R (M,F′) represents weighting the aggregator; SUM() represents summing along dimension 0 to sum the results of the weighted aggregator, and the central node obtains the aggregated result c i ∈R (F′) , where F′ is the node feature dimension of the current layer; f) Output the aggregated result c i To the message update component.
4. The method for constructing a defense model against graph adversarial attacks according to claim 3, wherein The AGG im includes a sum aggregator, a mean aggregator, a maximum aggregator, and a minimum aggregator.
5. The method for constructing a defense model against graph adversarial attacks according to claim 1, characterized in that, The construction steps of the gating sub-component are as follows: a) Receive the aggregated result c passed by the message aggregation component i ; b) Implement the feature filter f using the sigmoid function, as shown in Equation (9): f = sigmoid(w b @add(c i ,h i )) (9) In formula (9), w b ∈R (F′) is a learnable parameter, and F ′ is the node feature dimension of the current layer; add is an element-wise addition of vectors; c) Filter the aggregation result c using the feature filter f i as shown in Equation (10), where u i = f * c i (10) In formula (10), f is the feature filter, and c i is the aggregation result obtained by processing the node message aggregation component. * represents the multiplication of two elements at the corresponding positions of two matrices, and u i ∈R (F′) represents the node features after feature filtering, and F ′ is the node feature dimension of the current layer; d) Output the node feature u after feature filtering i To the loop judgment component.
6. The method for constructing a defense model against graph adversarial attacks according to claim 1, characterized in that The loop judgment sub-component receives the node feature u after feature filtering transmitted by the gating sub-component i , and determines whether the current layer number k is the last layer. If so, it outputs the node feature u after feature filtering i to the adaptive residual mechanism sub-component. If it is not the last layer, the current layer number k is incremented by 1, and a message is returned to the message generation component for continued processing.
7. The method for constructing a defense model against graph adversarial attacks according to claim 1, characterized in that The adaptive residual mechanism sub-component uses the adaptive residual mechanism to perform importance assignment on the original node features and the node features u after filtering the features of the last layer, as shown in Equation (11): i as shown in Equation (11): In formula (11), a ∈ R (1) is a learnable parameter, is the initial feature of node v i and is the feature of the node after being processed by the message update component in the (K - 1)-th layer.
8. The method for constructing a defense model against graph adversarial attacks according to claim 1, wherein Feed the output result of the adaptive residual mechanism sub-component into the classification layer for classification.
Citation Information
Patent Citations
An anti-attack defense method for a feature map attention mechanism and application
CN109948658A
Collaborative filtering recommendation algorithm based on graph convolution attention mechanism
CN112905900A