Three-level interactive fusion graph similarity learning method
Through the three-level interactive fusion graph similarity learning method, node embedding learning and graph interaction learning are enhanced, and rich graph interaction features are generated, which solves the problem that the existing technology is difficult to comprehensively capture multi-level and multi-grained information interaction relationships in the graph structure, significantly improving the accuracy and performance of graph similarity learning.
Patent Information
- Application Number
- CN202510068313.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-06
AI Technical Summary
When calculating graph similarity, it is difficult for the prior art to fully capture multi-level and multi-grained information interaction relationships in the graph structure, and the generated interaction features are not rich enough, which affects the accuracy of graph similarity learning.
A three-level interactive fusion graph similarity learning method is proposed. By enhancing the three stages of node embedding learning, dual graph interactive learning and similarity score prediction, fine-grained node-node interaction information, cross-level node-graph interaction information and global graph-graph interaction information are ordered to generate rich graph interaction features.
By orderly fusion of three levels of information, the information flow and continuity of the learning stage are improved, the performance of graph similarity learning is significantly improved, the generated interaction features are richer, and complex relationships between graphs can be expressed more accurately.
Smart Images

Figure CN120107626A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graph similarity calculation, and in particular to a three-level interactive fusion graph similarity learning method. Background Art
[0002] Graph is a mathematical structure that is widely used to model and analyze complex relationships in the real world. In recent years, with the increase in data complexity and correlation, graph-based applications have gained significant attention in many fields. The study of graphs not only helps to understand object relationships, but also promotes the solution of complex problems such as recommendation systems, anomaly detection, and molecular property prediction. In the study of graphs, calculating the similarity between two graphs is a core issue involving a variety of application scenarios, such as chemical molecular structure comparison and traffic network optimization. In order to measure the similarity between paired graphs, graph edit distance (GED) and maximum common subgraph (MCS) are usually used as the main metrics. However, in practice, the cost of accurately calculating the two core operations of GED or MCS between two graphs is very high and has been proven to be an NP-complete problem.
[0003] Due to the contradiction between the importance of accurate calculation of graph similarity and its high computational cost, current research has proposed various approximate algorithms to estimate graph similarity in a fast and heuristic way, which can be divided into (1) traditional algorithmic methods; (2) data-driven neural network methods. Traditional algorithmic methods usually require quite complex design and implementation based on discrete optimization or combinatorial search. For example, A*-Beamsearch (Beam) uses heuristic search to improve the efficiency of the solution, the Hungarian algorithm is used to solve the optimal matching problem and has polynomial time complexity, and VJ achieves efficient association of local structures through feature matching strategies. Compared with traditional algorithms, the purpose of data-driven neural network methods is to design a neural network-based model to optimize parameters by minimizing the loss between the predicted similarity score and the ground truth during training. The trained model can be used to quickly estimate the similarity of any set of input graphs. The computational cost of graph similarity is reduced while ensuring high accuracy.
[0004] In recent studies, models based on graph neural networks (GNNs) have been used for GED approximation and graph similarity estimation, and have been shown to have excellent accuracy while significantly improving the computational speed compared to traditional algorithms. Most existing end-to-end GNN-based graph similarity estimation models can be considered to consist of three parts: (1) graph embedding learning; (2) graph interaction learning; and (3) similarity score estimation. In order to learn the similarity score between two input graphs, some early simple methods transform complex graph structure data into low-dimensional vector representations (graph embeddings) through graph embedding learning, and then combine the graph embedding vectors of the two input graphs to predict their similarity scores. Although this method is effective, it has obvious limitations. By directly mapping the global-level single representation obtained after graph embedding learning to the graph similarity score, it ignores the interaction learning between graph pairs and does not fully model the fine-grained local structural relationship of the input graphs. Recently, in the study of fault-tolerant graph matching technology, the method of using graph matching network (GMN) to capture the interaction information between input graphs and generate interaction features has shown greater advantages. Modeling the similarity score based on these interaction features has been shown to significantly improve the performance of graph similarity learning. These works consider capturing fine-grained node-node interaction information or cross-level node-graph interaction information to identify and capture fine-grained structural similarity features in graph pairs. Although these studies have achieved remarkable results, there are still three key issues that have not received enough attention. (1) Currently, most graph similarity learning studies only focus on a single level of graph interaction learning, such as fine-grained node-node interaction, holistic graph-graph interaction, or cross-level node-graph interaction. However, this single interaction learning method has obvious limitations. It is difficult to fully capture the multi-level and multi-granular information interaction relationships in the graph structure, and it is also difficult to fully express the overall semantics of the graph. Therefore, it is particularly important to study how to effectively integrate multi-level different interaction information in the graph similarity learning process. (2) Previous studies on graph similarity learning all rely on GNN to learn node embeddings. However, GNN mainly updates node representations by aggregating local neighbor information, which makes it limited in modeling complex relationships between remote nodes. As the number of iterations increases, the node representation may be over-smoothed, reducing the discrimination, thereby affecting the accuracy of graph similarity learning. This problem is particularly obvious in scenarios where the global structure is crucial. (3) In most current studies, graph interaction learning methods are effective but relatively simple, and the generated interaction features are not rich enough. Graph-level interaction modeling usually generates graph embeddings by aggregating node embeddings and calculates similarity using inner products or cosine functions. Although this method can capture global information, it often ignores fine-grained local interaction features. Node embedding-based methods focus on fine-grained node-to-node or cross-level node-graph relationship modeling, but the generated interaction features are usually relatively simple, such as node similarity matrices or overall graph embeddings.This single representation is difficult to fully express the complex relationships between graphs, especially in terms of fine-grained local interactions and global information fusion, which limits its applicability in complex scenarios. Summary of the invention
[0005] The purpose of the present invention is to provide a three-level interactive fusion graph similarity learning method to address the above-mentioned problems. In three different learning stages, three levels of graph interaction information are sequentially integrated to promote graph similarity learning performance. Among them, attention is paid to enhancing the learning of node embedding, and rich graph interaction features are generated through a more effective graph interaction learning mode to promote the modeling of graph similarity relationships.
[0006] The technical solution of the present invention is as follows:
[0007] A three-level interactive fusion graph similarity learning method, characterized by comprising:
[0008] Enhanced node embedding learning: learn original node embeddings through multi-layer GIN with skip connections, capture fine-grained node-node interaction information through a style-based multi-head attention mechanism, and use the interaction information to improve the quality of node embeddings;
[0009] Dual graph interaction learning: A coarse-grained and fine-grained aggregation network is used to combine different attention mechanisms at two different levels of coarse-grained and fine-grained to generate graph embedding features of two granularities and fuse them; a node-graph interaction comparison network is used to generate new node embeddings by fusing cross-node-graph interaction information, and finally learns the comparative features of the original node embeddings and the new node embeddings;
[0010] Similarity score prediction: By fusing the comparative features and aggregated features generated by the graph interaction learning module, a fully connected layer is used to perform global graph-graph interaction modeling based on the multi-level features of the two input graphs to determine the relationship between the input graphs and map them into graph-graph similarity scores.
[0011] Furthermore, the coarse-grained and fine-grained aggregation network includes coarse-grained aggregation and fine-grained aggregation; the coarse-grained aggregation aggregates the node embeddings in the entire graph in a coarse-grained manner to generate a graph embedding that combines global information, and learns the weight of each node according to the similarity metric through the global context-aware attention mechanism, specifically including:
[0012]
[0013] in, represents the graph embedding generated by the GA module, Represents the node embedding matrix after node embedding learning is the embedding of node i in row i, σ(·) is the Sigmoid activation function, is a learnable weight matrix, is a global context variable, calculated as follows:
[0014]
[0015] c contains the global structural information and feature information of the graph. The global contextual attention weight of each node embedding can be obtained by the inner product of c and each node embedding. The corresponding node attention weight is applied to the node embedding to adjust the contribution of each node in the generated graph embedding, so that the nodes that are more relevant to the global information contribute more.
[0016] Furthermore, the fine-grained aggregation captures the fine-grained interaction features between each node in a graph and the entire graph, generates a new node embedding that fuses the node-graph interaction information, and aggregates it into a fine-grained interaction graph embedding; specifically, it includes:
[0017] Update node embeddings by cross-graph node attention, and model the similarity relationship between any nodes in two graphs based on the multi-head attention mechanism:
[0018]
[0019] in, Represent the input graph G respectively 1 , G 2 Node features The updated result on the hth head, h∈{1,2,…,h FA}, h FA Represents the number of heads; Represent the input graph G 1 The query, key and value on the hth head in the graph are obtained by multiplying their node features and weight matrices. 2 Using a similar calculation;
[0020] For each input graph, all single-head outputs of the corresponding input graph in the cross-graph node attention are connected to obtain the multi-head node feature output: Output multiple features of different input graphs Input into the feedforward neural network to obtain the input graph node features after cross-graph node attention enhancement:
[0021] G is updated by the cross-graph node attention. 1 and G 2 Node Features Aggregation to customize corresponding graph embedding
[0022]
[0023] Here, Agg(·) represents an aggregation function that generates graph embedding using node features.
[0024] Furthermore, one of global context-aware attention, maximum pooling, and average pooling is used as an aggregation function in the coarse-grained aggregation, and maximum pooling is used as an aggregation function in the fine-grained aggregation.
[0025] Furthermore, the node-graph interactive comparison network includes:
[0026] Node embedding update: Update the node embeddings of a pair of input graphs obtained after enhanced node embedding learning Repeat extension h C times, h C is the number of NGIC heads;
[0027] The weight matrix of the corresponding head Applied to the node embedding copy to linearly transform the node features and generate a node embedding representation specific to each head. The operation on the h-th head is as follows:
[0028]
[0029] On each corresponding head, similarity calculation is performed on the linearly transformed node embeddings to capture the mutual relationship between nodes across the graph and generate features that describe the degree of association between nodes:
[0030]
[0031] in and They represent the node embedding matrix on the hth head. and The embeddings of the i-th and j-th nodes in , represents the similarity score between them;
[0032] From G 1 From the perspective of 2 The weighted average of the similarity scores of all nodes in G with the current node 2 The global information update G 1 Node embedding in:
[0033]
[0034] in, Represents G 1 The updated representation of the i-th node in the h-th head combines the updated representation from G 2global information; similarly, Representation graph G 2 The jth node in the graph merges the hth head with the 1 Node representation after global information update;
[0035] Node embedding comparison: Compare the node embeddings before and after the update to capture and learn the difference features between the two; including:
[0036]
[0037] in, It represents the similarity score of the p-th view under the h-th head, which is a scalar. The symbol ⊙ represents the element-by-element multiplication operation. and Represent the input feature vectors, is the learnable weight vector of the pth view under the hth head. When considering multiple views, let the total number of views be p G , then the trainable weight matrix on each head is On each head, multiple perspectives are combined to obtain a p C Comparative characteristics of dimensions
[0038] From G 1 From the perspective of each head, a multi-view comparison function f is used c For G 1 The original node embedding of the i-th node in Combined with G 2 Enhanced node embedding after the whole graph information is updated Compare them to get the comparative characteristics between them A similar operation is applied to G 2 , to capture G 2 The original embedding of the jth node in With enhanced node embedding Comparative features of The above operation is expressed as:
[0039]
[0040] For the h-th head of the two input graphs, after performing a comparison operation on each original node embedding and the enhanced node embedding, these newly generated node comparison features are collected as G 1 and G 2 The comparative feature matrices are Then connect each head G 1 , G 2 Compare the feature matrix of all nodes to get the multi-head node comparison feature output: For the multi-head node comparison features of two input graphs, the SRM module is used to calibrate the weights of the node comparison feature matrix on each head, so that the features of different heads are appropriately emphasized or weakened, and the multi-head node comparison features after weight calibration are obtained. Then, in order to reduce the dimension of the head, the multi-head node comparison features are input into the feedforward neural network respectively to obtain the final node comparison feature matrix
[0041] Global comparison feature extraction: Use bidirectional LSTM to aggregate the node comparison feature matrix of each input graph to obtain the global comparison features of each input graph.
[0042]
[0043] in, is the node comparison feature matrix of each input graph, and the node comparison feature corresponding to each node The order of arrangement is random;
[0044] Concatenate the hidden vectors of the BiLSTM in both the forward and backward directions as the global comparison features for each input graph
[0045] Furthermore, the multi-layer GIN adopts a jumping strategy to splice the node features and feature dimensions obtained in each iteration, fully integrating multi-layer feature information; the calculation of GIN updating node embedding includes:
[0046]
[0047] in, Respectively represent the node features of node i at the lth (l≥1) layer and the l-1th layer, MLP (l) is the multilayer perceptron at layer l, ∈ (l) is a learnable coefficient, Represents the neighbor set of node i at layer l-1 The sum of all node features in ;
[0048] The MLP structure in GIN is:
[0049]
[0050] Where x represents the input feature vector, and are the weight matrices of the two linear transformations of the lth layer, representing the bias term of the corresponding linear layer, and is the bias term of the two linear transformations of the lth layer, ReLU(·) represents the activation function, and BatchNorm represents batch normalization;
[0051] Concatenate the iterative results of multiple GIN layers in the feature dimension:
[0052]
[0053] Among them, h i It represents the final node representation obtained after skip-GIN learning, and CONCAT represents the feature concatenation operation, that is, the feature vectors of different layers are connected in series in the feature dimension.
[0054] Furthermore, the style-based multi-head attention mechanism captures fine-grained node-node interaction information through the following steps:
[0055] The node embedding generated by the multi-layer GIN layer through jump connection is fused and re-represented by the feedforward neural network to reduce the node feature dimension size, generate a more compact and semantically rich node representation, and identify the node embedding matrix of the whole graph as Where N represents the number of nodes in the graph with the larger number of nodes in the two input graphs, and d represents the feature dimension of the node embedding;
[0056] Use multi-head attention mechanism to dynamically model dependencies between remote nodes;
[0057] SRM recalibrates the head weights in the multi-head attention mechanism, allowing the model to automatically adjust the importance of different heads.
[0058] Furthermore, the multi-head attention mechanism performs interactive learning between nodes in the graph, and calculates the query of the h-th head through linear transformation Key and Value as follows:
[0059]
[0060] in, is the weight matrix of the linear transformation on the hth head, h∈{1,2,…,h S}, and h S Represents the number of heads in MHA;
[0061] The calculation of the attention between nodes is the scaled dot product between matrices. Using the obtained attention Query, Key, Value matrix, the output of the h-th head is expressed as:
[0062]
[0063] Then concatenate the results of each single-head output to get the final output:
[0064]
[0065] Furthermore, the SRM includes a style pooling and a style integration module. The SRM uses style pooling to summarize features across spatial dimensions, extracts style features from the features of each head, and then estimates the recalibration weight of each head through a head-independent style integration module, and finally recalibrates the feature map to emphasize or suppress features of different heads; recalibrating the head weights in the multi-head attention mechanism through SRM includes:
[0066] The SRM module As input, extract style information and generate weights for each head based on the style information as follows:
[0067]
[0068] Among them, α S The specific calculation is as follows:
[0069] α S =StyleI(StyleP(H S )),
[0070] StyleP(x)=[AvgPool(x),StdPool(x)],
[0071] StyleI(x)=σ(BatchNorm(W S ·x+b S )),
[0072] Among them, StyleP(·) and StyleI(·) represent the style pooling and style integration modules in the SRM module, AvgPool(·) and StdPool(·) represent the average pooling and standard deviation pooling, respectively. S and b S represents the weight and bias of the linear transformation, σ(·) is the Sigmoid activation function;
[0073] Multi-node features It is input into the feedforward neural network to reduce the dimensionality of the head, and after making a residual connection with the node embedding matrix H, it is input into the layer normalization module to obtain the updated node features.
[0074] Furthermore, the similarity score prediction includes:
[0075] For each input graph, we first concatenate the outputs from different interactive learning modules to generate the final feature representation for each input graph.
[0076] The final feature representation of each input graph is concatenated and passed as input to multiple standard fully connected layers;
[0077] Use the sigmoid activation function to normalize the final scalar value and convert it into the final similarity score between the two input images;
[0078] The final operation to calculate the graph similarity score is:
[0079]
[0080] The obtained model is in the training set Each element of the training consists of a pair of input graphs and a true similarity score, and the mean square error is used as the final loss function of the model:
[0081]
[0082] Compared with the prior art, the present invention has the following beneficial effects:
[0083] 1. A new graph similarity calculation model TIFN is proposed. For the first time, three different levels of graph interaction information are integrated in an orderly manner. First, fine-grained node-node interaction information is integrated to enhance node embedding, and then cross-level node-graph interaction information is integrated for graph interaction learning. Finally, global-level graph-graph interaction learning is integrated to map features into similarity scores. By integrating three levels of information in an orderly manner, the information flow and continuity of the learning stage are improved, so that the interaction information integrated in the early stage can more effectively promote the effect of later learning.
[0084] 2. A learning model for enhancing node embedding is proposed. In this model, the MLP of GIN is designed to enhance its ability to represent node features. The skip-connection strategy is used to fuse important node features learned in the iterative process of multi-layer GIN to balance the problem of node embedding tending to be smooth. A style-based multi-head attention module (SBMHA) is proposed, which combines the multi-head attention mechanism with the style-based recalibration module (SRM) to capture node-node interaction information to enhance node features.
[0085] 3. A dual graph interaction learning model is proposed to study the importance of rich graph interaction features in graph and similarity learning, and to generate rich graph interaction features in the graph interaction learning process to promote graph similarity learning performance; a coarse-fine granularity aggregation (CFGA) module is designed to learn and fuse graph-level embedding features at both coarse and fine granularities; a node-graph interaction comparison (NGIC) module is designed to enhance node embedding through node-graph interaction information, and then use multiple perspective comparison functions to aggregate the differences before and after the enhancement to obtain global comparison features. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 A framework model diagram of a three-level interactive fusion graph similarity learning method. DETAILED DESCRIPTION
[0087] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0088] The features and performance of the present invention are further described in detail below in conjunction with the embodiments.
[0089] See also Figure 1 , a three-level interactive fusion graph similarity learning method, including:
[0090] Enhanced node embedding learning: Skip-GIN is used to learn the original node embedding, and the style-based multi-head attention module (SBMHA) mechanism is used to capture fine-grained node-node interaction information, and the interaction information is used to improve the quality of node embedding;
[0091] Dual graph interaction learning: The coarse-grained and fine-grained aggregation module (CFGA) combines different attention mechanisms at two different levels of coarse-grained and fine-grained to generate graph embedding features of two granularities and fuse them; the node-graph interactive comparison module (NGIC) is used to generate new node embeddings by fusing cross-node-graph interaction information, and finally learns the comparative features of the original node embeddings and the new node embeddings;
[0092] Similarity score prediction: By fusing the comparative features and aggregated features generated by the graph interaction learning module, a fully connected layer is used to perform global graph-graph interaction modeling based on the multi-level features of the two input graphs to determine the relationship between the input graphs and map them into graph-graph similarity scores.
[0093] The three learning stages of this application integrate three levels of interactive learning modes. The enhanced node embedding learning stage explores effective solutions for generating input graph node embeddings and integrates fine-grained node-node interaction information to enhance node embedding. The dual-graph interaction learning stage captures cross-level node-graph interaction information for graph interaction modeling. In the similarity score prediction stage, global-level graph-graph interaction learning is used to map features to similarity scores.
[0094] Multi-layer GIN uses a jumping strategy to concatenate the node features and feature dimensions obtained in each iteration, fully integrating multi-layer feature information; GIN updates the calculation of node embedding including:
[0095]
[0096] in, Respectively represent the node features of node i at the lth (l≥1) layer and the l-1th layer, MLP (l) is the l-th layer of the multilayer perceptron, ∈ (l) is a learnable coefficient, Represents the neighbor set of node i at layer l-1 The sum of all node features in ;
[0097] In order to enhance the nonlinear expression ability of GIN and enable it to learn the complex features of nodes and graph structures more effectively, the MLP structure in GIN is:
[0098]
[0099] Where x represents the input feature vector, and are the weight matrices of the two linear transformations of the lth layer, representing the bias terms of the corresponding linear layer, and is the bias term of the two linear transformations in the lth layer, ReLU(·) represents the activation function, and BatchNorm represents batch normalization; the two-layer MLP structure combines ReLU and BatchNorm operations, which not only has powerful nonlinear expression capabilities but also ensures good stability. Combined with the feature addition and aggregation mechanism of GIN, it significantly improves GIN's ability to learn node embedding.
[0100] The iteration results of multiple GIN layers are concatenated in the feature dimension to balance the problem of node embedding becoming smoother due to the increase in the number of iterations:
[0101]
[0102] Among them, h i It represents the final node representation obtained after skip-GIN learning, and CONCAT represents the feature concatenation operation, that is, the feature vectors of different layers are connected in series in the feature dimension.
[0103] Although Skip-GIN alleviates the problem of node smoothing caused by multiple iterations, GIN mainly focuses on the aggregation of node features in the local neighborhood, making it difficult to effectively capture the global relationship between remote nodes. Moreover, simply concatenating node features is not enough to fully capture the global correlation information between nodes. Therefore, the SBMHA module is used to further focus on capturing the interaction information between nodes in the graph.
[0104] The style-based multi-head attention mechanism captures fine-grained node-node interaction information through the following steps:
[0105] The node embedding generated by the multi-layer GIN layer through the jump connection is fused and re-represented by the multi-layer features through the feed-forward neural network (FFN), the node feature dimension size is reduced, a more compact and semantically richer node representation is generated, and the node embedding matrix of the whole graph is identified as Where N represents the number of nodes in the graph with the larger number of nodes in the two input graphs, and d represents the feature dimension of the node embedding;
[0106] Use the multi-head attention mechanism (MHA) to dynamically model the dependencies between remote nodes;
[0107] SRM recalibrates the head weights in the multi-head attention mechanism, allowing the model to automatically adjust the importance of different heads.
[0108] The multi-head attention mechanism performs interactive learning between nodes in the graph, and calculates the query of the h-th head through linear transformation of H. Key and Value as follows:
[0109]
[0110] in, is the weight matrix of the linear transformation on the hth head, h∈{1,2,…,h S}, and h S Represents the number of heads in MHA;
[0111] The calculation of the attention between nodes is the scaled dot product between matrices. Using the obtained attention Query, Key, Value matrix, the output of the h-th head is expressed as:
[0112]
[0113] Then concatenate the results of each single-head output to get the final output:
[0114]
[0115] SRM includes style pooling and style integration modules. SRM uses style pooling to summarize features across spatial dimensions, extracts style features from the features of each head, and then estimates the recalibration weights of each head through a head-independent style integration module, and finally recalibrates the feature map to emphasize or suppress the features of different heads; recalibrating the head weights in the multi-head attention mechanism through SRM includes:
[0116] The SRM module As input, extract style information and generate weights for each head based on the style information as follows:
[0117]
[0118] Among them, α S The specific calculation is as follows:
[0119] α S =StyleI(StyleP(H S )),
[0120] StyleP(x)=[AvgPool(x),StdPool(x)],
[0121] StyleI(x)=σ(BatchNorm(W S ·x+b S )),
[0122] Among them, StyleP(·) and StyleI(·) represent the style pooling and style integration modules in the SRM module, AvgPool(·) and StdPool(·) represent the average pooling and standard deviation pooling, respectively.S and b S represents the weight and bias of the linear transformation, σ(·) is the Sigmoid activation function;
[0123] Multi-node features It is input into the feedforward neural network to reduce the dimensionality of the head, and after making a residual connection with the node embedding matrix H, it is input into the layer normalization module to obtain the updated node features.
[0124] SBMHA combines MHA and SRM to model the relationship between nodes in the graph, making up for the shortcomings of GIN in dealing with remote dependencies between nodes. By introducing the SRM module in MHA, the importance of each head is dynamically adjusted to achieve more refined feature fusion. Combining multi-layer GIN with SBMHA for learning node embedding provides a more obvious and effective node feature representation for the next step of graph similarity calculation.
[0125] The coarse-grained aggregation module includes coarse-grained aggregation (GA) and fine-grained aggregation (FA). Coarse-grained aggregation aggregates the node embeddings in the entire graph in a coarse-grained manner to generate a graph embedding that combines global information, and learns the weight of each node according to the similarity metric through the global context-aware attention mechanism, including:
[0126]
[0127] in, represents the graph embedding generated by the GA module, Represents the node embedding matrix after node embedding learning is the embedding of node i in row i, σ(·) is the Sigmoid activation function, is a learnable weight matrix, is a global context variable, calculated as follows:
[0128]
[0129] c contains the global structural information and feature information of the graph. The global contextual attention weight of each node embedding can be obtained by the inner product of c and each node embedding. The corresponding node attention weight is applied to the node embedding to adjust the contribution of each node in the generated graph embedding, so that the nodes that are more relevant to the global information contribute more.
[0130] Fine-grained aggregation captures the fine-grained interaction features between each node in a graph and the entire graph, generates new node embeddings that fuse node-graph interaction information, and aggregates them into fine-grained interaction graph embeddings; specifically, it includes:
[0131] Update node embeddings by cross-graph node attention, and model the similarity relationship between any nodes in two graphs based on the multi-head attention mechanism:
[0132]
[0133] in, Represent the input graph G respectively 1 , G 2 Node features The updated result on the hth head, h∈{1,2,…,h FA}, h FA Represents the number of heads; Represent the input graph G 1 The query, key and value on the hth head in the graph are obtained by multiplying their node features and weight matrices. 2 Using a similar calculation;
[0134] For each input graph, all single-head outputs of the corresponding input graph in the cross-graph node attention are connected to obtain the multi-head node feature output: Output multiple features of different input graphs Input into the feedforward neural network to obtain the input graph node features after cross-graph node attention enhancement:
[0135] G is updated by the cross-graph node attention. 1 and G 2 Node Features Aggregation to customize corresponding graph embedding
[0136]
[0137] Here, Agg(·) represents an aggregation function that generates graph embedding using node features.
[0138] In coarse-grained aggregation, one of global context-aware attention, maximum pooling, and average pooling is used as the aggregation function, while in fine-grained aggregation, maximum pooling is used as the aggregation function.
[0139] In CFGA, through the GA module, the context-aware attention mechanism is introduced to aggregate node embeddings to generate G 1 and G 2 Coarse-grained global graph embedding of In the FA module, CGNA is used for fine-grained node-graph interactions to enhance node embeddings, and the maximum pooling is used as the aggregation function to generate fine-grained interaction graph embeddings. G1 , G 2 The coarse-grained global graph embedding is fused with the fine-grained interaction graph embedding, and the unified aggregation of information is achieved through the addition operation: This fusion strategy can not only fully integrate the information of cross-graph interactions, but also retain the global structural characteristics of each graph, so that the final generated graph embedding reflects multi-level graph information interaction and strengthens feature expression.
[0140] Most of the existing work on capturing the interaction information between graph pairs based on node embedding focuses on modeling the similarity relationship between nodes. However, this node-node level interaction cannot fully capture the global semantic information of the entire graph in many cases, especially in scenarios with complex graph structures and dense relationships between nodes. Relying only on node-node similarity often ignores the interaction features across the entire graph, resulting in insufficient expression of the graph. Therefore, in order to make up for this limitation, the NGIC module is proposed.
[0141] The node-graph interactive comparison module includes:
[0142] Node embedding update: Update the node embeddings of a pair of input graphs obtained after enhanced node embedding learning Repeat extension h C times, h C is the number of NGIC heads;
[0143] The weight matrix of the corresponding head Applied to the node embedding copy to linearly transform the node features and generate a node embedding representation specific to each head. The operation on the h-th head is as follows:
[0144]
[0145] On each corresponding head, similarity calculation is performed on the linearly transformed node embeddings to capture the mutual relationship between nodes across the graph and generate features that describe the degree of association between nodes:
[0146]
[0147] in and They represent the node embedding matrix on the hth head. and The embeddings of the i-th and j-th nodes in , represents the similarity score between them;
[0148] From G 1 From the perspective of 2 The weighted average of the similarity scores of all nodes in G with the current node2 The global information update G 1 Node embedding in:
[0149]
[0150] in, Represents G 1 The updated representation of the i-th node in the h-th head combines the updated representation from G 2 global information; similarly, Representation graph G 2 The jth node in the graph merges the hth head with the 1 Node representation after global information update;
[0151] Node embedding comparison: Compare the node embeddings before and after the update to capture and learn the difference features between the two; including:
[0152]
[0153] in, It represents the similarity score of the p-th view under the h-th head, which is a scalar. The symbol ⊙ represents the element-by-element multiplication operation. and Represent the input feature vectors, is the learnable weight vector of the pth view under the hth head. When considering multiple views, let the total number of views be p C , then the trainable weight matrix on each head is On each head, multiple perspectives are combined to obtain a p 1 Comparative characteristics of dimensions
[0154] From G 2 From the perspective of each head, a multi-view comparison function f is used c For G 1 The original node embedding of the i-th node in Combined with G 2 Enhanced node embedding after the whole graph information is updated Compare them to get the comparative characteristics between them A similar operation is applied to G 2 , to capture G 2 The original embedding of the jth node in With enhanced node embedding Comparative features of The above operation is expressed as:
[0155]
[0156] For the h-th head of the two input graphs, after performing a comparison operation on each original node embedding and the enhanced node embedding, these newly generated node comparison features are collected as G 1 and G 2 The comparative feature matrices are Then connect each head G 1 , G 2 Compare the feature matrix of all nodes to get the multi-head node comparison feature output: For the multi-head node comparison features of two input graphs, the SRM module is used to calibrate the weights of the node comparison feature matrix on each head, so that the features of different heads are appropriately emphasized or weakened, and the multi-head node comparison features after weight calibration are obtained. Then, in order to reduce the dimension of the head, the multi-head node comparison features are input into the feedforward neural network respectively to obtain the final node comparison feature matrix
[0157] Global comparison feature extraction: Use bidirectional LSTM to aggregate the node comparison feature matrix of each input graph to obtain the global comparison features of each input graph.
[0158]
[0159] in, is the node comparison feature matrix of each input graph, and the node comparison feature corresponding to each node The order of arrangement is random;
[0160] Concatenate the hidden vectors of the BiLSTM in both the forward and backward directions as the global comparison features for each input graph In order to reduce the impact of node input arrangement on the BiLSTM aggregation effect, before the node comparison matrix is input into the aggregator, the comparison features of each node are first randomly arranged, thereby reducing the impact of different arrangement orders on the BiLSTM aggregator.
[0161] Similarity score prediction includes:
[0162] For each input graph, we first concatenate the outputs from different interactive learning modules to generate the final feature representation for each input graph.
[0163] The final feature representation of each input graph is concatenated and passed as input to multiple standard fully connected layers;
[0164] Use the sigmoid activation function to normalize the final scalar value and convert it into the final similarity score between the two input images;
[0165] The final operation to calculate the graph similarity score is:
[0166]
[0167] The obtained model is in the training set Each element of the training consists of a pair of input graphs and a true similarity score, and the mean square error is used as the final loss function of the model:
[0168]
[0169] Experimental verification:
[0170] Dataset:
[0171] To evaluate the effect of the model, experiments were conducted on three datasets that are widely used for graph similarity search. Table 1 provides detailed information on the datasets. AIDS700nef comes from the antiviral screening compound database developed by NCI / NIH, which contains 700 graph sets screened from 43,687 compound structures. The Linux dataset comes from 48,747 program dependency graphs (PDGs) generated using CodeSurfer 2.1pl. After screening, 1,000 graphs were finally selected for analysis and research. The IMDB-Multi dataset consists of self-network graphs of 1,500 actors selected from the Internet Movie Database (IMDb) website. In each graph, the nodes represent actors, and if two actors co-star in a movie, they will be connected by an edge.
[0172] For the LINUX and AIDS700nef datasets, since their graph sizes are small and the maximum number of nodes is 10, the A* algorithm can be used to accurately calculate the GED between graph pairs. For the IMDB-Multi dataset, due to the large scale of the graphs, the number of nodes in the largest graph has reached 89, and the A* algorithm cannot reliably calculate the GED between graphs that are too large in a reasonable time. Therefore, three traditional GED approximation algorithms, Beam, Hungarian, and VJ, are used, and the minimum value calculated is used as the true GED of the dataset, because the GED values calculated by these approximate algorithms are greater than the true GED. During the training process, the ground truth graph similarity score is used as the evaluation indicator, so the ground truth GED needs to be converted into the ground truth graph similarity. The specific operations are as follows:
[0173]
[0174] Where |V 1 | and |V 2|Represents G 1 and G 2 The number of nodes in GED(G 1 ,G 2 ) represents G 1 and G 2 The ground truth GED.
[0175] Table 1 Detailed analysis of the dataset. Min|V|, Max|V| and Avg|V| represent the minimum, maximum and average number of nodes in the dataset, respectively. Min|E|, Max|E| and Avg|E| represent the minimum, maximum and average number of edges in the dataset, respectively. Min Degree, Max Degree, Avg Degree represent the minimum, maximum and average degree of nodes in the dataset, respectively. Node Type represents whether the dataset contains node labels and the number of labels. |D| and Pairs represent the total number of graphs and the number of graph pairs in the dataset, respectively.
[0176]
[0177] Experimental setup:
[0178] For each dataset, it is randomly divided into training set, validation set and test set, with the proportions of 60%, 20% and 20% respectively. In the experiment, the model was built and the experiment was conducted based on PyTorch and PyTorch Geometric. During the training process, a learning rate of 5e-4 was used, and 5000 iterations were performed. The validation process started from the 4500th iteration, and the batch size was set to 128. AdamW was selected as the optimizer. All experiments were conducted on a Linux server equipped with 2 Intel Xeon Gold 6133 processors (2.50GHz, 20 cores each) and NVIDIA RTX 4090 graphics card.
[0179] In the node embedding learning stage, four layers of GIN are used to learn node embedding, and the output dimensions of each layer are 64, 64, 32, and 16 respectively. In the graph interaction learning stage, for NGIC, the number of viewpoints p is set to CSet to 128, BiLSTM is used as the aggregation function by default, and its hidden state size is set to 128, as described in 3.4.2. In each input graph, the last two hidden vectors of the BiLSTM in the forward and backward directions are concatenated to obtain a 256-dimensional vector as the global comparison feature of each input graph. In the FA module of CFGA, max pooling is used as the aggregation function by default, and the graph-level embedding dimensions generated by the FA module and the CA module are both set to 256, which is consistent with the hidden state dimension of the BiLSTM aggregator in NGIC. In the similarity score prediction stage, a 4-layer standard fully connected network is used to fuse the multi-level interaction features of the two input graphs and interact between graphs.
[0180] Baseline method
[0181] The baseline methods are divided into two categories: traditional graph edit distance approximation algorithms and data-driven neural network based methods.
[0182] The first category of baseline methods includes three classic graph edit distance (GED) calculation algorithms. (1) Beam Search is an effective heuristic search algorithm that balances computational efficiency and solution quality by limiting the width of the search in the search space. (2) Hungarian algorithm is a classic algorithm for solving the maximum matching of bipartite graphs, especially suitable for the minimum assignment problem in weighted graphs. (3) The Viterbi-Johnson (VJ) Algorithm combines dynamic programming with heuristic search to optimize the matching problem of graph nodes and edges, and can efficiently calculate the GED between two graphs.
[0183] The second category of baselines includes the following 12 neural network models. (1) SimGNN learns node embedding based on GCN, combines histogram features with neural tensor network (NTN) for matching strategy, and calculates the similarity score of graphs. (2) GMN improves on the traditional graph embedding model and proposes an attention-based cross-graph matching mechanism. The node feature propagation process incorporates the cross-information of the two graphs to more effectively extract graph similarity features. (3) GraphSim avoids generating graph-level embeddings, uses CNN to extract features from the similarity matrix transformed from multi-level node embeddings, and uses standard image processing techniques to model the similarity relationship between graphs. (4) MGMN captures the cross-level and graph-level interactive information between graph pairs based on node embedding through a multi-level graph matching network, and effectively integrates multi-level information to calculate graph similarity. (5) GOTSim defines the optimal allocation target based on node embedding, and solves the optimal allocation problem through a differentiable algorithm to measure the similarity between two graphs. (6) H2MN constructs the graph as a hypergraph and combines hypergraph convolution for graph embedding learning, and then learns graph similarity by capturing the semantic similarity of the substructure of the graph through hyperedge pooling and subgraph matching. (7) EGSC uses knowledge distillation technology, combined with slow learning and fast reasoning strategies, to propose a new graph similarity calculation method to improve the reasoning speed while ensuring the calculation accuracy. (8) GREED does not directly predict the graph similarity score, but uses neural networks to learn graph embedding and designs functions to predict GED, providing a more flexible and efficient solution for graph matching and graph similarity calculation tasks. (9) CGMN combines contrastive learning with graph matching networks to achieve efficient and accurate graph similarity calculation through self-supervised learning. (10) NA_GSL introduces multiple attention mechanisms in different stages such as node embedding, graph interaction modeling, and similarity matrix alignment, which effectively improves the accuracy and performance of graph similarity calculation. (11) DeepSIM is a two-branch graph similarity calculation model that optimizes the graph-level feature interaction in NTN by introducing the distance and direction attributes of high-dimensional space, uses global average pooling instead of fake node insertion to unify node embedding, and combines CNN with attention mechanism to extract node features. (12) CLSim combines GCN and MLP to strengthen node embedding learning, uses vector directionality to refine graph-level feature extraction, and adopts cross-layer feature extraction technology to fuse node-level information and graph-level embedding for graph similarity calculation.
[0184] For articles that provide open source code, we used the same hyperparameters as those provided in the original paper to implement all baselines. For articles that do not provide hyperparameters, we carefully tuned the parameters to obtain the best results. For articles that do not provide code, we directly quoted the best experimental results in the original paper for comparison.
[0185] Evaluation Metrics:
[0186] In order to evaluate the model more comprehensively in the experiment, the following evaluation indicators are used:
[0187] (1) The Mean Squared Error (MSE) is a metric commonly used to evaluate the performance of regression models. It aims to measure the difference between the model's predicted value and the actual value. The calculation process first squares the difference (error) between each predicted value and the actual value, and then takes the average of all squared errors to quantify the model's prediction accuracy.
[0188] (2) Spearman's rank correlation coefficient (ρ) measures the monotonic relationship between two variables by comparing the rank differences of the data points. It is applicable to any monotonic relationship, not necessarily a linear relationship. Its value ranges from -1 (perfect negative correlation) to 1 (perfect positive correlation), and 0 indicates no monotonic correlation.
[0189] (3) Kendall's rank correlation coefficient (τ) is similar to ρ. It is a nonparametric statistical method that measures the correlation between two variables by comparing the rank order consistency of all data pairs. Its value is also between -1 and 1, where 1 indicates a perfect positive correlation, -1 indicates a perfect negative correlation, and 0 indicates no correlation. Compared with ρ, τ is more robust and is particularly suitable for situations where the sample size is small or there are many tied ranks in the data.
[0190] (4) Precision at k (p@k) is a commonly used metric in information retrieval and recommendation systems, used to measure the accuracy of a model in the top k recommended items or search results. Specifically, it measures the proportion of the top k results that are related to the true results. In the experiment, p@10 and p@20 were used to evaluate the model effect.
[0191] result:
[0192] This application evaluates the model on three different GED datasets. For each dataset, this application uses 5 evaluation metrics and compares the model of this application with two types of baseline models and a total of 15 state-of-the-art graph similarity calculation methods.
[0193] The experimental results are shown in Table 2. The model of this application shows the best performance in most evaluation indicators of all data sets. Especially on the LINUX and AIDS700nef data sets, the model of this application has achieved the best performance in all evaluation indicators. Specifically, on the LINUX data set, compared with NA_GSL, the model of this application has improved by an average of 9.59%, among which the improvement of the MSE indicator is particularly significant, reaching 45.79%. In addition, on the AIDS700nef data set, compared with GREED, this application also achieved a 7.67% improvement in the MSE indicator.
[0194] For the IMDB-Multi dataset, the experimental results once again demonstrated the strong competitiveness of the method of this application compared with other baseline methods. Specifically, on the MSE indicator, the method of this application achieved the best performance of 0.541, which is better than other models. Among the remaining four evaluation indicators, the method of this application achieved the best in two indicators, and was slightly inferior to DeepSim in the other two indicators. Among them, the highest values of 0.952 and 0.871 were achieved on Spearman'sρ and p@20 respectively; on Kendall'sτ and p@10, 0.876 and 0.862 ranked second, and the performance was close to the best. Judging from the five evaluation indicators, the method of this application always shows the best or suboptimal performance, which further verifies its advantages and reliability in graph similarity calculation.
[0195] Describe the performance of the first and second category baselines, and explain which techniques are effective in improving the results.
[0196] The experimental analysis results show that the first type of baseline model (classical algorithm) performs relatively poorly in all indicators, especially in the MSE indicator, with large errors (such as VJ's MSE as high as 63.863), reflecting its limitations in modeling complex data distributions and nonlinear relationships. In contrast, the second type of baseline model (data-driven neural network method) shows significant advantages in various evaluation indicators by introducing advanced technologies such as graph neural network (GNN), graph interaction learning and attention mechanism.
[0197] The method of this application can still show strong adaptability and stability when dealing with graph data with different characteristics. For various types of graph data, the method of this application can effectively capture the potential similarity of the graph and ensure excellent performance in calculating graph similarity. The model of this application is always better than other baseline models on the three data sets, which shows the superiority and wide applicability of the method of this application in the field of graph similarity calculation, and provides strong support and innovative solutions for various research and applications.
[0198] This application believes that TIFN has achieved comprehensive improvements in the five evaluation indicators of the three datasets compared with other baseline methods, which is mainly attributed to the following three points: (1) The model learns node embeddings that are more suitable for graph similarity calculation; (2) The model learns graph-level embeddings and global comparison features respectively through two graph interaction modules, providing rich multi-level feature expressions for graph similarity learning; (3) The model integrates low-level node-node interactions, cross-level node-graph interactions, and global graph-graph interactions in three different learning stages, which can more accurately capture the complex structural relationships in the graph.
[0199] Table 2 Performance comparison of our method with existing baseline methods on evaluation metrics. Bold and blue are used to highlight the best and second-best results.
[0200]
[0201]
[0202] Ablation experiment:
[0203] This application sets up ablation experiments in TIFN to evaluate the effectiveness of key components. For all data sets, this application sets up the same component ablation comparison. Specifically, (1) HDMIN-w / o skip-connection, this application verifies the superiority of using the skip connection strategy to combine the outputs of multi-layer GIN in the enhanced node embedding learning stage. This application designs to replace the skip connection with residual connection to add the outputs of multi-layer GIN for comparison (2) HDMIN-w / o GA and HDMIN-w / o FA, this application conducts ablation experiments on GA and FA components respectively. These two modules correspond to the coarse-grained and fine-grained graph-level embedding generation in CFGA, aiming to verify the difference between them and evaluate the effectiveness of the final graph embedding method obtained by fusing these two granularity information; (3) HDMIN-w / o CFGA and HDMIN-w / o NGIC, this application explores the effectiveness of rich interactive features for graph similarity learning. For the two-layer graph interaction learning stage, this application sets up ablation of the two graph interaction learning modules of CFGA and NGIC to verify the complementarity and effectiveness of the two modules. (4) HDMIN-w / o SRM, this application removes the SRM component in the SBMHA and the SRM component in the NGIC module for experiments to demonstrate the effectiveness of SRM in calibrating multi-head weight operations.
[0204] The experimental results are shown in Table 3. When the present application uses residual connections instead of skip connections to integrate the output of multi-layer GIN in the enhanced node embedding learning stage, the performance of graph similarity learning decreases significantly, especially on the LINUX and AIDS datasets. The MSE is reduced by 58.63% and 31.85% respectively, verifying the superiority of the original node embedding scheme of multi-layer GIN learning based on the skip connection architecture proposed in the present application. The experiments of TIFN-w / o FA and TIFN-w / o GA show that compared with the method of simply generating graph-level embedding by aggregating node embeddings in previous work, CFGA can learn more comprehensive and rich graph-level embedding features by combining coarse-grained and fine-grained learning strategies. In addition, it was observed in the experiment that the graph similarity learning performance of the TIFN-w / o CFGA model with the CFGA module removed further decreased, which further verified the complementarity of the graph-level embedding generated by GA and FA considering coarse and fine granularity respectively. The performance difference between TIFN-w / o CFGA and TIFN-w / o NGIC reveals that the dual-module approach of learning graph-level embedding features and global comparison features is effective in the graph interaction learning stage. The two features complementarily express graph interaction information from different dimensions, proving the importance of rich graph interaction features for subsequent graph similarity learning.
[0205] Table 3 Ablation experiments on key components of the model on LINX, AIDS700nef, and IMDBMulti
[0206]
[0207]
[0208] This application conducts graph interactive learning based on node embedding, so effective node embedding learning is very important for graph similarity learning. In order to evaluate the impact of different graph embedding learning techniques on node embedding learning, experiments are conducted on three datasets: LINUX, AIDS700nef, and IMDBMulti. Three currently popular graph neural network models are used to replace the GIN in the complete TIFN. They are: (1) Graph Convolutional Neural Network (GCN), which applies convolution operations to graph structures and learns node representation by aggregating and transforming the information of each node and its neighboring nodes. (2) Graph Attention Network (GAT), which is a neural network based on the attention mechanism. In the process of learning node representation, different weights are assigned to neighboring nodes through the attention mechanism, and important nodes are automatically focused on when aggregating information. (3) Graph Sampling and Aggregation (GraphSAGE), which learns node representation by sampling and aggregating the information of neighboring nodes in the local neighborhood of each node. Unlike traditional graph neural networks, GraphSAGE uses a fixed number of neighbors for sampling, thereby improving the processing capability of large-scale graphs.
[0209] The experimental results are shown in Table 4. For all indicators of the three datasets, the GIN method has the best overall performance, especially on the LINUX dataset. However, for the AIDS700nef dataset, the GraphSAGE method performs better in overall performance. This is because the dataset consists of heterogeneous graphs composed of different types of nodes and edges, and the GraphSAGE method has more advantages in processing heterogeneous graphs.
[0210] Table 4 Experimental results of different GNN learning node representations
[0211]
[0212]
[0213] The above-mentioned embodiments only express the specific implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the protection scope of the present application. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the technical solution concept of the present application, and these all belong to the protection scope of the present application.
Claims
1. A three-level interactive fusion graph similarity learning method, characterized in that: include: Enhanced node embedding learning: learn original node embeddings through multi-layer GIN with skip connections, capture fine-grained node-node interaction information through a style-based multi-head attention mechanism, and use the interaction information to improve the quality of node embeddings; Dual graph interaction learning: A coarse-grained and fine-grained aggregation network is used to combine different attention mechanisms at two different levels of coarse-grained and fine-grained to generate graph embedding features of two granularities and fuse them; a node-graph interaction comparison network is used to generate new node embeddings by fusing cross-node-graph interaction information, and finally learns the comparative features of the original node embeddings and the new node embeddings; Similarity score prediction: By fusing the comparative features and aggregated features generated by the graph interaction learning module, a fully connected layer is used to perform global graph-graph interaction modeling based on the multi-level features of the two input graphs to determine the relationship between the input graphs and map them into graph-graph similarity scores.
2. According to claim 1, a three-level interactive fusion graph similarity learning method is characterized in that: The coarse-grained and fine-grained aggregation network includes coarse-grained aggregation and fine-grained aggregation; the coarse-grained aggregation aggregates the node embeddings in the entire graph in a coarse-grained manner to generate a graph embedding that combines global information, and learns the weight of each node according to the similarity metric through the global context-aware attention mechanism, specifically including: in, represents the graph embedding generated by the GA module, Represents the node embedding matrix after node embedding learning is the embedding of node i in row i, σ(·) is the Sigmoid activation function, is a learnable weight matrix, is a global context variable, calculated as follows: c contains the global structural information and feature information of the graph. The global contextual attention weight of each node embedding can be obtained by the inner product of c and each node embedding. The corresponding node attention weight is applied to the node embedding to adjust the contribution of each node in the generated graph embedding, so that the nodes that are more relevant to the global information contribute more.
3. According to claim 2, a three-level interactive fusion graph similarity learning method is characterized in that: The fine-grained aggregation captures the fine-grained interaction features between each node in a graph and the entire graph, generates a new node embedding that fuses the node-graph interaction information, and aggregates it into a fine-grained interaction graph embedding; specifically, it includes: Update node embeddings by cross-graph node attention, and model the similarity relationship between any nodes in two graphs based on the multi-head attention mechanism: in, Represent the input graph G respectively 1 , G 2 Node features The updated result on the hth head, h∈{1,2,…,h FA }, h FA Represents the number of heads; Represent the input graph G 1 The query, key and value on the hth head in the graph are obtained by multiplying their node features and weight matrices. 2 Using a similar calculation; For each input graph, all single-head outputs of the corresponding input graph in the cross-graph node attention are connected to obtain the multi-head node feature output: Output multiple features of different input graphs Input into the feedforward neural network to obtain the input graph node features after cross-graph node attention enhancement: G is updated by the cross-graph node attention. 1 and G 2 Node Features Aggregation to customize corresponding graph embedding Here, Agg(·) represents an aggregation function that generates graph embedding using node features.
4. The three-level interactive fusion graph similarity learning method according to claim 3 is characterized in that: In the coarse-grained aggregation, one of global context-aware attention, maximum pooling, and average pooling is used as the aggregation function, and in the fine-grained aggregation, maximum pooling is used as the aggregation function by default.
5. The three-level interactive fusion graph similarity learning method according to claim 3 is characterized in that: The node-graph interactive comparison network includes: Node embedding update: Update the node embeddings of a pair of input graphs obtained after enhanced node embedding learning Repeat extension h C times, h C is the number of NGIC heads; The weight matrix of the corresponding head Applied to the node embedding copy to linearly transform the node features and generate the node embedding representation of each head. The operation on the h-th head is as follows: On each corresponding head, similarity calculation is performed on the linearly transformed node embeddings to capture the mutual relationship between nodes across the graph and generate features that describe the degree of association between nodes: in and They represent the node embedding matrix on the hth head. and The embeddings of the i-th and j-th nodes in , represents the similarity score between them; From G 1 From the perspective of 2 The weighted average of the similarity scores of all nodes in G with the current node 2 The global information update G 1 Node embedding in: in, Represents G 1 The updated representation of the i-th node in the h-th head combines the updated representation from G 2 global information; similarly, Representation graph G 2 The jth node in the graph merges the hth head with the 1 Node representation after global information update; Node embedding comparison: Compare the node embeddings before and after the update to capture and learn the difference features between the two; including: in, It represents the similarity score of the p-th view under the h-th head, which is a scalar. The symbol ⊙ represents the element-by-element multiplication operation. and Represent the input feature vectors, is the learnable weight vector of the pth view under the hth head. When considering multiple views, let the total number of views be p C , then the trainable weight matrix on each head is On each head, multiple perspectives are combined to obtain a p C Comparative characteristics of dimensions From G 1 From the perspective of each head, a multi-view comparison function f is used c For G 1 The original node embedding of the i-th node in Combined with G 2 Enhanced node embedding after the whole graph information is updated Compare them to get the comparative characteristics between them A similar operation is applied to G 2 , to capture G 2 The original embedding of the jth node in With enhanced node embedding Comparative features of The above operation is expressed as: For the h-th head of the two input graphs, after performing a comparison operation on each original node embedding and the enhanced node embedding, these newly generated node comparison features are collected as G 1 and G 2 The comparative feature matrices are Then connect each head G 1 , G 2 Compare the feature matrix of all nodes to get the multi-head node comparison feature output: For the multi-head node comparison features of two input graphs, the SRM module is used to calibrate the weights of the node comparison feature matrix on each head, so that the features of different heads are appropriately emphasized or weakened, and the multi-head node comparison features after weight calibration are obtained. Then, in order to reduce the dimension of the head, the multi-head node comparison features are input into the feedforward neural network respectively to obtain the final node comparison feature matrix Global comparison feature extraction: Use bidirectional LSTM to aggregate the node comparison feature matrix of each input graph to obtain the global comparison features of each input graph. in, is the node comparison feature matrix of each input graph, and the node comparison feature corresponding to each node The order of arrangement is random; Concatenate the hidden vectors of the BiLSTM in both the forward and backward directions as the global comparison features for each input graph 6. The three-level interactive fusion graph similarity learning method according to claim 1, characterized in that: The multi-layer GIN adopts a skipping strategy to concatenate the node features and feature dimensions obtained in each iteration, fully integrating multi-layer feature information; The calculation of GIN update node embedding includes: in, Respectively represent the node features of node i at the lth (l≥1) layer and the l-1th layer, MLP (l) is the multilayer perceptron at layer l, ∈ (l) is a learnable coefficient, Represents the neighbor set of node i at layer l-1 The sum of all node features in ; The MLP structure in GIN is: Where x represents the input feature vector, and are the weight matrices of the two linear transformations of the lth layer, representing the bias term of the corresponding linear layer, and is the bias term of the two linear transformations of the lth layer, ReLU(·) represents the activation function, and BatchNorm represents batch normalization; Concatenate the iterative results of multiple GIN layers in the feature dimension: Among them, h i It represents the final node representation obtained after skip-GIN learning, and CONCAT represents the feature concatenation operation, that is, the feature vectors of different layers are connected in series in the feature dimension.
7. A three-level interactive fusion graph similarity learning method according to claim 1 or 6, characterized in that: The style-based multi-head attention mechanism captures fine-grained node-node interaction information through the following steps: The node embedding generated by the multi-layer GIN layer through jump connection is fused and re-represented by the feedforward neural network to reduce the node feature dimension size, generate a more compact and semantically rich node representation, and identify the node embedding matrix of the whole graph as Where N represents the number of nodes in the graph with the larger number of nodes in the two input graphs, and d represents the feature dimension of the node embedding; Use multi-head attention mechanism to dynamically model dependencies between remote nodes; SRM recalibrates the head weights in the multi-head attention mechanism, allowing the model to automatically adjust the importance of different heads.
8. The three-level interactive fusion graph similarity learning method according to claim 7 is characterized in that: The multi-head attention mechanism performs interactive learning between nodes in the graph, and calculates the query of the hth head through linear transformation Key and Value as follows: in, is the weight matrix of the linear transformation on the hth head, h∈{1,2,…,h S }, and h S represents the number of heads in MHA; The calculation of the attention between nodes is the scaled dot product between matrices. Using the obtained attention Query, Key, Value matrix, the output of the h-th head is expressed as: Then concatenate the results of each single-head output to get the final output:
9. The three-level interactive fusion graph similarity learning method according to claim 8, characterized in that: The SRM includes a style pooling and a style integration module. The SRM uses style pooling to summarize features across spatial dimensions, extracts style features from the features of each head, and then estimates the recalibration weight of each head through a head-independent style integration module, and finally recalibrates the feature map to emphasize or suppress features of different heads. Recalibrating the head weights in the multi-head attention mechanism through SRM includes: The SRM module As input, extract style information and generate weights for each head based on the style information as follows: Among them, α S The specific calculation is as follows: α S =StyleI(StyleP(H S )), StyleP(x)=[AvgPool(x),StdPool(x)], StyleI(x)=σ(BatchNorm(W S ·x+b S )), Among them, StyleP(·) and StyleI(·) represent the style pooling and style integration modules in the SRM module, AvgPool(·) and StdPool(·) represent the average pooling and standard deviation pooling, respectively. S and b S represents the weight and bias of the linear transformation, σ(·) is the Sigmoid activation function; Multi-node features It is input into the feedforward neural network to reduce the dimensionality of the head, and after making a residual connection with the node embedding matrix H, it is input into the layer normalization module to obtain the updated node features.
10. The three-level interactive fusion graph similarity learning method according to claim 1, characterized in that: The similarity score prediction includes: For each input graph, we first concatenate the outputs from different interactive learning modules to generate the final feature representation for each input graph. The final feature representation of each input graph is concatenated and passed as input to multiple standard fully connected layers; Use the sigmoid activation function to normalize the final scalar value and convert it into the final similarity score between the two input images; The final operation to calculate the graph similarity score is: The obtained model is in the training set The training is performed on i=1,2,…,|D|, where each element contains a pair of input graphs and a true similarity score, and the mean square error is used as the final loss function of the model:
Citation Information
Cited By
Drug resistance prediction system and method based on tensor enhancement graph similarity
CN120429657A
A drug resistance prediction system and method based on tensor enhancement graph similarity
CN120429657B