Graph similarity calculation method based on position-structure graph neural network

By introducing position-structure graph neural network, local-global encoder and weight-based feature fusion module in graph similarity calculation, the problem of neglecting global position relationship and lacking global information capture in the prior art is solved, and a more accurate and robust graph similarity calculation is achieved.

CN119942156APending Publication Date: 2025-05-06XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510035815.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing GNN-based graph similarity calculation method ignores the global positional relationship of nodes, lacks global information capture capabilities, and insufficient utilization of hierarchical information, resulting in inaccurate feature representation, affecting the model's inference performance.

Method used

The graph similarity calculation method based on the location-structure graph neural network is adopted, and node embeddings with different levels of information are generated through a local-global encoder. Combined with the weight-based feature fusion module and the score regression module, the difference information between the graphs and dynamically adjust the hierarchical information.

Benefits of technology

It significantly enhances the expression ability of node embedding, improves the efficiency of utilization of hierarchical information, and improves the accuracy and robustness of the model in graph similarity calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942156A_ABST
    Figure CN119942156A_ABST
Patent Text Reader

Abstract

The invention discloses a graph similarity calculation method based on a position-structure graph neural network. The method comprises the following steps of: 1, inputting an undirected graph pair into a local-global encoder, and generating a node embedding set with different levels of information after passing through the local-global encoder; step 2, a weight-based feature fusion module interacts with the two node embedding sets to capture difference information between graphs so as to generate two graph-level embedding vectors, and the graph-level embedding vectors are vectors representing overall information of the graphs; 3, a score regression module models the similarity relation of the two graphs according to the two graph-level embedded vectors, and finally outputs the similarity score of the undirected graph pair. The method aims at enhancing the expression ability of node embedding and improving the utilization efficiency of hierarchical information, so that the accuracy of the model in graph similarity calculation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of graph similarity calculation, and in particular relates to a graph similarity calculation method based on a position-structure graph neural network. Background Art

[0002] Graph similarity calculation is an important research direction in the field of graph mining and graph analysis, and is widely used in the fields of chemical molecule property prediction, social network analysis, bioinformatics, etc. As one of the commonly used graph similarity measurement methods, Graph Edit Distance (GED) is defined as the minimum number of operations required to transform graph G1 into graph G2, including the addition, deletion, or replacement of nodes or edges. GED provides an intuitive framework to quantify the similarity between graphs. However, the calculation process of GED is generally considered to be an NP-hard problem. As the scale of the graph increases, the computational complexity of GED increases exponentially, which limits its feasibility in large-scale applications.

[0003] In recent years, with the rapid development of Graph Neural Networks (GNN), many GNN-based graph similarity calculation algorithms have emerged. These methods usually use GNN to convert GED into a learnable similarity score, and directly learn the mapping between graph pairs and their similarity scores through an end-to-end network framework.

[0004] As a general framework, existing GNN-based graph similarity calculation methods usually use two encoders with the same structure and shared parameters to extract and aggregate the information of each graph in the graph pair, while introducing a feature fusion module to capture the similarity representation between graphs, and finally regress the similarity score through a multi-layer perceptron (MLP). In this framework, the design of the encoder and feature fusion modules plays an important role. Existing research focuses on optimizing these two modules, but these technologies still have the following major defects:

[0005] ① Ignoring the global position relationship of nodes: In the prior art, encoders usually use GNNs. The message passing mechanism of GNN fundamentally determines that GNN is structure-based, that is, the node representation only depends on the local topological structure of the graph. However, the global position relationship of nodes is often ignored, which limits the distinguishing ability and expressiveness of feature representation. Take chemical molecules as an example: in the same molecule, two atoms with the same neighbors are expected to have similar representations. However, since the two atoms have different positions in the molecule, their functions and effects may also be different, so the similarity between them is limited. Therefore, ignoring the global position relationship will lead to inaccuracy in feature representation, which in turn affects the reasoning performance of the model and the accuracy of the final result.

[0006] ② Lack of global information capture capability: Most existing technologies obtain node embeddings by stacking a small number of GNN layers in the encoder module. This design makes node embeddings mainly perceive local information and fail to capture broader global details in the graph. However, there may be key features in the entire graph that are determined by global relationships or subtle differences, which are difficult to accurately capture by relying solely on local information. This limitation weakens the model's ability to express the overall features of the graph, thereby affecting the overall performance of the model.

[0007] ③ Insufficient use of hierarchical information: Existing technologies usually simply concatenate the graph-level embeddings of each level extracted from the feature fusion module as the input of the final MLP, which means that the model treats the features of each level as equal. However, encoders at different levels aggregate information from neighbors of different ranges: shallow encoders mainly capture local features, while deep encoders usually integrate features of a larger range or even the global range. Therefore, the information encoded in shallow features and deep features is significantly different, and their impact on downstream tasks should also be different. Simple concatenation may not fully utilize the differences in these hierarchical features, thereby affecting the prediction results of the model. Summary of the invention

[0008] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a graph similarity calculation method based on a position-structure graph neural network, which aims to enhance the expressive power of node embedding and improve the utilization efficiency of hierarchical information, thereby improving the accuracy of the model in graph similarity calculation.

[0009] In order to achieve the above object, the technical solution adopted by the present invention is:

[0010] A graph similarity calculation method based on a position-structure graph neural network comprises the following steps;

[0011] Step 1: Input an undirected graph pair to the local-global encoder, and generate two node embedding sets with different levels of information after passing through the local-global encoder;

[0012] Step 2: A weight-based feature fusion module is used to interact with the two node embedding sets to capture the difference information between graphs, and then generate two graph-level embedding vectors. The graph-level embedding vector is a vector that represents the overall information of the graph.

[0013] Step 3: Use the score regression module to model the similarity relationship between the two graphs based on the two graph-level embedding vectors, and finally output the similarity score of the undirected graph pair.

[0014] The undirected graph pair in step 1 is a combination of two undirected graphs to be compared. Each graph in the graph pair aggregates the local structural information and global semantic features of the graph through a local-global encoder, thereby generating respective node embedding sets with different levels of information.

[0015] In step 1, the local-global encoder is composed of three positional and structural graph neural network (PS-GNN) encoders, one feed-forward neural network (FFN) and two Transformer encoders in sequence, and each layer of encoder receives the output of the previous layer as the input of the current layer, and uses the output of the current layer as the input of the next layer. PS-GNN aggregates local information through a message passing mechanism, Transformer encoder aggregates global information through a global attention mechanism, and FFN coordinates the feature representation of the two encoders through feature mapping.

[0016] The specific steps of step 1 are:

[0017] Step (1):

[0018] Given an undirected graph pair (G1, G2), where each graph G i , i∈{1,2}, are all undirected graphs, let graph G i With n i nodes, whose adjacency matrix is ​​A i , the degree matrix is ​​D i , the initial feature matrix is ​​X i , then graph G i The position encoding of node u is calculated by the following formula:

[0019]

[0020] Where W i =A i D i -1 is the random walk matrix, the elements in the matrix (W i ) vu represents the probability of transferring from node v to node u in one step; (W i j ) uu represents the probability of returning to itself after j steps from node u; k is the maximum step length. According to the above formula, we can get graph G i The node position encoding matrix

[0021] The initial feature matrix Xi and the node position encoding matrix P i The concatenated matrix will undergo a linear transformation to unify the input dimension and obtain the initial node embedding Right now:

[0022]

[0023] Where W(0) and b(0) are both trainable parameters;

[0024] Combining the information transmission method of position encoding and graph isomorphism network GIN, PS-GNN is constructed, and its mathematical expression is as follows:

[0025]

[0026] in represents the neighbor set of node v, represents the embedding of node v output by the l-th layer PS-GNN, It represents the embedding of the neighbor node u of node v output by the l-th layer PS-GNN;

[0027] Embed the initial node As input, after passing through 3 layers of PS-GNN, we get a set of node embeddings with different levels of local information. represents the node embedding matrix output by PS-GNN at layer l;

[0028] Use FFN for feature mapping, that is, the output of the last layer of PS-GNN The matrix obtained after feature mapping by FFN is used as the input of the Transformer encoder:

[0029]

[0030] Step (2):

[0031] Given a node embedding The Transformer encoder first updates the node embedding by transferring information across the entire graph through a self-attention mechanism:

[0032]

[0033] in, denote the query matrix, key matrix, and value matrix of the Transformer encoder of the l-3 layers, respectively, and W Q (l) , W K (l) , W v (l)is a trainable parameter; Indicates the dimensions of the key matrix.

[0034] The step (2) is specifically:

[0035] First, by querying the matrix Q (l) and key matrix K (l) , calculate the similarity score between each pair of nodes, and then use the softmax(·) function to convert the similarity score into attention weights, and then apply the attention weights to the value matrix V (l) Perform weighted summation on the nodes to obtain the updated node embedding matrix.

[0036] The Transformer encoder introduces residuals and uses layer normalization (LayerNorm) to normalize each node embedding to stabilize the training process, namely:

[0037]

[0038] The Transformer encoder performs independent nonlinear changes to the embedding of each position through a feed-forward neural network (FFN), and again combines residual connections and layer normalization to stabilize the training process, ultimately generating updated node embeddings.

[0039]

[0040] exist After two layers of Transformer encoders, we get a node embedding set with different levels of global information.

[0041] In step 2, the weight-based feature fusion module first performs a pooling operation on the node embeddings at each level, and converts the node embeddings into graph-level embeddings at the corresponding level;

[0042] The graph-level embeddings are then interacted through a DifferentialAttention (DiffAtt) layer to capture the difference information between graphs.

[0043] Finally, by using the feature fusion layer of Multi-head Self-Attention (MHSA), adaptive fusion weights are generated, and the graph-level embeddings of each level are dynamically adjusted through the weights, and finally fused into a comprehensive graph-level embedding.

[0044] The step 2 is specifically as follows:

[0045] Step (1): Through global context attention pooling, embed the nodes generated by encoders at each level Convert to graph-level embedding of the corresponding level

[0046] Step (2): The graph-level embedding is interacted through the DiffAtt layer to obtain a difference-enhanced graph-level embedding and

[0047] Step (3): Utilize the multi-head self-attention mechanism to adaptively assign weights to the difference-enhanced graph-level embeddings at different levels and fuse them into a comprehensive graph-level embedding.

[0048] The step (1) is specifically:

[0049] Given a node embedding in Representation graph G i The embedding of node u in the l-th layer encoder, the global context attention pooling calculation steps are as follows:

[0050] Attention score calculation: For each input element Calculate its attention score

[0051]

[0052] Where v, W h , b attn is a learnable parameter;

[0053] Attention weight calculation: Use the softmax(·) function to convert the attention score Convert to attention weight

[0054]

[0055] Weighted sum: by attention weights right Weighted summation to get the global context vector graph-level embedding

[0056]

[0057] The resulting global context vector graph-level embedding That is graph G i The graph-level embedding vector at the l-th encoder layer.

[0058] The step (2) is specifically:

[0059] In the DiffAtt layer, we first calculate the graph-level embedding vectors of the same level of graphs G1 and G2 The absolute difference of , and this representation is input into a multi-layer perceptron (MLP) to fine-tune and enhance the difference information, that is:

[0060]

[0061] These differences are then converted into attention weight vectors using the softmax(·) function:

[0062]

[0063] Among them, a (l) Each element in the vector represents the weight assigned to each dimension of the graph-level embedding, and the temperature factor t regulates the smoothness of the difference;

[0064] Next, the weight vector a (l) Respectively with graph-level embedding and Perform element-wise multiplication to obtain the difference-enhanced graph-level embedding and

[0065]

[0066] The step (3) is specifically:

[0067] G i The difference enhancement graph-level embedding is concatenated into a matrix in Representation graph G i For the difference-enhanced graph-level embedding vector corresponding to the encoder at the lth layer, the process of allocating weights using the multi-head self-attention mechanism is shown in the following mathematical formula:

[0068] MultiHead(Q,K,V)=Concat(head1,...,head h )W o

[0069] U′ i =LayerNorm(U i +MultiHead(Q,K,V)

[0070]

[0071] Here, head j =Attention(Q j , K j , V j ),in They represent the query matrix, key matrix, and value matrix of the j-th head respectively. is a trainable parameter; h is the number of heads to be determined by the model; Wo is a trainable parameter;

[0072] After assigning weights through the multi-head self-attention mechanism, the resulting matrix Each row in is the weighted graph-level embedding of each level. By adding up the rows in , we can get the fused comprehensive graph-level embedding Right now

[0073] The score regression module in step 3 uses a multilayer perceptron (MLP) as a regression head to map the extracted comprehensive graph-level embedding to a scalar value as the similarity score of the input graph pair;

[0074] That is, the comprehensive graph-level embedding of graph G1 and graph G2 and After concatenation, a multi-layer perceptron (MLP) is used to map it to the similarity score of the image pair, that is,

[0075]

[0076] A graph similarity calculation system based on a position-structure graph neural network for implementing the method comprises a local-global encoder, a weight-based feature fusion module, and a score regression module;

[0077] The local-global encoder module aggregates local and global information of the graph based on the structure and features of the input graph to generate a set of node embeddings with different levels of information;

[0078] The weight-based feature fusion module captures the difference information between graphs through the interaction between graph pairs and generates graph-level embedding;

[0079] The score regression module predicts the similarity score of the input graph pair based on the graph-level embedding information.

[0080] Beneficial effects of the present invention:

[0081] The present invention combines position coding with traditional GNN and designs a PS-GNN module. The PS-GNN module provides global position awareness for each node by introducing position coding. While aggregating local topological information using the GNN message passing mechanism, the collective position coding enhances global characteristic modeling, allowing the model to effectively distinguish nodes with similar neighbor structures but different global positions. This design significantly enhances the expressive power of node embedding and improves the adaptability and robustness of the model to graphs with complex structures.

[0082] The local-global information capture framework designed by the present invention uses GNN to efficiently capture the local topological structure information of nodes on the one hand, and uses the Transformer encoder to directly transmit the node information of the whole graph on the other hand, thereby capturing the global context information and long-distance node dependencies. After combining the two, GNN makes up for the deficiency of Transformer in modeling local details, while Transformer expands the perception range of GNN and enhances the global feature capture capability of the model.

[0083] The weight-based hierarchical information fusion strategy proposed in this invention dynamically adjusts the ratio of shallow layers (local information) and deep layers (global information) through adaptive weights, so that the model can effectively focus on the key feature layer, suppress irrelevant or redundant information, and make full use of the feature expression capabilities of each layer, thereby enhancing the stability and robustness of the model. In addition, the weight distribution provides a quantitative basis for the importance of each layer of features, which helps to analyze the actual contribution of each feature layer in a specific task and enhance the interpretability of the model.

[0084] In summary, the present invention effectively solves the problems of the prior art ignoring the global position relationship of nodes, lacking the ability to capture global information, and insufficient use of hierarchical information through the position-structure graph neural network (PS-GNN) module, the local-global information capture framework, and the weight-based hierarchical information fusion strategy. The method proposed by the present invention can achieve better reasoning effects in both simple and complex scenarios, and has a certain degree of interpretability. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] Figure 1 This is the overall architecture diagram of the model of the present invention.

[0086] Figure 2 Schematic diagram of the local-global encoder structure of the present invention.

[0087] Figure 3 Schematic diagram of the feature fusion module structure. DETAILED DESCRIPTION

[0088] The present invention will be further described in detail below in conjunction with the accompanying drawings.

[0089] The technical solution of the present invention relates to a graph similarity calculation method based on a position-structure graph neural network. In order to facilitate a more intuitive understanding of the structure, working principle and interaction between the modules of the entire invention, the technical solution is described in detail as follows in conjunction with the accompanying drawings.

[0090] From the attached Figure 1As can be seen in the figure, the core structure of the whole invention is mainly divided into three modules: local-global encoder, weight-based feature fusion module, and the final score regression module. Each module cooperates with each other to form the graph similarity calculation method based on position-structure graph neural network proposed in the present invention.

[0091] ① Local-Global Encoder Module: This module aggregates the local and global information of the graph based on the structure and features of the input graph to generate a set of node embeddings with different levels of information.

[0092] In the attached Figure 2 In the figure, we can see that the local-global encoder module consists of 3 Positional and Structural Graph Neural Network layers (PS-GNN), 1 Feed-Forward Neural Network layer (FFN) and 2 Transformer encoder layers in sequence. Each encoder layer receives the output of the previous layer as the input of the current layer, and uses the output of the current layer as the input of the next layer. PS-GNN aggregates local information through the message passing mechanism, Transformer encoder aggregates global information through the global attention mechanism, and FFN coordinates the feature representation of the two encoders through feature mapping.

[0093] Finally, the module collects the node embeddings generated by encoders at each layer as the output of the current module.

[0094] ② Weight-based feature fusion module: This module aims to capture the difference information between graphs through the interaction between graph pairs, and finally generate graph-level embeddings that can represent each input graph.

[0095] In the attached Figure 3 In the figure, we can see that the module first performs a pooling operation on the node embeddings of each level of the input graph pair, converting them into graph-level embeddings of the corresponding level. Subsequently, these graph-level embeddings interact through the Differential Attention (DiffAtt) layer to capture the difference information between graphs. Finally, by using the feature fusion layer of Multi-head Self-Attention (MHSA), adaptive fusion weights are generated, and the graph-level embeddings of each level are dynamically adjusted through the weights, and finally fused into a comprehensive graph-level embedding.

[0096] ③ Score regression module: The main function of this module is to predict the similarity score of the input image pair based on the image-level embedding information. This module uses a multilayer perceptron (MLP) as a regression head to map the extracted features to a scalar value as the similarity score of the input image pair.

[0097] Principle description: This invention adopts a local-global encoder module to extract and aggregate the local and global information of each graph, and introduces a weighted feature fusion module to capture the similarity representation between graphs and dynamically fuse the hierarchical information. Finally, the sum information of the graph pair is mapped to a scalar as the similarity score of the graph pair through the score regression module.

[0098] ① Principle of the local-global encoder module: This module introduces position encoding in the node representation, aiming to capture the global position relationship of the node in the graph. By combining position encoding with node features, the global position information and neighbor information of the node can be effectively integrated during the message passing process, thereby generating node embeddings with both local structural information and global position information. PS-GNN and Transformer encoders are the core components of message passing: PS-GNN integrates position encoding into the neighbor information aggregation process through a structured message passing mechanism, strengthening the perception of the topological structure between nodes; Transformer encoders use a global self-attention mechanism to dynamically adjust the weights of global message passing to further capture complex global dependencies. Since the GNN and Transformer encoders have different perception ranges, in order to ensure that the local characteristics of GNN and the global characteristics of Transformer do not conflict during joint optimization, a feature mapping layer is introduced between the two modules, which consists of a feedforward neural network FFN.

[0099] (1) Position encoding and PS-GNN principle:

[0100] Given an undirected graph G = (V, E), where V represents the node set of the graph, E represents the edge set of the graph, the adjacency matrix of the graph is set to A, and the degree matrix is ​​set to D, then the position encoding of the node u of the graph G is calculated by the following formula:

[0101]

[0102] Where W = AD -1 is the random walk matrix, the elements in the matrix (W) vu represents the probability of transferring from node v to node u in one step; (W j ) uu represents the probability of returning to itself after j steps from node u; k is the maximum step length;

[0103] Through position encoding and combining the information transmission method of graph isomorphic network GIN, PS-GNN is designed, and its mathematical expression is as follows:

[0104]

[0105] in represents the neighbor set of node v, represents the embedding of node v in the l-th layer of PS-GNN, It represents the node embedding of node v’s neighbor node u in the l-th layer of PS-GNN;

[0106] In addition to the random walk position coding used in the present invention, index-based position coding and shortest path-based position coding may also be used.

[0107] Index-based position encoding: Assign a unique identifier defined by the node index to each node (artificially or adaptively), and encode the node's position through the unique identifier.

[0108] Average shortest path based position encoding: Encode the position of a node based on the average shortest path distance from the node to all other nodes.

[0109] (1) Transformer encoder principle:

[0110] Given a node embedding H (l) , the Transformer encoder first updates the node embedding by transferring the full-graph information through the self-attention mechanism:

[0111]

[0112] Where Q = H (l) W Q , K=H (l) W k , V=H (l) W V For query matrix, key matrix and value matrix. In order to stabilize the model learning, the Transformer encoder introduces residual connection and layer normalization (LayerNorm):

[0113] O (l) =LayerNorm(H (l) +Attention(Q,K,V)

[0114] Considering that the self-attention mechanism may not be able to fit complex processes well enough, the Transformer encoder then adds a feed-forward neural network (FFN) to enhance the model’s capabilities, and combines residual connections and layer normalization to generate updated node embeddings H (l+1) :

[0115] H (l+1) =LayerNorm(o (l) +FFN(O (l) ))

[0116] The effect is achieved by stacking more GNN layers instead of using Transformer to capture global information, because as the number of GNN layers increases, the more information it can perceive.

[0117] Use Graph Transformer directly. Graph Transformer is a Transformer with a graph structure. It combines the global attention mechanism of Transformer and the characteristics of GNN considering the topological properties of the graph, and can capture local and global information at the same time.

[0118] ② Principle of weight-based feature fusion module: In order to reduce the computational complexity of the model, the feature fusion module mainly uses coarse-grained graph-level embedding information for interaction, rather than fine-grained node embedding. Therefore, firstly, the node embeddings of each level are converted into graph-level embeddings of the corresponding level through pooling operations. In order to make the graph-level embeddings more expressive, global contextual attention pooling is used. Specifically, given a node embedding H = [h1, h2, ..., h n ] T , where h t Representing the node embedding of node t, the global context attention pooling calculation steps are as follows:

[0119] 1. Attention score calculation: For each input element h t , calculate its attention score e t :

[0120] e t =v T tanh(W h h t +b attn )

[0121] Where v, W h , b attn For each h t are all independent learnable parameters (i.e., non-shared).

[0122] 2. Attention weight calculation: Use the softmax(·) function to convert the score into attention weight:

[0123]

[0124] 3. Weighted summation: The global context vector c is obtained by weighted summation of ht through attention weights:

[0125]

[0126] The resulting global context vector c is the required graph-level embedding.

[0127] Next, these graph-level embeddings are interacted through the DiffAtt layer. The DiffAtt layer aims to emphasize the structural differences between the two graphs and guide the attention mechanism through these difference information. This method amplifies the structural differences between the graphs while suppressing the common structural features, enhancing the model's ability to distinguish between graphs. Specifically: In the DiffAtt layer, first calculate the graph G i and Figure G j The graph-level embedding vector h of the same level i ,h j This computational approach produces a new representation that captures the differences between graph-level embeddings. Next, this representation is fed into a multi-layer perceptron (MLP) to fine-tune and enhance the difference information, i.e., h diff =MLP(|h i -h j |), and use the softmax(·) function to convert these differences into attention weights, that is, α = softmax(h diff ·t -1 ). The value of α determines the weight assigned to each dimension of the graph-level embedding, where a higher value indicates that the graph will show greater differences in that dimension. The temperature factor t can adjust the smoothness of this difference. Subsequently, α is element-wise multiplied with the graph-level embedding to obtain the enhanced graph-level embedding and This process selectively enhances the structural differences between graphs while weakening the shared structural features.

[0128] Finally, in the feature fusion layer, a multi-head self-attention mechanism is used to adaptively assign weights to graph-level embeddings at different levels and fuse them into a comprehensive graph-level embedding.

[0129] Specifically: Given a hierarchical graph-level embedding matrix of graph G in represents the graph-level embedding vector corresponding to the t-th layer encoder. The calculation formula of the multi-head self-attention mechanism is as follows:

[0130] MultiHead(Q,K,V)=Concat(head1,...,head h )W o

[0131] O1=LayerNorm(U+MultiHead(Q, K, V))

[0132] O2=LayerNorm(O1+FFN(O1))

[0133] in Note that at this time O2 = [o 1 ,…,o m ] T Each row in represents the weighted graph-level embedding of each level. Therefore, we only need to add up the rows in O2 to get the fused comprehensive graph-level embedding h G ,Right now

[0134] The present invention uses a multi-head attention mechanism to dynamically allocate weights and fuse hierarchical information into a comprehensive information.

[0135] Another approach is to use a convolution-based fusion strategy, which treats features at different levels as "channels" and fuses them through one-dimensional convolution operations. The convolution kernel can automatically learn the weights of specific channels.

[0136] Use RNN and its variants (such as LSTM, GRU, Bi-LSTM, etc.) for fusion. This method regards features at different levels as sequences, and transfers information at different levels between different time steps through the recurrence mechanism of RNN, thereby achieving the effect of hierarchical information fusion.

[0137] ③ Principle of score regression module: This module converts graph G i and Figure G j The sum of graph-level embeddings and After concatenation, a multi-layer perceptron (MLP) is used to map it to the similarity score of the image pair, that is,

[0138]

[0139] A graph similarity calculation method based on a position-structure graph neural network comprises the following steps;

[0140] Step 1: Input an undirected graph pair, which is the combination of two undirected graphs to be compared. Each graph in the graph pair is aggregated with local structural information and global semantic features through a local-global encoder to generate its own node embedding set with different levels of information;

[0141] The undirected graph pair in step 1 is a combination of two undirected graphs to be compared. Each graph in the graph pair aggregates the local structural information and global semantic features of the graph through a local-global encoder, thereby generating respective node embedding sets with different levels of information.

[0142] In the step 1, the local-global encoder includes three positional and structural graph neural network (PS-GNN) encoders, one feed-forward neural network (FFN) and two Transformer encoders;

[0143] The PS-GNN encoder in step 1 aggregates local information through a message passing mechanism to capture the neighborhood features and local structural relationships of nodes;

[0144] The Transformer encoder aggregates global information through a global attention mechanism to obtain the global context and long-range dependency characteristics of the entire graph;

[0145] The FFN coordinates the feature representations of the two encoders through feature mapping to avoid conflicts during the optimization process.

[0146] The specific steps of step 1 are:

[0147] Step (1):

[0148] Given an undirected graph pair (G1, G2), where each graph G i , i∈{1,2}, are all undirected graphs, let graph G i With n i nodes, whose adjacency matrix is ​​A i , the degree matrix is ​​D i , the initial feature matrix is ​​X i , then graph G i The position encoding of node u is calculated by the following formula:

[0149]

[0150] Where W i =A i D i -1 is the random walk matrix, the elements in the matrix (W i ) vu represents the probability of transferring from node v to node u in one step; (W i j ) uu represents the probability of returning to itself after j steps from node u; k is the maximum step length. According to the above formula, we can get graph G i The node position encoding matrix

[0151] In order to make the initial node embedding contain its own feature information and position information, the initial feature matrix X i and the node position encoding matrix P i Then, in order to facilitate subsequent operations and calculations, the concatenated matrix will undergo a linear transformation to unify the input dimension, thereby obtaining the initial node embedding Right now:

[0152]

[0153] Where W(0) and b(0) are both trainable parameters;

[0154] Combining the information transmission method of position encoding and graph isomorphism network GIN, PS-GNN is designed, and its mathematical expression is as follows:

[0155]

[0156] in represents the neighbor set of node v, represents the embedding of node v output by the l-th layer PS-GNN, It represents the embedding of the neighbor node u of node v output by the l-th layer PS-GNN;

[0157] Embed the initial node As input, after passing through 3 layers of PS-GNN, we get a set of node embeddings with different levels of local information. represents the node embedding matrix output by PS-GNN at layer l;

[0158] Since the ranges perceived by PS-GNN and Transformer encoders are different, in order to ensure that the local characteristics of PS-GNN and the global characteristics of Transformer encoder do not conflict during optimization, FFN is used for feature mapping, that is, the output of the last layer of PS-GNN is The matrix obtained after feature mapping by FFN is used as the input of the Transformer encoder:

[0159]

[0160] Step (2):

[0161] Given a node embedding The Transformer encoder first updates the node embedding by transferring information across the entire graph through a self-attention mechanism:

[0162]

[0163] in, denote the query matrix, key matrix, and value matrix of the Transformer encoder of the l-3 layers, respectively, and W Q (l) , W K (l) , W V (l) is a trainable parameter; Indicates the dimensions of the key matrix.

[0164] First, by querying the matrix Q (l) and key matrix K (l) The similarity scores between each pair of nodes are calculated by matrix multiplication. The similarity scores are then converted into attention weights using the softmax(·) function. The attention weights are then applied to the value matrix V (l) Perform weighted summation on the nodes to obtain the updated node embedding matrix.

[0165] In order to stabilize the model learning, the Transformer encoder introduces residual connections to avoid the problem of gradient disappearance, and uses layer normalization (LayerNorm) to standardize each node embedding to stabilize the training process, that is:

[0166]

[0167] In addition, in order to improve the feature expression ability of the model, the Transformer encoder adds a feedforward neural network (FFN) to perform independent nonlinear changes on the embedding of each position, and again combines residual connections and layer normalization to stabilize the training process, and finally generates updated node embeddings

[0168]

[0169] Therefore, in After two layers of Transformer encoders, we can get a node embedding set with different levels of global information.

[0170] In summary, after the initial node embeddings of graphs G1 and G2 pass through 3 layers of PS-GNN, 1 layer of FFN, and 2 layers of Transformer encoders, we can get the node embedding set of graph G1: And the node embedding set of graph G2

[0171] Step 2: The weighted feature fusion module interacts with the two node embedding sets generated in the previous step to capture the difference information between the graphs, and then generates two graph-level embedding vectors to represent the overall information of the two comparison graphs respectively;

[0172] In step 2, the weight-based feature fusion module first performs a pooling operation on the node embeddings of each level and converts them into graph-level embeddings of the corresponding level;

[0173] These graph-level embeddings are then interacted through a DifferentialAttention (DiffAtt) layer to capture the difference information between graphs.

[0174] Finally, by using the feature fusion layer of Multi-head Self-Attention (MHSA), adaptive fusion weights are generated, and the graph-level embeddings of each level are dynamically adjusted through the weights, and finally fused into a comprehensive graph-level embedding.

[0175] The step 2 is specifically as follows:

[0176] First, the nodes generated by the encoders at each level are embedded into the global context attention pool Convert to graph-level embedding of the corresponding level

[0177] Furthermore, the conversion is specifically:

[0178] Given a node embedding in Representation graph G i The embedding of node u in the l-th layer encoder. The calculation steps of global context attention pooling are as follows:

[0179] Attention score calculation: For each input element Calculate its attention score

[0180]

[0181] Where v, W h , b attn is a learnable parameter;

[0182] Attention weight calculation: Use the softmax(·) function to convert the attention score Convert to attention weight

[0183]

[0184] Weighted sum: by attention weights right Weighted summation to get the global context vector

[0185]

[0186] The resulting global context vector That is graph G i The graph-level embedding vector at the l-th encoder layer.

[0187] Through the above steps, you can get the graph-level embedding set of graph G1 And the graph-level embedding set of graph G2

[0188] Next, the graph-level embedding is interacted through the DiffAtt layer;

[0189] In the DiffAtt layer, we first calculate the graph-level embedding vectors of the same level of graphs G1 and G2 The absolute difference of , and this representation is input into a multi-layer perceptron (MLP) to fine-tune and enhance the difference information, that is:

[0190]

[0191] These differences are then converted into attention weight vectors using the softmax(·) function:

[0192]

[0193] Among them, α (l) Each element in the vector represents the weight assigned to each dimension of the graph-level embedding, and the temperature factor t regulates the smoothness of the differences.

[0194] Next, the weight vector α (l) Respectively with graph-level embedding and Perform element-wise multiplication to obtain the difference-enhanced graph-level embedding and

[0195]

[0196] Through the above steps, we can obtain the difference enhanced graph-level embedding set of graph G1 And the difference enhanced graph-level embedding set of graph G2

[0197] Finally, the multi-head self-attention mechanism is used to adaptively assign weights to the difference-enhanced graph-level embeddings at different levels and fuse them into a comprehensive graph-level embedding.

[0198] Specifically:

[0199] G i The difference enhancement graph-level embedding is concatenated into a matrix in Representation graph Gi The difference enhancement graph-level embedding vector corresponding to the encoder at the lth layer. The process of allocating weights using the multi-head self-attention mechanism is shown in the following mathematical formula:

[0200] MultiHead(Q,K,V)=Concat(head1,...,head h )W o

[0201] U i ′=LayerNorm(U i +MultiHead(Q,K,V)

[0202]

[0203] Here, head j =Attention(Q j , K j , V j ),in They represent the query matrix, key matrix, and value matrix of the j-th head respectively. is a trainable parameter; h is the number of heads to be determined by the model; W o is a trainable parameter.

[0204] After assigning weights through the multi-head self-attention mechanism, the resulting matrix Each row in is the weighted graph-level embedding of each level. By adding up the rows in , we can get the fused comprehensive graph-level embedding Right now

[0205] Through the above steps, we can get the comprehensive graph-level embedding of graph G1 And the comprehensive graph-level embedding of graph G2

[0206] Step 3: The score regression module models the similarity relationship between the two graphs based on the two graph-level embedding vectors generated in the previous step, and finally outputs the similarity score of the undirected graph pair.

[0207] The score regression module in step 3 uses a multilayer perceptron (MLP) as a regression head to map the extracted comprehensive graph-level embedding to a scalar value as the similarity score of the input graph pair.

[0208] The step 3 is specifically as follows:

[0209] Comprehensive graph-level embedding of graphs G1 and G2 and After concatenation, a multi-layer perceptron (MLP) is used to map it to the similarity score of the image pair, that is,

[0210]

[0211] The action relationship between each module is mainly reflected in the transmission of information flow:

[0212] The local-global encoder module receives the node features and topological structure of the input graph, extracts local information through the message passing mechanism, and integrates global information with the help of the global attention mechanism to generate a node embedding set containing multi-level information as the input of subsequent modules.

[0213] The weight-based feature fusion module receives the node embedding set output from the local-global encoder module, fuses the node embedding features using an adaptive weight mechanism, and generates comprehensive graph-level features representing the input graph, further capturing the similarity information between graph pairs.

[0214] The score regression module inputs the graph-level embedding generated by the weighted feature fusion module into the regression model, maps it through a multi-layer perceptron (MLP), and finally predicts the similarity score of the graph pair.

[0215] Combined with the attached figure, the interaction between these modules clearly shows the complete process of the model from graph pair input to similarity score output. Through clear functional division and information transmission between modules, the modules gradually collaborate to complete the accurate calculation of graph similarity.

[0216] The position-structured graph neural network (PS-GNN) module, on the one hand, combines appropriate position encoding with node features to introduce global position awareness; on the other hand, it combines position encoding with the message passing mechanism of GNN to extract local topological information and capture global position information, while preventing the position information from being submerged in deep propagation.

[0217] The local-global information capture framework introduces a feature mapping module between the GNN and Transformer encoder to ensure the alignment of features across modules.

[0218] Weight-based hierarchical information fusion automatically adjusts the importance ratio of shallow features and deep features according to task requirements through adaptive weights.

Claims

1. A graph similarity calculation method based on position-structure graph neural network, characterized in that: The steps include: Step 1: Input an undirected graph pair to the local-global encoder, and generate two node embedding sets with different levels of information after passing through the local-global encoder; Step 2: A weight-based feature fusion module is used to interact with the two node embedding sets to capture the difference information between graphs and generate two graph-level embedding vectors. Step 3: Use the score regression module to model the similarity relationship between the two graphs based on the two graph-level embedding vectors, and finally output the similarity score of the undirected graph pair.

2. According to claim 1, a graph similarity calculation method based on position-structure graph neural network is characterized in that: The undirected graph pair in step 1 is a combination of two undirected graphs to be compared. Each graph in the graph pair aggregates the local structural information and global semantic features of the graph through a local-global encoder, thereby generating respective node embedding sets with different levels of information. The local-global encoder is composed of three position-structure graph neural network encoders, one feedforward neural network and two Transformer encoders in sequence. Each layer of encoder receives the output of the previous layer as the input of the current layer, and uses the output of the current layer as the input of the next layer. The position-structure graph neural network aggregates local information through a message passing mechanism, the Transformer encoder aggregates global information through a global attention mechanism, and the feedforward neural network coordinates the feature representations of the two encoders through feature mapping.

3. A graph similarity calculation method based on position-structure graph neural network according to claim 2, characterized in that: The specific steps of step 1 are: Step (1): Given an undirected graph pair (G1, G2), where each graph G i ,i∈{1,2}, are all undirected graphs, let graph G i With n i nodes, whose adjacency matrix is ​​A i , the degree matrix is ​​D i , the initial feature matrix is ​​X i , then graph G i The position encoding of node u is calculated by the following formula: Where W i =A i D i -1 is the random walk matrix, the elements in the matrix (W i ) vu represents the probability of transferring from node v to node u in one step; (W i j ) uu represents the probability of returning to itself after j steps from node u; k is the maximum step length. According to the above formula, we get graph G i The node position encoding matrix The initial feature matrix X i and the node position encoding matrix P i The concatenated matrix will undergo a linear transformation to unify the input dimension and obtain the initial node embedding Right now: Where W (0) and b (0) All are trainable parameters; Combining the information transmission method of position encoding and graph isomorphism network GIN, a position-structure graph neural network is constructed, and its mathematical expression is as follows: in represents the neighbor set of node v, represents the embedding of node v output by the l-th layer PS-GNN, It represents the embedding of the neighbor node u of node v output by the l-th layer PS-GNN; Embed the initial node As input, after passing through 3 layers of PS-GNN, we get a set of node embeddings with different levels of local information. represents the node embedding matrix output by PS-GNN at layer l; Use a feedforward neural network for feature mapping, that is, the output of the last layer of position-structure graph neural network The matrix obtained after feature mapping by the feedforward neural network is used as the input of the Transformer encoder: Step (2): Given a node embedding The Transformer encoder first updates the node embedding by transferring information across the entire graph through a self-attention mechanism: in, denote the query matrix, key matrix, and value matrix of the Transformer encoder of the l-3 layers, respectively, and W Q (l) , W K (l) , W V (l) is a trainable parameter; Indicates the dimensions of the key matrix.

4. The method for calculating graph similarity based on position-structure graph neural network according to claim 3, characterized in that: The step (2) is specifically: First, by querying the matrix Q (l) and key matrix K (l) , calculate the similarity score between each pair of nodes, and then use the softmax(·) function to convert the similarity score into attention weights, and then apply the attention weights to the value matrix V (l) Perform weighted summation on the nodes to obtain the updated node embedding matrix; The Transformer encoder introduces residuals and uses layer normalization to normalize each node embedding to stabilize the training process, namely: The Transformer encoder performs independent nonlinear changes to the embedding of each position through a feedforward neural network, and again combines residual connections and layer normalization to stabilize the training process, ultimately generating updated node embeddings. exist After two layers of Transformer encoders, we get a node embedding set with different levels of global information.

5. The method for calculating graph similarity based on position-structure graph neural network according to claim 4, characterized in that: The step 2 is specifically as follows: Step (1): Through global context attention pooling, embed the nodes generated by encoders at each level Convert to graph-level embedding of the corresponding level Step (2): Interact the graph-level embedding through the difference attention layer to obtain a difference-enhanced graph-level embedding and Step (3): Utilize the multi-head self-attention mechanism to adaptively assign weights to the difference-enhanced graph-level embeddings at different levels and fuse them into a comprehensive graph-level embedding.

6. The method for calculating graph similarity based on position-structure graph neural network according to claim 5, characterized in that: The step (1) is specifically: Given a node embedding in Representation graph G i The embedding of node u in the l-th layer encoder, the global context attention pooling calculation steps are as follows: Attention score calculation: For each input element Calculate its attention score Where v, W h , b attn is a learnable parameter; Attention weight calculation: Use the softmax(·) function to convert the attention score Convert to attention weight Weighted sum: by attention weights right Weighted summation to get the global context vector graph-level embedding The resulting global context vector graph-level embedding That is graph G i The graph-level embedding vector at the l-th encoder layer.

7. The method for calculating graph similarity based on position-structure graph neural network according to claim 5, characterized in that: The step (2) is specifically: In the difference attention layer, the graph-level embedding vectors of the same level of graph G1 and graph G2 are first calculated The absolute difference of , and this representation is input into a multilayer perceptron to fine-tune and enhance the difference information, namely: These differences are then converted into attention weight vectors using the softmax(·) function: Among them, α (l) Each element in the vector represents the weight assigned to each dimension of the graph-level embedding, and the temperature factor t regulates the smoothness of the difference; Next, the weight vector α (l) Respectively with graph-level embedding and Perform element-wise multiplication to obtain the difference-enhanced graph-level embedding and 8. The method for calculating graph similarity based on position-structure graph neural network according to claim 5, characterized in that: The step (3) is specifically: G i The difference enhanced graph-level embeddings are concatenated into a matrix in Representation graph G i For the difference-enhanced graph-level embedding vector corresponding to the encoder at the lth layer, the process of allocating weights using the multi-head self-attention mechanism is shown in the following mathematical formula: MultiHead(Q,K,V)=Concat(head1,…,head h )W O IN' i =LayerNorm(U i +MultiHead(Q,K,V)) head j =Attention(Q j ,K j ,V j ),in They represent the query matrix, key matrix, and value matrix of the j-th head respectively. is a trainable parameter; h is the number of heads to be determined by the model; W O is a trainable parameter; After assigning weights through the multi-head self-attention mechanism, the resulting matrix Each row in is the weighted graph-level embedding of each level. By adding up the rows in , we can get the fused comprehensive graph-level embedding Right now 9. The method for calculating graph similarity based on position-structure graph neural network according to claim 8, characterized in that: The score regression module in step 3 uses a multi-layer perceptron as a regression head to map the extracted comprehensive graph-level embedding to a scalar value as the similarity score of the input graph pair; That is, the comprehensive graph-level embedding of graph G1 and graph G2 and After concatenation, a multi-layer perceptron is used to map it to the similarity score of the image pair, that is, 10. A graph similarity calculation system based on a position-structure graph neural network for implementing the method according to any one of claims 1 to 9, characterized in that: Includes a local-global encoder, a weighted feature fusion module, and a score regression module; The local-global encoder module aggregates local and global information of the graph based on the structure and features of the input graph to generate a set of node embeddings with different levels of information; The weight-based feature fusion module captures the difference information between graphs through the interaction between graph pairs and generates graph-level embedding; The score regression module predicts the similarity score of the input graph pair based on the graph-level embedding information.