Graph data similarity measurement method and device, medium and equipment
By defining graph similarity in complex space and combining multi-level joint graph embedding network and cored Softmax operator, the problems of high time complexity and missing fine-grained difference information in graph similarity calculation are solved, and efficient and accurate graph similarity prediction is achieved.
Patent Information
- Application Number
- CN202510341398.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
The existing graph similarity calculation methods have problems with high time complexity and lack of fine-grained difference information when calculating the similarity between large graphs. Traditional methods cannot calculate accurate values within a reasonable time, and the calculation overhead of existing graph neural network methods is too high when processing complex graph structures.
The complex space rotation module and multi-fraction fusion mechanism are adopted to build a parallel node embedding network and neural tensor network, and combined with the complex space rotation module, the graph data is mapped to complex vector space, and the rotation similarity between graphs is defined using Euler's identity, and the multi-level joint graph embedding network and the nucleated Softmax operator reduces the algorithm complexity.
It improves the accuracy and efficiency of graph similarity calculation, can capture subtle differences between graphs more accurately, reduces the computation time complexity, and significantly improves the accuracy of similarity prediction.
Smart Images

Figure CN120277425A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graph data processing, and in particular, to a method, apparatus, medium, and device for measuring the similarity of graph data. Background Art
[0002] A graph is the object of study in graph theory and is a commonly used data structure in computer science. Algorithms related to graphs are also fundamental algorithms in computer science. The graph model can represent general and complex data structures, which enables the graph model to be widely applied in computational problems in many different fields, such as social network analysis, bioinformatics, and recommendation systems. In these problems, there are a large number of problems that require descriptions of the similarity degrees between different objects. Therefore, graph similarity calculation has become a core step in many graph-related machine learning tasks. These tasks include brain data analysis, network traffic detection, and user behavior recognition. In this context, graph edit distance and maximum common subgraph have become widely used graph similarity measurement methods. However, calculating the exact GED and MCS of two graphs is NP-hard and requires exponential time complexity in the worst case. Currently, there is no algorithm that can calculate the exact similarity value between large graphs within a reasonable time.
[0003] In recent years, graph neural networks, as a deep learning method based on graph structures, have become a powerful class of deep learning models for learning node representations in graph data. The core idea of the GNN model is to encode graph information into vector form and perform end-to-end learning of node representations through a neural network. Then, the learned node representations can be applied to various downstream tasks, such as directly used for node classification or aggregated into graph-level vectors for graph classification. Graph neural networks can effectively process complex graph tasks in an end-to-end manner, bringing new ideas to traditional similarity measurement methods. By using GNN to learn the similarity score between two graphs, it can be further divided into two categories: an embedding-based graph similarity calculation model and a matching-based graph similarity calculation model.
[0004] The embedding-based graph similarity calculation model embeds the entire graph as a graph-level vector and then calculates the similarity between the vectors as the similarity between the two graphs. The limitation of this method is that it does not pay attention to more fine-grained difference information, resulting in low prediction accuracy. Summary of the Invention
[0005] Based on this, in order to solve the technical problems in the prior art, the present invention provides a method, apparatus, medium, and device for measuring the similarity of graph data.
[0006] The present invention provides a method for measuring the similarity of graph data, including:
[0007] Construct a neural network, including: a first node embedding network module and a second node embedding network module in parallel, a neural tensor network module connected to the output ends of the first node embedding network module and the second node embedding network module simultaneously, a complex space rotation module connected to the output ends of the first node embedding network module, the second node embedding network module and the neural tensor network module simultaneously, and a similarity calculation module connected to the output ends of the neural tensor network module and the complex space rotation module;
[0008] Collect graph data to construct a data set, and use the data set to train the neural network to obtain a metric model for measuring the similarity of two graph data;
[0009] Input the first graph data and the second graph data to be measured into the metric model. Respectively perform embedding operations on the first graph data and the second graph data through the parallel first node embedding network module and second node embedding network module to obtain a first global feature and a second global feature; match the first global feature and the second global feature through the neural tensor network module to obtain a pseudo-similarity score of the first global feature and the second global feature; map the first global feature and the second global feature to the complex vector space through the complex space rotation module, and use the pseudo-similarity score as the rotation relationship between the first global feature and the second global feature in the complex vector space to calculate the similarity of the first global feature and the second global feature to obtain a rotation similarity score; fuse the pseudo-similarity score and the rotation similarity score through the similarity calculation module and map them to the similarity score of the first graph data and the second graph data.
[0010] Furthermore, the matching of the first global feature and the second global feature to obtain the pseudo-similarity score of the first global feature and the second global feature is achieved through the following formula:
[0011]
[0012] where f p (·) is the pseudo-similarity score; is the first global feature; is the second global feature; the superscript T represents the transpose; W, V and b are learnable parameters in the neural tensor network, and K is a hyperparameter; is the Hadmard product.
[0013] Furthermore, the mapping of the first global feature and the second global feature to the complex vector space, and using the pseudo-similarity score as the rotation relationship between the first global feature and the second global feature in the complex vector space to calculate the similarity of the first global feature and the second global feature to obtain the rotation similarity score is achieved through the following formula:
[0014]
[0015] Among them, and are the first global feature, the second global feature, and the pseudo-similarity score projected into the complex vector space respectively; f s (·) is the rotational similarity score.
[0016] Furthermore, the embedding operation on the first graph data and the second graph data to obtain the first global feature and the second global feature specifically includes:
[0017] Performing node embedding operations on the first graph data and the second graph data based on the Graph Transformer network introducing the self-attention mechanism and the kernelized Softmax operator, and respectively generating the first node embedding and the second node embedding where N and M are the numbers of nodes of the first graph data and the second graph data respectively, and D is the node feature dimension; and are the node features of node i of the first graph data and node j of the second graph data respectively;
[0018] Aggregating the first node embedding and the second node embedding after the graph pooling operation into the first graph-level embedding and the second graph-level embedding respectively through the global addition aggregation function; obtaining the first global feature and the second global feature by performing multi-level feature extraction on the first graph-level embedding and the second graph-level embedding.
[0019] Furthermore, the graph pooling operation specifically includes:
[0020] Projecting the first node embedding H 1 and the second node embedding H 2 onto one dimension to obtain the corresponding projection scalars y 1 and y 2 , and respectively returning the largest k node indices in y 1 and y 2 by using the node sorting operation, and forming the pooled first node embedding and the second node embedding based on the returned node indices.
[0021] Furthermore, mapping the fused pseudo-similarity score and rotational similarity score to the similarity score of the first graph data and the second graph data specifically includes:
[0022] Fusing the pseudo-similarity score and the rotational similarity score to obtain the fused similarity score H s :
[0023]
[0024] Mapping the similarity score H through a two-layer feedforward network and the sigmoid functions Mapped to a similarity score P out :
[0025] P out = σ(MLP(H S ))
[0026] Where σ(·) is the sigmoid function; MLP(·) is a multi-layer perceptron feed-forward network.
[0027] The present invention provides a similarity measurement device for graph data, including:
[0028] A model construction module for constructing a neural network, including: a parallel first node embedding network module and a second node embedding network module, a neural tensor network module connected to the output ends of the first node embedding network module and the second node embedding network module at the same time, a complex space rotation module connected to the output ends of the first node embedding network module, the second node embedding network module and the neural tensor network module at the same time, and a similarity calculation module connected to the output ends of the neural tensor network module and the complex space rotation module;
[0029] A model training module for collecting graph data to construct a data set, and using the data set to train the neural network to obtain a measurement model for measuring the similarity of two graph data;
[0030] A similarity calculation module for inputting the first graph data and the second graph data to be measured into the measurement model, performing embedding operations on the first graph data and the second graph data respectively through the parallel first node embedding network module and the second node embedding network module to obtain a first global feature and a second global feature; matching the first global feature and the second global feature through the neural tensor network module to obtain a pseudo-similarity score of the first global feature and the second global feature; mapping the first global feature and the second global feature to the complex vector space through the complex space rotation module, and calculating the similarity of the first global feature and the second global feature with the pseudo-similarity score as the rotation relationship between the first global feature and the second global feature in the complex vector space to obtain a rotation similarity score; and mapping the similarity score after fusing the pseudo-similarity score and the rotation similarity score to the similarity score of the first graph data and the second graph data through the similarity calculation module.
[0031] The present invention provides a computer-readable storage medium, where the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned similarity measurement method for graph data is implemented.
[0032] The present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned similarity measurement method for graph data is implemented.
[0033] The above at least one technical solution adopted by the present invention can achieve the following beneficial effects:
[0034] In the similarity measurement method of graph data provided by the present invention, by introducing a complex space rotation module and a multi-fraction fusion mechanism, the problem of missing fine-grained difference information in traditional graph-level vector similarity calculation is effectively solved. Specifically, after mapping the global features to complex vectors by the complex space rotation module, the local structural differences between nodes (such as small changes in edge weights and subgraph topologies) are captured through rotation operations, thereby enhancing the sensitivity to subtle differences between graphs. At the same time, the fusion of the pseudo-similarity scores generated by the neural tensor network and the complex space rotation similarity scores can comprehensively consider the dual perspectives of global feature matching and local structure alignment, avoiding the over-compression of complex structure information by a single graph-level vector. Finally, through the collaborative optimization of multi-dimensional features, the accuracy of similarity prediction is significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0036] Figure 1 is a schematic diagram of the overall model architecture of SRGSim provided by the present invention;
[0037] Figure 2 is a schematic diagram of the overall model architecture of GTPool provided by the present invention;
[0038] Figure 3 is a schematic diagram of the comparison of running times provided by the present invention, Figure 3 in which (a) is Running time on IMDBand LINUX, Figure 3 in which (b) is Running time on [3,200], Figure 3 in which (c) is Running time on [20,200], Figure 3 in which (d) is Running time on [50,200];
[0039] Figure 4 is a schematic diagram of the influence of the rotation operation provided by the present invention;
[0040] Figure 5 is a schematic diagram of the influence of the kernelized Softmax operator provided by the present invention;
[0041] Figure 6 is a schematic diagram of the sensitivity analysis of hyperparameters provided by the present invention, Figure 6In (a), it is the graph pooling ratio r. Figure 6 In (b), it is the tensor slices dimension K. Detailed implementation manners
[0042] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0043] The graph similarity calculation model based on embedding embeds the entire graph into a graph-level vector, and then calculates the similarity between the vectors as the similarity between two graphs. The limitation of this method is that it does not pay attention to more fine-grained difference information, resulting in low prediction accuracy. The graph similarity calculation model based on matching embeds each node into a low-dimensional vector and measures the similarity between two graphs through node-level interaction or graph-level interaction. The limitation of this method is that the node comparison process has at least a quadratic time complexity, thus consuming a large amount of time. Therefore, to design an efficient framework for learning the similarity between graphs of different scales, it is necessary to consider how to process different graph structures and how to establish meaningful comparisons between different scales. This is a challenging and attractive task.
[0044] To address the above challenges, a new method called Scalable Rotational Graph Similarity Calculation (abbreviated as SRGSim) is proposed in the present invention to jointly learn graph representations and graph similarity metric functions in an end-to-end manner. The motivation comes from Euler's identity e iθ = cosθ + isinθ, which indicates that a unitary complex number can be regarded as a rotation on the complex plane. Specifically, by mapping graph representations to the complex vector space and defining the similarity between graphs as the rotation between two graphs, the similarity between an arbitrary pair of structured graphs can be expressed more accurately. Before that, a multi-level joint graph embedding network GTPool was designed to obtain graph features from multiple perspectives using a multi-view module. That is, GTPool maps each graph to an embedding vector through a Graph Transformer network and a GraphPooling operation. At the same time, a multi-level joint module is proposed to concatenate the outputs of different GTPool network layers, and the Softmax aggregation operator is used to aggregate the graph-level embeddings to provide a global summary of the graph.
[0045] Although this method provides a novel idea for solving the graph similarity calculation problem, it still needs to solve two problems.
[0046] Problem 1: How to reduce the algorithmic complexity of the Graph Transformer. The Graph Transformer network needs to pass the information of each node to its neighbor nodes for message aggregation. This results in a computational overhead that grows linearly with the number of nodes. For graph similarity calculation, this significantly increases the calculation time. The key to solving the problem lies in reducing the algorithmic complexity of the Graph Transformer without sacrificing accuracy.
[0047] Problem 2: How to define the similarity score between graphs in the complex number space. Existing graph similarity calculation methods usually measure the similarity between two graphs through node-level interaction or graph-level interaction, but there are limitations in capturing the complex relationships of graphs. By defining the similarity score in the complex number space, these limitations can be overcome. The key to solving the problem lies in how to map the representation of graphs to the complex vector space, define the similarity score between graphs, and perform similarity measurement in this domain.
[0048] Based on this, the following two strategies are proposed to solve the above problems. To solve Limitation 1, the algorithmic complexity of the Graph Transformer is reduced to linear by a kernelized Gumbel-Softmax operator, which seamlessly integrates random feature mapping for extracting the latent structure in all instance nodes. To define the similarity score between graphs in the complex number space, i.e., to solve Limitation 2, a neural tensor network suitable for inferring the relationship between two graphs is introduced, which explicitly correlates the pseudo-similarity between two graphs. Specifically, given two graphs G1 and G2, the learned graph representation is mapped to the complex number space, and the pseudo-similarity score is defined as the rotation from G1 to G2.
[0049] Embodiment 1
[0050] The following details the method for measuring the similarity of graph data in this embodiment, which specifically includes the following steps:
[0051] S1: Construct a neural network, including: a parallel first node embedding network module and second node embedding network module, a neural tensor network module connected to the output ends of the first node embedding network module and the second node embedding network module at the same time, a complex number space rotation module connected to the output ends of the first node embedding network module, the second node embedding network module, and the neural tensor network module at the same time, and a similarity calculation module connected to the output ends of the neural tensor network module and the complex number space rotation module at the same time.
[0052] Graph similarity calculation aims to measure the similarity between two graphs to reveal the structural relationships and feature similarities between them. Traditional graph similarity calculation seeks appropriate distance metrics to reflect the similarity between two graphs. For example, the graph edit distance takes the minimum cost of the edit operations required to transform the input graph into the target graph as the distance between the graphs. Compared with the graph edit distance, the advantage of the distance metric based on the maximum common subgraph is that it does not require defining edit operations and their costs. However, traditional graph similarity calculation mainly relies on the original graph architecture and is learning-free. Different from traditional graph similarity calculation methods, graph kernel similarity calculation represents the entire graph structure as a vector containing the number of its basic substructures and inserts it into the Hilbert space using the inner product. But the graph kernel method has a high time complexity and depends on substructures, so it cannot efficiently process graphs with a large number of nodes.
[0053] In recent years, with the development of graph neural networks, it has provided new opportunities for graph pair similarity calculation. SimGNN first regarded graph similarity learning as a regression task, using histogram features and neural tensor networks to model node-level and graph-level interactions; GraphSim was extended based on SimGNN, directly matching two sets of node embeddings, generating vector representations for the nodes of the two graphs with different layers of GCNs to capture complex node-level interactions; MGMN found that recent research on GSL mainly considered graph-to-graph (G2G) or node-to-node (N2N) interactions and ignored cross-layer interactions. To solve this problem, MGMN added a node-graph cross-level interaction layer that can compare each node with all the nodes of another graph one by one; H2MN learned the representation of graphs from the perspective of hypergraphs and matched each hyperedge as a subgraph; PSimGNN obtained subgraphs through graph partitioning and selected Top-k from them for comparison. But these methods adopt a cross-graph mechanism for graph similarity learning, which is computationally expensive. Considering these limitations, CGSim utilizes the complementary information of the two input graphs to capture pairwise relationships in the contrastive learning framework. But these methods adopt a cross-graph mechanism for graph similarity learning, which is computationally expensive. Considering these limitations, CGSim utilizes the complementary information of the two input graphs to capture pairwise relationships in the contrastive learning framework.
[0054] The proposed SRGSim effectively calculates the similarity between graphs while maintaining high accuracy. However, it differs from existing work in three aspects: 1) Different from methods that consider graph-level or node-level interactions, it captures different levels of features between two graphs in the complex vector space for the first time, showing better performance and efficiency than existing work. 2) A multi-level joint graph embedding network GTPool is designed, which combines the advantages of the Graph Transformer network and the Graph Pooling operation to capture the global summary of the graph from multiple perspectives; 3) The algorithm complexity of the Graph Transformer is reduced from quadratic to linear through the kernelized Gumbel-Softmax operator, greatly reducing the running time cost.
[0055] A brief introduction to the problem formulation is given. Given a pair of input graphs (G1, G2) ∈ G, the purpose of graph similarity learning is to learn the similarity score y = s(G1, G2) ∈ Y of these two graphs. Among them, G1 = {V1, E1, X1}, where V1 is the set of nodes; E1 is the set of edges, used to represent the relationship between nodes; X1 is the node feature X1 = R N×D . Similarly, G2 = {V2, E2, X2}, where V2 is the set of nodes; E2 is the set of edges, used to represent the relationship between nodes; X2 is the node feature X2 = R M×D . It should be noted that for the graph-graph classification task, the similarity score y is the classification label, that is, y ∈ Y = {0, 1}; for the graph-graph regression task, the similarity score y is a measure of the similarity of the graph pair, that is, y ∈ Y = [0, 1]. Taking any two graphs as input, and then outputting a score representing their similarity through a calculation and learning process. The output similarity score can not only be used to judge whether they belong to the same category or have similar characteristics, but also be used for various downstream tasks related to graph similarity in the future. The important symbols and notations used in this invention are shown in Table 1.
[0056] Table 1 Symbol Summary
[0057]
[0058] The proposed method SRGSim is introduced in detail. This is an end-to-end neural network-based method. The overall model architecture of SRGSim is as Figure 1As shown in the figure, the model uses the GTPool network to convert the entire graph into a vector representation to capture the information in the graph. Then, two strategies are proposed to model the similarity between two graphs. One is to use the neural tensor network to model the similarity between graphs; the other is to map the graph embedding into the complex space and define the similarity score learned by the tensor neural network as the rotation between graph pairs in the complex vector space to express the similarity between graphs. Finally, the two strategies are combined and input into the fully connected neural network to obtain the final similarity score. In the following subsections, the multi-level joint graph embedding network GTPool is first introduced, and then how to use the neural tensor network to model the similarity between two graphs in the complex vector space is introduced.
[0059] GTPool: Graph representation learning is the process of representing nodes in graph data as a set of low-dimensional vectors. Graph representation learning plays a key role in learning the graph similarity between two graphs. The Graph Transformer network architecture is adopted to capture the topological structure, attribute information, and high-order feature information of the graph. However, as the number of nodes increases, the computational overhead of the Graph Transformer network grows linearly. To this end, the algorithm complexity of the Graph Transformer is reduced to linear through the kernelized Softmax operator. Next, the graph pooling operation is introduced to coarsen the graph, reduce the parameters of SRGSim, and capture the hierarchical structure information of the graph. Finally, in order to obtain the features of the graph from multiple perspectives, the Softmax aggregation operator is used to aggregate the information output by different GTPool network layers. The GTPool consists of three parts: 1) Node embedding layer; 2) Graph pooling layer; 3) Graph-level embedding aggregation layer. The overall model architecture of GTPool is as Figure 2 shown.
[0060] Node embedding layer: In the node embedding layer, the Graph Transformer network architecture is considered, and the self-attention mechanism is applied to graph learning to generate node embeddings for the nodes in G1 and G2 and Specifically, given the node features, the attention coefficient of each edge from j to i in graph G is calculated:
[0061]
[0062] and are learnable parameters, which perform linear transformation operations on the source node feature x i and the target node feature x j and then calculate the attention coefficient α ij .
[0063] After obtaining the graph attention coefficients, message aggregation is performed from the target node j to the source node i:
[0064]
[0065] Here, W i (k) represents a learnable parameter matrix. It is worth noting that the Graph Transformer network architecture shares parameters when training on input graphs G1 and G2, and the required number of layers depends on a specific practical application.
[0066] Graph pooling layer: To make full use of the node feature information and capture the complex information inside the graph, inspired by the graph pooling operations in recent years, it is considered to apply it to the graph representation learning model. This will help improve the performance and generalization ability of the model. Specifically, a trainable projection vector p is adopted. By projecting all node features onto one dimension, the scalar projection y of X on p is obtained. Then, the node sorting operation rank(·) is used to return the indices idx of the top k nodes in y. The key step of this operation is to retain the indices of the top-ranked nodes according to importance to form the feature matrix X′ and adjacency matrix A′ of the new graph. The selection rule for the graph pooling operation is:
[0067]
[0068] idx = rank(y, k)
[0069] X′ = (X tanh(y)) idx
[0070] A′ = A idx,idx
[0071] where X = {H 1 , H 2}.
[0072] Graph embedding layer: After calculating the node embeddings H of each graph using the graph pooling operation, these node embeddings need to be aggregated to form the graph-level embedding H (G) corresponding to each graph. In this step, different aggregation functions can be selected to better capture the global features of the graph, such as global sum aggregation, global average aggregation, and global max aggregation. It is worth noting that in global sum aggregation, global sum aggregation emphasizes the overall contribution of each node in the graph and is more sensitive to those nodes with obvious centrality in the graph. By adding the node embeddings, the focus is placed on the global structure of the graph, which is very beneficial for capturing the overall features of the graph. Therefore, in the model, global sum aggregation is selected as the aggregation function.
[0073]
[0074] To effectively aggregate the multi-level features of graphs, multiple layers of RGSim networks are stacked, where l represents the number of network layers. In these layers, the Softmax aggregation operation is used to integrate the graph embedding representations obtained from each layer, as follows.
[0075]
[0076] Here, t is a continuous variable, called the inverse temperature. The SoftMax function has been studied in many machine learning fields considering the temperature parameter t. When t is low, the SoftMax function behaves similarly to average aggregation, i.e., the probabilities of the elements in the output distribution tend to be equal. However, when t is high, the SoftMax function behaves more like max aggregation, i.e., the element with the highest probability in the output distribution dominates. It is worth noting that the goal is to dynamically learn the stable t of the SoftMax aggregation operator. This means that t is not a static hyperparameter but a parameter that is dynamically adjusted during the model training process to find the optimal aggregation strategy, balancing the characteristics of average aggregation and max aggregation. This dynamic adjustment mechanism can help the model better adapt to the characteristics of the data, thereby improving the performance and generalization ability of the model.
[0077] Rotation layer: Graph embeddings can better express complex relationships and features. Learning graph-level interactions can be an important supplement to the graph similarity between two graphs. To calculate the similarity between two graphs under a richer mathematical representation, the graph representation is mapped to the complex vector space, and then the pseudo-similarity score h is defined s as the rotation from G1 to G2. In other words, given ( H s ), it is expected that:
[0078]
[0079] where is the Hadmard product. Due to the rotation property of this model, it is called rotation similarity calculation. According to the above definition, the function for calculating the similarity between two graphs is defined as:
[0080]
[0081] where the superscript i is the Euler's formula e iθThe imaginary number in \(e^{i\theta}=\cos\theta + i\sin\theta\), where \(e\) is the base of the natural logarithm; by defining each similarity as a rotation in the complex vector space, RGS im can model and infer the graph similarity between two graphs. At the same time, the rotation from \(G2\) to \(G1\) is also considered. It is worth noting that to solve Problem 2, a neural tensor network is proposed to model the pseudo-similarity \(H\) between two graphs. s .
[0082]
[0083] where \(W\), \(V\), and \(b\) are learnable parameters. \(K\) is a hyperparameter that controls the number of interaction (similarity) scores generated by the model for each pair of graph embeddings.
[0084] Output layer: To predict the final graph similarity score, the similarity scores output by the rotation layer and the pseudo-similarity scores obtained from the neural tensor network are concatenated to obtain a more comprehensive and accurate graph similarity score.
[0085]
[0086] \(\sigma(\cdot)\) is the sigmoid function; \(MLP(\cdot)\) is a multi-layer perceptron feed-forward network. This concatenation operation can integrate the advantages of both methods, making full use of the spatial transformation of the rotation layer and the advanced feature extraction ability of the neural tensor network, thus providing more accurate results. Finally, using \(H\) S as the input of the prediction layer, the prediction layer consists of a two-layer feed-forward network and a sigmoid function to generate the probability distribution for molecular property prediction:
[0087] \(P\) out \(=\sigma(MLP(H\) S ))
[0088] \(MLP(\cdot)\) is a multi-layer perceptron feed-forward network, and \(\sigma\) represents the sigmoid function, which converts the output to a value between 0 and 1, representing the final probability distribution for molecular property prediction.
[0089] S2: Collect graph data to construct a dataset, and use the dataset to train the network model to obtain a metric model for calculating the similarity of two graph data.
[0090] The binary cross-entropy loss function is used to train the supervised model. Binary cross-entropy loss is a commonly used loss function for measuring the difference between the model's prediction results and the true labels. Its form is as follows:
[0091]
[0092] where \(L(\cdot)\) is the loss function, \(N\) is the number of samples, \(y\) iis the true label (0 or 1) of the i-th sample, is the predicted probability of the i-th sample.
[0093] Model summary.
[0094] The model adopts an end-to-end learning method to learn the graph representation and the graph similarity metric function to calculate the graph similarity. First, the algorithm complexity of the Graph Transformer is reduced from O(N 2 ) to O(N) through the kernelized Softmax operator, significantly reducing the time consumption of the node embedding layer. Then, a multi-level joint graph embedding network GTPool is proposed to obtain the global summary of the graph from multiple perspectives. Finally, the graph representation is mapped to the complex vector space, and the similarity score between two graphs is calculated from the rotation angle, and the pseudo-similarity score learned by the tensor neural network is defined as the rotation in the complex vector space. Subsequent experiments will demonstrate the superior performance of the proposed framework.
[0095] S22: Experiments.
[0096] S221: Datasets.
[0097] Classification datasets: For the graph-graph classification task of calculating the similarity score between two binary functions, a benchmark dataset generated by two popular open-source softwares FFmpeg and OpenSSL is used to evaluate the proposed model. For the FFmpeg and OpenSSL datasets, each graph represents the control flow graph of a binary function, where the graph nodes represent basic blocks (a basic block is a sequence without jump instructions), the edges represent the control flow paths between these basic blocks, and each node is initialized with 6 block-level digital features. Considering the impact of the graph size on the performance of the graph matching network, the dataset is further divided into three sub-datasets according to the size range of the input graph pairs (i.e., [3,200], [20,200], and [50,200]). For example, in FFmpeg [50,200], "50" represents the minimum CFG size and "200" represents the maximum CFG size.
[0098] It should be noted that although there are many benchmark datasets for general graph classification tasks, these datasets cannot be directly used for the graph classification task because two input graphs with the same label cannot be regarded as "similar". This is because the general graph classification task only assigns one label to each graph, while the graph classification task learns the binary similarity label (i.e., similar or not similar) for two graph pairs rather than one graph.
[0099] Clustering Datasets: For the graph-graph clustering task of learning two graph distances, the models were evaluated on two benchmark datasets, LINUX1000 and IMDB1500. The LINUX dataset contains 48,747 program dependency graphs generated by the Linux kernel, where the nodes are statements, the edges are dependencies between statements, and the nodes are unlabeled. For the IMDB dataset, there is an edge if two people appear in the same movie. Each dataset contains a set of input graphs and their ground-truth GED scores, which are calculated by the A* algorithm. To convert the ground-truth GED scores into similarity scores, the same method as SimGNN was adopted to normalize them to the range (0, 1].
[0100] For fair comparison, the same training / validation / test split as in previous work was followed. The detailed statistics of these datasets can be found in Table 2.
[0101] Table 2 Dataset Statistics for Classification and Clustering Tasks
[0102]
[0103] S222: Baselines
[0104] The SRGSim was compared with the latest graph neural network baselines. This set of methods includes SimGNN, GraphSim, GMN, MGMN, H2MN, and CGSim. These methods were compared in graph-graph classification and regression tasks. Notably, all experiments were repeated ten times, and the average and standard deviation of the experimental results were reported, with the best ones in bold.
[0105] S223: Evaluation Metrics.
[0106] To comprehensively evaluate the performance of the models on the graph-graph regression task, five metrics were adopted for fair comparison. Mean-Square Error (MSE), which measures the mean variance between the predicted similarity and the ground truth similarity; Spearman's rank correlation coefficient (ρ) and Kendall's rank correlation coefficient (τ) are used to evaluate the rank correlation between the predicted results and the true ranking results; Precision@k is used to evaluate the accuracy of the model in the top-k results, where k = 10, 20. For the graph-graph classification task, the Area Under the Curve (AUC) was used to measure the classification performance of the model. For MSE, the smaller the better; while for ρ, τ, p@k, and AUC, the larger the better.
[0107] S224: Experimental Settings.
[0108] Build RGSim using PyTorch and PyTorch Geometric, and adopt the Adam optimizer to optimize the model parameters. In the graph embedding layer, use four layers of GTPool. Set the output dimension of each GTPool layer to 48, and run 400 iterations with a learning rate of 5e-4 to train the model. The MLP consists of three fully connected layers, uses ReLU as the activation function, and adds a dropout layer after each layer. During the training process, adopt the early stopping criterion for model optimization, that is, if the validation loss does not decrease for 10 consecutive epochs, stop the training, and select the best model according to the lowest validation loss. It should be noted that all experiments were conducted on a machine equipped with an NVIDIA GeForce MX450 GPU and an Intel i5-11300H 3.1GHz CPU.
[0109] S225: Performance of the model
[0110] S2251: Graph classification.
[0111] For the graph-graph classification task of detecting whether two binary functions are similar, evaluate the performance of each model by measuring the mean and standard deviation of the area under the ROC curve (AUC). It can be clearly seen from Table 3 that RGSim significantly achieves state-of-the-art performance on all six sub-datasets of the FFmpeg and OpenSSL datasets. Especially when the graph scale increases, RGSim can still maintain good performance and robustness compared to other state-of-the-art methods. The experimental results show that the introduction of the complex vector space can more accurately describe the similarity, enabling the model to calculate the similarity between two graphs under a richer mathematical representation.
[0112] Table 3 Experimental performance in graph classification
[0113]
[0114] S2251: Graph clustering.
[0115] Table 4 Experimental performance in graph clustering
[0116]
[0117]
[0118] For the regression task of calculating the graph edit distance between two graphs, the mean squared error (MSE), Spearman's rank correlation coefficient (ρ), Kendall's rank correlation coefficient (τ), and precision at k (p@k) are used to evaluate the model. This is the same as previous work. Table 4 summarizes all the results on the AIDS700 and LINUX1000 datasets. Generally speaking, similar conclusions can be drawn as in the graph-graph classification task. From all the evaluation metrics, the model performs significantly better than the SimGNN, GMN, and GraphSim baseline models on the AIDS700 and LINUX1000 datasets. On the other hand, compared with MGMN, RGSim obtains better results than MGMN. The results emphasize the importance of defining the similarity score between graphs in the complex space, and introducing the complex space can more accurately describe the graph pair similarity.
[0119] To further verify the efficiency of the proposed SRGSim, the running times between different methods are shown in Figure 3 . It is worth noting that in the regression task, SRGSim shows obvious performance advantages. Compared with the state-of-the-art graph neural network baselines, SRGSim runs approximately 2 to 5 times faster on all datasets. For the graph-graph classification task, the best-performing baselines MGMN and CGSim are selected for comparison, and the running times for training 150 epochs are recorded. In addition, CGSim is the baseline with the lowest time complexity among the GSL baselines. Figure 3 It shows that SRGSim consumes much less time in GSL. Compared with CGSim, the speed of RGSim is generally improved by 2 to 4 times.
[0120] It is worth mentioning that the excellence of SRGSim's running time stems from two important innovations in the model architecture. By avoiding the traditional node-to-node interaction matching mechanism and defining the similarity score between graphs in the complex space, a more efficient analysis of graph similarity is achieved. This key method not only improves the running time but also maintains the quality of the results, thus achieving a balance between efficiency and effectiveness. In addition, the algorithm complexity of GraphTransformer is reduced to linear by the kernelized Softmax operator, significantly reducing the computational time. The following ablation experiments will analyze the impact of the rotation operation and the kernelized operator on the model.
[0121] S226: Ablation study.
[0122] The graph similarity learning model SRGSim proposed in this invention consists of three main steps, namely the calculation of rotation graph similarity, the kernelized Softmax operator, and the multi-level joint graph embedding network GTPool. To verify the effectiveness of the rotation graph similarity calculation, the kernelized Softmax operator, and the multi-level joint graph embedding network GTPool in the model, ablation experiments will be designed in this section to replace the calculation method in the original steps of ERMat and conduct experiments with the graph classification task as the goal. The ablation experiment of ERMat will be introduced in detail below.
[0123] S2261: Investigate the impact of rotation operations
[0124] In the first section, it was proposed that modeling similarity as rotation in the complex vector space can calculate the similarity between two graphs under a richer mathematical representation. For this reason, the impact of this rotation operation on ERMat is further verified, and the Figure 4 experimental results are given. Specifically, on the premise of the same experimental settings, the rotation-based graph similarity calculation is replaced with the traditional real vector space to model the similarity between two graphs. This variant form of ERMat is denoted as w / o-Mat to verify whether more accurate similarity calculation between two graphs can be obtained by using the rotation operation.
[0125] It can be seen from Figure 4 that mapping the graph representation to the complex vector space can more accurately measure the similarity between graph pairs and perform well in various situations. An interesting observation is that in the case where the dataset has a large number of nodes (such as OpenSSL[50,200] and FFmpeg[50,200]), the performance of the complex vector space modeling is better than that of the real vector space modeling. This shows that the model can model similarity as a rotation operation in the complex vector space to improve the quality of graph similarity learning, especially when dealing with graph pairs with a large number of nodes.
[0126] S2262: Investigate the impact of the kernelized Softmax operator
[0127] To further study the impact of the kernelized Softmax operator on the model performance and efficiency, two variants are considered: 1) "w / o-ks" replaces Softmax with Gumbel-Softmax (Gumbel-Softmax is used by NodeFormer for node classification tasks in large networks); 2) "traditional-soft" uses the traditional Softmax operator.
[0128] An important phenomenon was observed that when the Gumbel component was added, the performance of the model decreased instead. This may be because the Gumbel component was introduced into the model to solve the vanishing gradient problem that occurred when calculating node attention in large datasets (up to 2M nodes). However, in the work, the datasets processed contain at most 200 nodes, which is a relatively small scale and far from sufficient to trigger the vanishing gradient problem compared with large datasets. Therefore, when the Gumbel component was added, it introduced a certain degree of randomness, and on small-scale datasets, this randomness may have an adverse effect, making it more difficult for the model to generalize to unseen data. The kernelized softmax operator can reduce the computational consumption without sacrificing accuracy. The schematic diagram of the impact of the kernelized Softmax operator is as Figure 5 shown.
[0129] S2262: Investigating the Impact of GTPool
[0130] Finally, the impact of the node embedding layer on classification and regression tasks in the GTPool model was investigated. With the same experimental settings as before, the Graph Transformer (GT) was replaced with four different node embedding methods, namely GCN, SAGE, GIN, and GGNN. Note that their output dimensions were kept consistent with GTPool (i.e., 48) in the experiment, and no hyperparameters of the four GNN models were fine-tuned.
[0131] Table 5 presents the classification task results of different GNN modules in SRGSim. According to the experimental results, it was found that in the ERMat model, the performance impact of various node embedding methods on the classification task was relatively small, and the fluctuation range of the experimental results was only between 1% and 3%. The GTPool model was not very sensitive to the choice of GNN models in the node embedding layer, but choosing GraphTransformer as the node embedding method still had a certain superiority. Generally speaking, in the SRGSim model, the choice of node embedding method had a small impact on the performance of the classification task, and the GTPool model was to some extent robust and could maintain stable performance under different node embedding methods.
[0132] S227: Parameter Sensitivity Analysis
[0133] In this study, it was explored how the number of layers, representation dimension, and graph pooling ratio of the neural network affected the performance of SRGSim in graph classification and graph regression tasks. The experimental settings were kept consistent, and only the number of layers, representation dimension, and graph pooling ratio of GTPool in the node embedding layer of SRGSim were adjusted.
[0134] S2271: Impact of the Number of Layers on the Performance of SRGSim
[0135] Specifically, the number of layers of the GTPool network was changed from 1, 2, 3, 4 to 5, and the experimental results of the graph-to-graph classification task were summarized in Table VIII. It can be seen from Table VIII that for FFmpeg and OpenSSL, ERMat with a larger number of GTPool layers (i.e., 4 and 5 layers) has better and relatively stable performance on all sub-datasets, while ERMat with a smaller number of GTPool layers (i.e., 1 or 2 layers) has poor performance on some sub-datasets. For example, ERMat with 1 layer performs very poorly on the [20,200] and [50,200] sub-datasets of the two datasets; ERMat with 2 layers performs poorly on the [20,200] and [50,200] sub-datasets of FFmpeg and the [50,200] sub-dataset of OpenSSL.
[0136] These observations indicate that the number of GTPool layers required by ERMat depends on different datasets or different tasks. Therefore, in order to avoid over-tuning this hyperparameter (i.e., the number of GTPool layers) on different datasets and tasks, and considering resource consumption, four-layer GCN is used as the default value in the node embedding layer of GTPool.
[0137] S2272: Influence of Embedding Dimension on SRGSim Performance
[0138] In this study, the influence of the GTPool network in ERMat on performance under different numbers of perspectives was explored. While keeping other experimental settings unchanged, only the number of perspectives of GTPool was adjusted (i.e., 16, 32, 48, 64, 80, 96). It can be clearly seen from Table 10 that for the graph-to-graph classification task, when the number of perspectives increases from 16 to 64, the AUC score of GTPool does not increase accordingly. However, when the number of perspectives exceeds 64, the AUC score of GTPool decreases as the number of perspectives increases. Similar results for the graph-to-graph regression task can be seen in Table IX of the appendix. Therefore, it is concluded that the performance of ERMat is not sensitive to the number of perspectives (from 16 to 64), and the dimension is default set to 48.
[0139] Table 7 Influence of Different Embedding Dimensions on Classification Performance
[0140]
[0141] S2273: Influence of Pooling Ratio on SRGSim Performance
[0142] In this study, how the graph pooling ratio affects the performance of ERMat in graph classification and graph regression tasks was explored. As Figure 6As shown in (a), when the graph pooling ratio increases from 0.80 to 1, the performance increases steadily. However, when the graph pooling ratio is less than 0.8, ERMat performs poorly on the dataset. This indicates that the pooling ratio r plays a crucial role in the model performance. Too small an r value will destroy the graph structure and lead to performance degradation. It is worth noting that even without using the pooling mechanism (i.e., r = 1.0), the performance will still be affected.
[0143] S2274: Impact of Tensor Slice Size on SRGSim Performance
[0144] During the graph similarity modeling process, the neural tensor network is used to directly relate two graph vectors across multiple dimensions, where K is the tensor slice (dimension size). To explore the impact of the tensor slice K on the ERMat performance at different dimensions, the dimension size of the tensor slice was adjusted (i.e., 8, 16, 32, 48, 64, 80). From Figure 6 (b), it can be clearly seen that for the graph-graph classification task (the results of the graph-graph regression task are shown in the appendix), when the dimension increases from 16 to 48, the AUC score of ERMat only has a slight fluctuation. Generally speaking, for all datasets of the two tasks, within the range of 16 to 48, the performance of different tensor slice dimension sizes is very similar. An interesting observation is that too small or too large tensor slices perform poorly in these two tasks.
[0145] S3: Input the first graph data and the second graph data to be measured into the measurement model. Through the parallel first node embedding network module and second node embedding network module, perform embedding operations on the first graph data and the second graph data respectively to obtain the first global feature and the second global feature; through the neural tensor network module, learn the interaction between the first global feature and the second global feature to obtain the pseudo-similarity score of the first global feature and the second global feature; through the complex space rotation module, map the first global feature and the second global feature to the complex space, and define the pseudo-similarity score as the rotation between graph pairs in the complex vector space to express the similarity between the first graph data and the second graph data, obtaining the rotation similarity score; through the similarity calculation module, map the pseudo-similarity score and the rotation similarity score to the similarity probability of the first graph data and the second graph data.
[0146] Based on the graph data similarity measurement method shown, graph similarity calculation is to establish a matching relationship of nodes and edges between two graph structures according to structural similarity. With its good flexibility and robustness, this technology plays a very important role in fields such as computer vision, industrial manufacturing, and bioinformatics. The present invention proposes a new efficient rotation graph similarity calculation method (SRGSim) for the efficiency problem faced by existing algorithms in processing cross-level interactions and the deficiency of the graph representation learning module in capturing complex graph structure information.
[0147] The core idea of SRGSim is to utilize Euler's identity to map graph embeddings into the complex vector space and define the similarity between graphs as the rotation between two graphs, thus avoiding the high-cost interactions at the node level and improving the computational efficiency. Specifically, first, a multi-level joint graph embedding network GTPool is designed using Graph Transformer and Graph Pooling. A multi-perspective module is introduced to encode different perspectives of the graph and fuse them with the original graph embedding. By combining information from different perspectives, ERMat can obtain a more comprehensive and accurate graph representation. In addition, the algorithmic complexity of Graph Transformer is reduced to linear through a kernelized Softmax operator, thus greatly improving the efficiency of graph similarity calculation. Finally, the graph representation learned by GTPool is mapped into the complex space, and the similarity between graphs is defined as the rotation between two graphs using the rotation property of Euler's identity, avoiding the high-cost calculation problems brought by existing algorithms when using cross-level interactions to learn the graph similarity between two graphs. However, it should be noted that the similarity between graph pairs is not clear in advance. To solve this problem, a neural tensor network is used to explicitly associate the pseudo-similarity between two graphs as rotation. Through comprehensive experiments on four widely used graph datasets, it is demonstrated that, compared with the state-of-the-art baseline models, ERMat significantly improves the accuracy and efficiency of graph similarity learning in classification and regression tasks.
[0148] In future work, it is planned to attempt to combine the model with emerging technologies such as multi-modal graph embedding and graph reinforcement learning. For example, by using multi-modal graph embedding technology, these different forms of information are embedded into the same low-dimensional space, and a reinforcement learning strategy is introduced to dynamically adjust the model parameters to adapt to different graph structures and similarity measurement tasks, providing more solutions for solving complex graph similarity problems.
[0149] The above is the similarity measurement method for graph data provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding similarity measurement device for graph data, including:
[0150] A model construction module for constructing a neural network, including: a parallel first node embedding network module and second node embedding network module, a neural tensor network module connected to the output ends of the first node embedding network module and the second node embedding network module simultaneously, a complex space rotation module connected to the output ends of the first node embedding network module, the second node embedding network module, and the neural tensor network module simultaneously, and a similarity calculation module connected to the output ends of the neural tensor network module and the complex space rotation module simultaneously.
[0151] A model training module, which is used to collect graph data to construct a data set, and use the data set to train a neural network to obtain a metric model for measuring the similarity between two graph data.
[0152] A similarity calculation module, which is used to input the first graph data and the second graph data to be measured into the metric model, and perform embedding operations on the first graph data and the second graph data respectively through a parallel first node embedding network module and a second node embedding network module to obtain a first global feature and a second global feature; match the first global feature and the second global feature through a neural tensor network module to obtain a pseudo-similarity score of the first global feature and the second global feature; map the first global feature and the second global feature to a complex vector space through a complex space rotation module, and use the pseudo-similarity score as the rotation relationship between the first global feature and the second global feature in the complex vector space to calculate the similarity between the first global feature and the second global feature to obtain a rotation similarity score; fuse the pseudo-similarity score and the rotation similarity score through the similarity calculation module and map them to the similarity score of the first graph data and the second graph data.
[0153] For the specific limitations of the graph data similarity measurement device, reference can be made to the limitations of the graph data similarity measurement method in the above text, which will not be elaborated here. Each module in the above graph data similarity measurement device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in the form of hardware or be independent of it, or be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0154] The present invention also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above-provided graph data similarity measurement method.
[0155] The present invention also provides the structure of a computer device. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 provided graph data similarity measurement method.
[0156] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0157] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded by the present invention.
Claims
1. A method for measuring the similarity of graph data, characterized in that, Comprising: Construct a neural network, including: a parallel first node embedding network module and second node embedding network module, a neural tensor network module connected to the output ends of the first node embedding network module and the second node embedding network module simultaneously, a complex space rotation module connected to the output ends of the first node embedding network module, the second node embedding network module, and the neural tensor network module simultaneously, and a similarity calculation module connected to the output ends of the neural tensor network module and the complex space rotation module; Collect graph data to construct a data set, and use the data set to train the neural network to obtain a metric model for measuring the similarity of two graph data; Input the first graph data and the second graph data to be measured into the metric model. Respectively perform embedding operations on the first graph data and the second graph data through the parallel first node embedding network module and second node embedding network module to obtain a first global feature and a second global feature; match the first global feature and the second global feature through the neural tensor network module to obtain a pseudo-similarity score of the first global feature and the second global feature; map the first global feature and the second global feature to the complex vector space through the complex space rotation module, and use the pseudo-similarity score as the rotation relationship between the first global feature and the second global feature in the complex vector space to calculate the similarity of the first global feature and the second global feature to obtain a rotation similarity score; Fuse the pseudo-similarity score and the rotation similarity score through the similarity calculation module and map them to the similarity score of the first graph data and the second graph data.
2. The similarity measurement method of graph data according to claim 1, characterized in that, The matching of the first global feature and the second global feature to obtain the pseudo-similarity score of the first global feature and the second global feature is achieved through the following formula: where, f p (·) is the pseudo-similarity score; is the first global feature; is the second global feature; the superscript T represents transpose; W, V, and b are learnable parameters in the neural tensor network, K is a hyperparameter; ο is the Hadmard product.
3. The similarity measurement method for graph data according to claim 2, wherein The mapping of the first global feature and the second global feature to the complex vector space, and using the pseudo-similarity score as the rotation relationship between the first global feature and the second global feature in the complex vector space to calculate the similarity of the first global feature and the second global feature to obtain the rotation similarity score is achieved through the following formula: Among them, and are the first global feature, the second global feature, and the pseudo-similarity score projected onto the complex vector space, respectively; f s (·) is the rotational similarity score.
4. The method for measuring the similarity of graph data according to claim 1, characterized in that The embedding operations on the first graph data and the second graph data to obtain the first global feature and the second global feature specifically include: The Graph Transformer network based on the introduction of self-attention mechanism and kernelized Softmax operator performs node embedding operations on the first graph data and the second graph data, and generates the first node embedding respectively and the second node embedding where N and M are the number of nodes of the first graph data and the second graph data respectively, and D is the node feature dimension; and are the node features of node i of the first graph data and node j of the second graph data respectively; Aggregate the first node embedding and the second node embedding after graph pooling operations into a first graph-level embedding and a second graph-level embedding respectively through a global addition aggregation function; obtain the first global feature and the second global feature through multi-level feature extraction of the first graph-level embedding and the second graph-level embedding.
5. The similarity measurement method of graph data according to claim 4, characterized in that The graph pooling operation specifically includes: Embed the first node into H respectively 1 and embed the second node into H 2 Project them into one dimension to obtain the corresponding projection scalars y 1 and y 2 , and use the node sorting operation to return y 1 and y 2 The indices of the k largest nodes in, and form the pooled first node embedding and second node embedding based on the returned node indices.
6. The method for measuring the similarity of graph data according to claim 3, characterized in that The fusion of the pseudo-similarity score and the rotation similarity score and mapping them to the similarity score of the first graph data and the second graph data specifically includes: Fuse the pseudo-similarity score and the rotation similarity score to obtain a fused similarity score H s : Map the similarity score H to the similarity score P through a two-layer feedforward network and the sigmoid function s out : P out = σ(MLP(H S )) Wherein, σ(·) is the sigmoid function; MLP(·) is a multi-layer perceptron feedforward network.
7. A similarity measurement device for graph data, characterized in that Comprising: A model construction module for constructing a neural network, including: a parallel first node embedding network module and a second node embedding network module, a neural tensor network module connected to the output ends of the first node embedding network module and the second node embedding network module simultaneously, a complex space rotation module connected to the output ends of the first node embedding network module, the second node embedding network module and the neural tensor network module simultaneously, and a similarity calculation module connected to the output ends of the neural tensor network module and the complex space rotation module; A model training module for collecting graph data to construct a data set, and using the data set to train the neural network to obtain a metric model for measuring the similarity of two graph data; The similarity calculation module is used to input the first graph data and the second graph data to be measured into the metric model, and perform embedding operations on the first graph data and the second graph data respectively through the parallel first node embedding network module and the second node embedding network module to obtain a first global feature and a second global feature; match the first global feature and the second global feature through the neural tensor network module to obtain a pseudo-similarity score of the first global feature and the second global feature; map the first global feature and the second global feature to the complex vector space through the complex space rotation module, and calculate the similarity of the first global feature and the second global feature with the pseudo-similarity score as the rotation relationship between the first global feature and the second global feature in the complex vector space to obtain a rotation similarity score; map the similarity score after fusing the pseudo-similarity score and the rotation similarity score to the similarity score of the first graph data and the second graph data through the similarity calculation module.
8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 6 above is implemented.
9. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in any one of claims 1 to 6 above is implemented.
Citation Information
Cited By
Model management system and method based on image similarity algorithm and database
CN122435295A