A Heterogeneous Information Similarity Matching Method and System Based on Graph Representation

Through a graph representation-based method, the node-level graph of heterogeneous graph is extracted and the vector is embedded in the heterogeneous graph and the global position coding and cross attention mechanism are used to solve the accuracy and stability of heterogeneous information similarity matching, and efficient heterogeneous data matching is achieved.

CN119337148BActive Publication Date: 2025-07-11ZHEJIANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411869600.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-07-11
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

Existing heterogeneous information similarity matching methods have limited effects when dealing with variable and complex heterogeneous data, and the steep learning curve leads to a high risk of inconsistency.

Method used

Using a graph representation method, the node-level graph embedding vectors of heterogeneous graphs are extracted, and the global position coding and cross attention mechanism are used to predict graph-graph similarity scores with multi-layer perception machines to achieve efficient and accurate matching of heterogeneous information.

Benefits of technology

Mapping heterogeneous information into a unified low-dimensional space reduces feature engineering dependence and improves the accuracy and stability of matching multiple types of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119337148B_ABST
    Figure CN119337148B_ABST
Patent Text Reader

Abstract

The present invention discloses a heterogeneous information similarity matching method and system based on graph representation, including the steps of: extracting a heterogeneous graph of heterogeneous information; converting the high-dimensional information of nodes in the heterogeneous graph into a low-dimensional information vector space, and obtaining a node-level graph embedding vector after splicing; extracting a similarity feature vector of graph-graph nodes in an alignment similarity matrix; using a cross-attention mechanism to aggregate the embedding vectors of graph pairs to obtain an aggregated graph-level embedding vector; inputting the similarity feature vector of graph-graph nodes and the aggregated graph-level embedding vector into a multi-layer perceptron to obtain a prediction result of the graph-graph similarity score. In the present invention, the graph-based representation method can map heterogeneous information into a unified low-dimensional space, making similarity matching more efficient and accurate, reducing the dependence on feature engineering, and having a good matching effect when dealing with various types of data and the relationships between them.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information processing, and particularly to a heterogeneous information similarity matching method and system based on graph representation. Background Art

[0002] Heterogeneous Information Similarity Matching has received extensive attention in recent years. With the development of big data technology, more and more practical application scenarios are filled with multi-source heterogeneous information. How to effectively perform similarity matching on these heterogeneous data has become a research focus in the fields of data mining and artificial intelligence. The existing research methods are mainly divided into traditional methods based on feature engineering and methods based on deep learning.

[0003] Traditional methods based on feature engineering: Early research mainly focused on converting heterogeneous data into a unified feature representation through feature engineering, such as text vectorization, image feature extraction, etc. Such methods rely on the knowledge of domain experts for feature construction. Although they have certain effects, they are easily affected by feature selection and data quality when dealing with complex heterogeneous information.

[0004] Methods based on deep learning: With the development of deep learning, many studies have begun to use neural networks to automatically learn features, especially for the representation of unstructured data such as images and texts. For example, Chinese Patent Publication No. CN118378683A discloses a heterogeneous information network representation learning method based on artificial intelligence, which relates to the technical field of network learning and includes: extracting each node of the heterogeneous information network, and based on each node and the corresponding node features for initial modeling to obtain an initial learning model; based on the network feature matching embedding method of the target heterogeneous information network, extracting and transforming each node and the node semantic relationship; combining the initial learning model with the extraction and transformation results to construct a first learning model, and performing model optimization to obtain a first optimized model; performing network representation on the target heterogeneous information network based on the first optimized model, and performing verification and evaluation. By combining the construction of the initial learning model with the network nodes and node features and the results of extracting and transforming the node semantic relationship according to the corresponding embedding method to construct the first learning model for network representation and verification and evaluation of the representation results, the network representation of the target heterogeneous information network can be made more accurate and effective. Such methods can reduce the dependence on feature engineering to a certain extent. However, when it comes to multiple types of data and the relationships between them, the effect is still limited.

[0005] CN113035319 A Method and system for automatically recommending a lifestyle management path for diabetic patients, comprising: acquiring data related to the lifestyle management of diabetic patients to form multi-source heterogeneous data; constructing a multi-modal knowledge graph for the lifestyle management of diabetic patients based on the multi-source heterogeneous data; complementing the relationships in the multi-modal knowledge graph; inputting the semantic information of the multi-modal knowledge graph with complemented relationships into a deep learning model for training and learning; constructing a sub-knowledge graph based on the pathological information and target lifestyle information of a certain diabetic patient, inputting the semantic information of the sub-knowledge graph into the trained deep learning model for graph similarity matching prediction, and extracting the part with the highest similarity to the sub-knowledge graph in the multi-modal knowledge graph as the recommended lifestyle management path for the diabetic patient.

[0006] The graph structure can effectively represent different types of nodes and edges, thereby expressing the relationships between complex heterogeneous information. In recent years, graph neural networks have shown great potential in the field of heterogeneous information similarity matching, but there is not much research. Summary of the Invention

[0007] In view of the variable and complex characteristics of heterogeneous information, the difficulty and accuracy of similarity matching need to be improved. The present invention provides a method for similarity matching of heterogeneous information based on graph representation to solve the problem of limited matching effect for variable and complex heterogeneous data features and the higher risk of inconsistency caused by a steeper learning curve.

[0008] To achieve the above object, the technical solution adopted by the present invention is:

[0009] A method for similarity matching of heterogeneous information based on graph representation, comprising the steps of:

[0010] S1, extracting the heterogeneous graph of heterogeneous information;

[0011] S2, converting the high-dimensional feature vectors of the nodes in the heterogeneous graph into low-dimensional feature vectors, and obtaining node-level graph embedding vectors after splicing the low-dimensional feature vectors;

[0012] S3, using global position encoding to ensure the uniqueness of the generated similarity matrix, and obtaining the similarity matrix after splicing the node-level graph embedding vectors after global position encoding; aligning the similarity matrix by supplementing zero embeddings, and extracting the similarity feature vectors of graph-graph nodes in the aligned similarity matrix;

[0013] S4, using a cross-attention mechanism to aggregate the embedding vectors of the graph pair to obtain an aggregated graph-level embedding vector;

[0014] S5, inputting the similarity feature vectors of graph-graph nodes in S3 and the aggregated graph-level embedding vector in S4 into a multi-layer perceptron to obtain the prediction result of the graph-graph similarity score.

[0015] In some specific embodiments, S1 specifically includes the steps of:

[0016] S11, determining the correspondence between elements and attributes in heterogeneous information and nodes, edges, and attributes in the graph;

[0017] S12, extracting elements and attributes in heterogeneous information, and converting them into nodes and edges to obtain a heterogeneous graph data structure.

[0018] In some specific embodiments, S2 specifically includes the steps of:

[0019] S21, meta-path extraction: traversing the meta-paths of nodes in the heterogeneous graph by fusing depth-first search and breadth-first search, and extracting the set of meta-paths of nodes in the heterogeneous graph;

[0020] S22, node content conversion: mapping and transforming node features of different types in the heterogeneous graph into the same latent vector space to obtain transformed node feature vectors;

[0021] S23, performing in-aggregate and out-aggregate on the meta-path to obtain node-level graph embedding vectors.

[0022] Further, S23 specifically includes the steps of:

[0023] Encoding the meta-path instances extracted in S21 to obtain node feature vectors, and obtaining the meta-path feature vectors of nodes through weighted summation; encoding all meta-path instances of nodes and performing weighted summation to obtain a set of meta-path feature vectors of nodes;

[0024] Performing mean aggregation on the same meta-path for multiple sets of meta-path feature vectors of nodes of the same type to obtain a meta-path specific node vector representation, fusing the specific node vector representations with the attention mechanism for weighted summation, and realizing the embedding mapping of the high-dimensional feature vectors containing all meta-path features of nodes to the vector space of low-dimensional information through additional linear transformation to obtain the low-dimensional feature vectors of nodes, and concatenating them to obtain node-level graph embedding vectors.

[0025] In some specific embodiments, using global position encoding to ensure the uniqueness of the generated similarity matrix specifically includes the steps of:

[0026] By marking seed nodes and key nodes to merge graph pairs, and using random walk, sorting the nodes according to the probability that the nodes return to themselves for global position encoding rearrangement.

[0027] In some specific embodiments, the method of supplementing zero embeddings in step 3 to align the similarity matrix specifically includes the steps of:

[0028] For node-level graph embeddings with different dimensions, by constructing dummy nodes, zero embeddings are supplemented in the graph embeddings with a smaller number of nodes for dimension complementation to achieve alignment of the similarity matrix.

[0029] In some specific embodiments, the step of extracting the similarity feature vector of nodes in step 3 specifically includes the steps of:

[0030] Using a convolutional neural network to convert the similarity calculation of the aligned similarity matrix as a pattern recognition problem for feature extraction, and obtaining the similarity feature vector of graph-graph nodes.

[0031] In some specific embodiments, S4 specifically includes the steps of:

[0032] S41, S4 specifically includes the steps of:

[0033] S41, calculating the cross-graph attention coefficient α i ∈V 1 between the node v in graph G1 j ∈V 2 and all other nodes v in graph G2 ij ; at the same time, calculating the cross-graph attention coefficient β j ∈V 2 between the node v in graph G2 i ∈V 1 and all other nodes v in graph G1 ji ;

[0034]

[0035] where: is an attention function for calculating the similarity score between two node embedding vectors, is the embedding vector of the node v i in graph G1; is the embedding vector of the node v j in graph G j ;

[0036] S42, in the view of the node v i ∈V 1 in G1, calculating the attention graph-level embedding vector of G2 by weighted averaging all node embeddings of G2; in the view of the node v j ∈V 2 in G2, calculating the attention graph-level embedding vector of G1 by weighted averaging all node embeddings of G1;

[0037]

[0038] S43, using a BiLSTM (Bidirectional Long Short-Term Memory Network) for the attention graph-level embedding vector of graph G1 and the attention graph-level embedding vector Aggregate them separately to obtain the aggregated graph-level embedding vectors of graph G1 and graph G2.

[0039] In some specific embodiments, S5 specifically includes the steps of:

[0040] Input the similarity feature vectors of graph-graph nodes obtained in S3 and the aggregated graph-level embedding vectors obtained in S4 into a multi-layer perceptron, and fuse them to execute the sigmoid activation function to obtain the graph-graph similarity score.

[0041] The present invention also provides a system for realizing similarity matching of heterogeneous information based on graph representation by using the above method, including a heterogeneous graph extraction module, a low-dimensional embedding vector module, a similarity feature extraction module, an aggregated graph-level embedding module, and a similarity score prediction module;

[0042] The heterogeneous graph extraction module is used to extract the heterogeneous graph of heterogeneous information;

[0043] The low-dimensional embedding vector module is used to transform the high-dimensional information of the nodes in the heterogeneous graph into a low-dimensional information vector space that aggregates the semantic and structural features of the neighborhood nodes and the self-node, and splices them to obtain the node-level graph embedding vector;

[0044] The similarity feature extraction module is used to use global position encoding to ensure the uniqueness of the generated similarity matrix, and after global position encoding, splice the node-level graph embedding vector to obtain the similarity matrix; align the similarity matrix by supplementing zero embeddings, and extract the similarity feature vectors of graph-graph nodes in the aligned similarity matrix;

[0045] The aggregated graph-level embedding module is used to aggregate the embedding vectors of the graph pair by using the cross-attention mechanism to obtain the aggregated graph-level embedding vector;

[0046] The similarity score prediction module is used to input the similarity feature vectors of graph-graph nodes and the aggregated graph-level embedding vectors into a multi-layer perceptron to obtain the graph-graph similarity score prediction result.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] In the present invention, the graph-based representation method can map heterogeneous information into a unified low-dimensional space, making similarity matching more efficient and accurate, reducing the dependence on feature engineering, and having a good matching effect when dealing with various types of data and the relationships between them. Description of the Drawings

[0049] Figure 1 It is a graph of the heterogeneous information similarity matching algorithm model based on graph representation in the present invention.

[0050] Figure 2Schematic diagram of meta-path extraction based on the fusion traversal of DFS and BFS in the present invention.

[0051] Figure 3 Global position encoding diagram in the present invention.

[0052] Figure 4 Schematic diagram of similarity score prediction in the present invention. Detailed implementation manners

[0053] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Those skilled in the art who make modifications or equivalent replacements based on the understanding of the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention shall be covered by the protection scope of the present invention.

[0054] Embodiment

[0055] A heterogeneous information similarity matching method based on graph representation, specifically as Figure 1 shown, including the steps:

[0056] S1. Extract the heterogeneous graph of heterogeneous information, and retain the semantics and relationships of the original heterogeneous information model in the heterogeneous graph; specifically, in this embodiment, taking the SysML model as an example, it includes the steps:

[0057] S11. Determine the corresponding relationships between the elements and attributes in the heterogeneous information and the nodes, edges and attributes in the graph; among them, the nodes can represent various elements in SysML, such as Block, Port, Activity, etc. The edges can represent the relationships between elements, such as Connector, Dependency (dependency relationship), Association (association relationship), etc.

[0058] S12. During the conversion process, in addition to mapping the elements and relationships, it is also necessary to convert the attribute information in the SysML model into the attributes of the nodes and edges in the graph. Extract the elements and attributes in the heterogeneous information, and convert them into nodes and edges to obtain the heterogeneous graph data structure. For example, it includes: Node attributes: such as the name of the Block, the direction of the Port (input or output), the operation name of the Activity, etc. Edge attributes: such as the type of the Connector (such as data flow, control flow), the weight of the Dependency, the implementation condition of the Realization, etc.

[0059] S2. Convert the high-dimensional information of the nodes in the heterogeneous graph into a low-dimensional information vector space that aggregates the semantics and structural features of the neighborhood nodes and the self-node; specifically,

[0060] S21, Meta-path extraction: As shown in Figure 2 , traverse the meta-paths of nodes in the heterogeneous graph by fusing depth-first search (DFS) and breadth-first search (BFS) to extract the set of meta-paths of nodes in the heterogeneous graph; for example, given a node V of type A in a heterogeneous graph, it is necessary to extract the maximum number of equal-length meta-paths that start with type A and end with type A nodes. BFS controls the step size s of the meta-path from a macroscopic perspective, and DFS generates a set of meta-paths that meet the extraction conditions according to different step numbers from a microscopic perspective. Subsequently, select the set with the largest number K of meta-paths in the step size set to obtain a set of meta-paths for a type A node , where , , , …… represent the number of meta-paths with step sizes of 1, 2, 3... n.

[0061] S22, Node content transformation: Map and transform the node features of different types in the heterogeneous graph into the same latent vector space to obtain the transformed node feature vectors; for a node v with node type A, the transformation process is as follows:

[0062]

[0063] where: x v is the original feature vector of node v, h v is the feature vector of node v in the mapping space after transformation, and W A is the parameter weight matrix of type A nodes. After this transformation, the mapping features of all nodes share the same dimension, providing conditions for the aggregation in the next part.

[0064] S23, Intra-meta-path aggregation: Encode the meta-path instances extracted in S21 to obtain node feature vectors, and obtain the meta-path feature vectors of nodes through weighted summation; encode all meta-path instances and perform weighted summation to obtain a set of meta-path feature vectors of nodes;

[0065] Given a meta-path instance , ; Intra-aggregation refers to learning the structural, semantic information of the target node and the context information of the neighborhood in the meta-path. In this embodiment, a special meta-path instance encoder is used to convert all node features into a single vector along the meta-path instance, and the process is as follows:

[0066]

[0067] where: h P(v,u) is the node feature vector, is the starting point feature in the meta-path, is the ending point feature in the meta-path, are all intermediate node features in the meta-path; the encoder is a linear average encoder, . Where W P is a linear transformation matrix, the element-wise average of the MEAN node vectors.

[0068] After the meta-path instance is vectorized, in order to enable different meta-path instances to promote the representation of the target node to different degrees, the attention mechanism is used to perform weighted summation on all relevant meta-paths of the target node:

[0069]

[0070] Where: is the attention vector of the meta-path, represents the concatenation of the target node and the expression vector for which the attention weight is to be solved. LeakyReLU is a commonly used activation function, obtaining the importance of the meta-path P(v, u) for the target node v - the attention weight .

[0071] Next, for unified comparison, softmax normalization processing is performed to obtain , integrating all meta-path-based neighbors and attention weights obtaining the normalized importance weight of each meta-path for the target node , which, summed with the node feature vector h P(v,u) through the activation function results in the meta-path feature vector of the embedded node .

[0072] In summary, given the feature vector of a target node after mapping and a set of meta-paths where both the starting point and the ending point are of type A , this module can generate a set of meta-path vectors of M target nodes , and each vector representation can be regarded as an aggregation of semantic information of node v.

[0073] S24. For meta-path external aggregation: For multiple sets of meta-path feature vectors of nodes of the same type, mean aggregation of the same meta-path is performed to obtain the meta-path specific node vector representation , and the attention mechanism is used to fuse and perform weighted summation on the set of meta-path feature vectors of the node to obtain the high-dimensional feature vector containing all meta-path features of the node ; the high-dimensional feature vector containing all meta-path features of the node The node feature vector is embedded and mapped to a vector space of low-dimensional information through an additional linear transformation to obtain the low-dimensional feature vector of the node , and the node-level graph embedding vector is obtained after splicing.

[0074] Specifically, since different meta-paths contribute differently to the current node representation, it is necessary to continue to aggregate the semantic information of all meta-paths through the attention layer. After aggregating the information of nodes and edges on each meta-path, a set of meta-path vectors of type-A nodes is obtained , where M is the number of meta-paths of type-A nodes. Given a meta-path P i , first, the mean aggregation is performed on the set of meta-path instance vectors of each specific type of node to obtain the meta-path specific node vector representation :

[0075]

[0076] where: M A and b A are both learnable parameters, |V A | is the number of type-A node vector sets, .

[0077] Then, the attention mechanism is used to fuse the feature vectors of the target node v in different meta-paths, and the meta-path specific node (type-A) vector representation obtains a scalar reflecting the correlation between the type-A node attention vector and the meta-path specific node (type-A) vector representation through the dot product operation , and the larger this value is, the higher the importance of the meta-path P i to the type-A node.

[0078] Then, the scalar of the correlation is converted into the importance weight β through the normalization function, and the importance weight β Pi and the meta-path feature vector of the node Pi are weighted and summed through the activation function to obtain the high-dimensional feature vector of the node containing all meta-path features . The formula is as follows: . The formula is as follows:

[0079]

[0080] where: is the parameterized attention vector of the type-A node.

[0081] Finally, an additional linear transformation with a non-linear function is adopted to map the high-dimensional feature vector of the node containing all meta-path features to a vector space with the desired output dimension to obtain the low-dimensional feature vector of the node :

[0082]

[0083] Wherein: is an activation function, and W o is a weight matrix.

[0084] For all nodes of graph G i after being vectorized through the above steps and concatenated, the node-level graph embedding of graph G i can be obtained, where n is the number of nodes in the graph. , where n is the number of nodes in the graph.

[0085] S3. Align the similarity matrix by supplementing zero embeddings, and extract the similarity feature vectors of graph-graph nodes in the aligned similarity matrix; and use global position encoding to ensure the uniqueness of the generated similarity matrix.

[0086] The node-level graph embedding obtained in step S24 may generate different similarity matrices S and feature vectors vec(S) due to different initial node orderings during concatenation. To ensure uniqueness, encoding rearrangement is required. And in the heterogeneous information similarity matching task, the two graphs cannot be separately position-encoded, but the global position encoding of the nodes needs to be calculated by combining the graph pair.

[0087] First, as Figure 3 shown, use the key nodes and predefined seed nodes to connect and merge the graph pair. The key nodes can be defined as the two nodes with the highest matching degree in the graph pair, and the seed nodes are the nodes where the two key nodes match their corresponding subgraphs G1 and G2. Use a k-step random walk method to perform encoding on the merged graph, defined as follows:

[0088]

[0089] Where: k is the step size, and K is the maximum step size; T is the state transition matrix of the random walk, and T k represents the k-th power of T. T ii is a random walk matrix that only focuses on the probability of node i returning to itself; finally, reverse-order the nodes according to the size of the probability that node i returns to itself, and the smaller the probability, the higher the ranking.

[0090] After determining the global encoding of the nodes, re-concatenate the node-level graph embeddings U i and U j to obtain the similarity matrix .

[0091] For two graphs with different numbers of nodes, the obtained node-level graph embedding dimensions are different and cannot be matrix-multiplied. The method of supplementing zero embeddings in step 3 to align the similarity matrices specifically includes the following steps: For node-level graph embeddings with different dimensions, zero embeddings are supplemented in the graph embedding with a smaller number of nodes by constructing dummy nodes to complete the dimension complementation and achieve the alignment of the similarity matrices.

[0092] The specific steps for extracting the similarity feature vectors of nodes in step 3 are as follows: The convolutional neural network is used to convert the similarity calculation of the aligned similarity matrices as a pattern recognition problem for feature extraction, and the similarity feature vectors of graph-graph nodes are obtained. 。

[0093] S4. The cross-attention mechanism is used to aggregate the embedding vectors of the graph pair to obtain the aggregated graph-level embedding vector; specifically including the following steps:

[0094] S41. Calculate the cross-graph attention coefficient α i ∈V 1 between the node v j ∈V 2 in graph G1 and all other nodes v ij ∈V j in graph G2; at the same time, calculate the cross-graph attention coefficient β 2 between the node v i ∈V 1 in graph G2 and all other nodes v ji in graph G1;

[0095]

[0096] where: is the attention function for calculating the similarity score between two node embedding vectors, is the embedding vector of the node v i in graph G1; is the embedding vector of the node v j in graph G j ; In this embodiment, the cosine function cosine is used, and other similarity measurement metrics can also be used.

[0097] S42. In the view of the node v i ∈V 1 in G1, calculate the attention graph-level embedding vector of G2 by weighted averaging all the node embeddings of G2; in the view of the node v j ∈V 2 in G2, calculate the attention graph-level embedding vector of G1 by weighted averaging all the node embeddings of G1;

[0098]

[0099] S43, use BiLSTM (Bidirectional Long Short-Term Memory Network) to perform attention graph-level embedding vectors on graph G1 and the attention graph-level embedding vectors of graph G2 respectively aggregate to obtain the aggregated graph-level embedding vectors of graph G1 and graph G2 ;

[0100] 。

[0101] S5, input the similarity feature vectors of graph-graph nodes in S3 and the aggregated graph-level embedding vectors in S4 into a multi-layer perceptron to obtain the prediction result of the graph-graph similarity score. As Figure 4 shown, use four standard fully connected layers to gradually project the dimension of the result vector onto a scalar of dimension 1. Since the expected similarity score should be in the range of [0,1], so the sigmoid activation function is executed to map the similarity score into this range.

[0102] Input the similarity feature vectors of graph-graph nodes obtained in S3 and the aggregated graph-level embedding vectors obtained in S4 into the multi-layer perceptron, and fuse to execute the sigmoid activation function to obtain the graph-graph similarity score ;

[0103] 。

[0104] Finally, the HGraphSim model is trained on a set of n triple data containing two input heterogeneous graph structures and two icon scalar similarity scores . And the mean squared error loss function is used to train the model, and the calculated similarity score is compared with the true similarity score y, and the loss function Loss is as follows:

[0105]

[0106] In this embodiment, specifically, taking the AIDS dataset and the LINUX dataset as examples, the above-mentioned graph representation method of the present invention is used to perform similarity matching on the information in the dataset. AIDS is a molecular graph representation dataset of AIDS antiviral active compounds. It contains 42,687 abbreviated compound structures, and the molecules are directly converted into graphs by representing atoms as nodes, covalent bonds as edges. LINUX consists of 48,747 program dependence graphs (PDGs) generated from the Linux kernel. Each graph represents a function, where a node represents a statement and an edge represents the dependence relationship between these two statements.

[0107] The similarity scoring criterion is calculated based on the predicted GED and the true GED. The graph edit distance (GED): The edit distance between two graphs G1 and G2 refers to the minimum number of operations required to transform G1 into G2. The operations include inserting or deleting nodes, inserting or deleting edges, relabeling nodes or edges. If we want to convert the GED into a scoring criterion for measuring the similarity of two graphs, a normalization function is needed to transform the GED into the range of 0 to 1.

[0108] To verify the effectiveness and stability of the proposed method, experiments on the GED similarity scoring criterion were conducted on the dataset using the method proposed in this paper and the baseline Noah method based on the above content. The evaluation metrics in the experimental results are shown in Table 1, where the method proposed in this paper is defined as HGraphSim.

[0109] Table 1 Similarity matching evaluation metrics for different datasets

[0110]

[0111] Among them, Noah refers to a method different from the end-to-end learning model. It combines the A* algorithm and the graph neural network to optimize the search direction of the A* algorithm and simultaneously learn the elastic adjustment of the search range size of the heuristic function.

[0112] Mean squared error (MSE): It measures the average squared difference between the calculated similarity score and the true similarity score, and is used to evaluate the stability of the model in calculating similarity.

[0113] Spearman's rank correlation coefficient (ρ), Kendall's rank correlation coefficient (τ): These two metrics measure the matching ratio between the calculated ranking results and the true ranking results.

[0114] Precision at k (p@k): This metric measures the matching ratio between the top k calculated results and the true top k results. It focuses on the top k results rather than the global ranking results. Specifically, in this evaluation, k values used are 10 and 20 respectively.

[0115] Among them, the symbol ↑ indicates that the higher the value, the better the evaluation effect of the indicator, while ↓ indicates that the lower the value, the better. It can be seen that the method proposed in this paper outperforms the baseline method in the performance of multiple indicators. Therefore, the following conclusions can be drawn: (1) HGraphSim has an excellent model structure and stable computing ability. Specifically, when obtaining heterogeneous graph node embeddings, it can well aggregate the domain structure and its own semantics, and can well extract the similarity features between graph-graph and node-node through the similarity matrix and cross-attention methods. (2) HGraphSim has a stronger ability to learn similarity scores in different samples. Whether in the AIDS molecular dataset or the program dependence graph dataset, its effect of completing the similarity matching task is better.

Claims

1. A method for matching the similarity of heterogeneous information based on graph representation, characterized in that, Including steps: S1. Extract the heterogeneous graph of heterogeneous information in the AIDS dataset or LINUX dataset from the SysML model; The AIDS dataset converts molecules into graphs in a direct way by representing atoms as nodes, covalent bonds as edges; each graph in the LINUX dataset represents a function, where a node represents a statement and an edge represents the dependency between two statements; S2. Convert the high-dimensional feature vectors of the nodes in the heterogeneous graph into low-dimensional feature vectors, and the concatenated low-dimensional feature vectors obtain the node-level graph embedding vectors; S3. Use global position encoding to ensure the uniqueness of the generated similarity matrix. After global position encoding, concatenate the node-level graph embedding vectors to obtain the similarity matrix; align the similarity matrix by supplementing zero embeddings, and extract the similarity feature vectors of graph-graph nodes in the aligned similarity matrix; S4. Adopt a cross-attention mechanism to aggregate the embedding vectors of the graph pair to obtain the aggregated graph-level embedding vectors; S5. Input the similarity feature vectors of graph-graph nodes in S3 and the aggregated graph-level embedding vectors in S4 into a multi-layer perceptron to obtain the prediction result of the graph-graph similarity score; S2 specifically includes steps: S21. Meta-path extraction: Use a combination of depth-first search and breadth-first search to traverse the meta-paths of the nodes in the heterogeneous graph, and extract the set of meta-paths of the nodes in the heterogeneous graph; S22. Node content conversion: Map and transform the node features of different types in the heterogeneous graph into the same latent vector space to obtain the transformed node feature vectors; S23. Aggregate inside and outside the meta-path to obtain the node-level graph embedding vectors; The specific steps of using global position encoding in S3 to ensure the uniqueness of the generated similarity matrix include: Merge the graph pairs by marking the seed nodes and key nodes, and use random walk to reorder the nodes according to the probability that the nodes return to themselves for global position encoding rearrangement; The specific steps of aligning the similarity matrix by supplementing zero embeddings in step 3 include: For node-level graph embeddings with different dimensions, supplement zero embeddings in the graph embedding with fewer nodes by constructing dummy nodes to complete the dimension complementation and achieve the alignment of the similarity matrix; S4 specifically includes steps: S41, calculate the node v in graph G1 i ∈V 1 and the cross - graph attention coefficient α j ∈V 2 between it and all other nodes v in graph G2 ij ; At the same time, calculate the node v in graph G2 j ∈V 2 and the cross - graph attention coefficient β i ∈V 1 between it and all other nodes v in graph G1 ji ; ; Wherein: is an attention function for calculating the similarity score between two node embedding vectors, is the embedding vector of node v in graph G1 i ; is the embedding vector of node v in graph G j ; j ; S42, in the view of node v of G1 i ∈V 1 calculate the attention map-level embedding vector of G2 by weighted averaging all node embeddings of G2 ; in the view of node v of G2 j ∈V 2 calculate the attention map-level embedding vector of G1 by weighted averaging all node embeddings of G1 ; ; S43, using BiLSTM to perform attention graph-level embedding vectors for graph G1 and the attention graph-level embedding vectors for graph G2 are respectively aggregated to obtain the aggregated graph-level embedding vectors for graph G1 and graph G2.

2. The method for matching the similarity of heterogeneous information based on graph representation according to claim 1, wherein S1 specifically includes steps: S11. Determine the correspondence between elements and attributes in the heterogeneous information and nodes, edges, and attributes in the graph; S12. Extract the elements and attributes in the heterogeneous information, and convert them into nodes and edges to obtain the heterogeneous graph.

3. The method for matching the similarity of heterogeneous information based on graph representation according to claim 1, characterized in that S23 specifically includes steps: Encode the meta-path instances extracted in S21 to obtain node feature vectors, and obtain the meta-path feature vectors of the nodes through weighted summation; encode all meta-path instances of the nodes and obtain the set of meta-path feature vectors of the nodes through weighted summation; Perform mean aggregation of the same meta-path for multiple sets of meta-path feature vectors of the same type of nodes to obtain the meta-path specific node vector representation, use the attention mechanism to fuse the weighted summation of the specific node vector representation, and realize the high-dimensional feature vector embedding of the nodes containing all meta-path features to be mapped to the vector space of low-dimensional information to obtain the low-dimensional feature vectors of the nodes, and the concatenated ones obtain the node-level graph embedding vectors.

4. The heterogeneous information similarity matching method based on graph representation according to claim 1, wherein The specific steps of extracting the similarity feature vectors of the nodes in step 3 include: The similarity calculation of the aligned similarity matrix is transformed into a feature extraction problem for pattern recognition using a convolutional neural network, and a similarity feature vector of graph-graph nodes is obtained.

5. The heterogeneous information similarity matching method based on graph representation according to claim 1, wherein S5 specifically includes the steps: The similarity feature vector of graph-graph nodes obtained in S3 and the aggregated graph-level embedding vector obtained in S4 are input into a multi-layer perceptron, and the sigmoid activation function is fused to obtain a graph-graph similarity score.

6. A system for realizing similarity matching of heterogeneous information based on graph representation by using the method according to any one of claims 1-5, characterized in that It includes a heterogeneous graph extraction module, a low-dimensional embedding vector module, a similarity feature extraction module, an aggregated graph-level embedding module, and a similarity score prediction module; The heterogeneous graph extraction module is used to extract a heterogeneous graph of heterogeneous information; The low-dimensional embedding vector module is used to transform the high-dimensional information of the nodes in the heterogeneous graph into a low-dimensional information vector space that aggregates the semantic and structural features of the neighborhood nodes and its own node, and the node-level graph embedding vector is obtained after splicing; The similarity feature extraction module is used to use global position encoding to ensure the uniqueness of the generated similarity matrix. After global position encoding, the node-level graph embedding vector is spliced to obtain a similarity matrix; the similarity matrix is aligned by supplementing zero embeddings, and the similarity feature vector of graph-graph nodes in the aligned similarity matrix is extracted; The aggregated graph-level embedding module is used to aggregate the embedding vectors of graph pairs using a cross-attention mechanism to obtain an aggregated graph-level embedding vector; The similarity score prediction module is used to input the similarity feature vector of graph-graph nodes and the aggregated graph-level embedding vector into a multi-layer perceptron to obtain a graph-graph similarity score prediction result.

Citation Information

Patent Citations

  • Diabetes patient lifestyle management path automatic recommendation method and system

    CN113035319A

  • Heterogeneous information network representation learning method based on artificial intelligence

    CN118378683A

  • Legal case similarity calculation method and system based on knowledge graph matching

    CN114092283A

  • Code similarity detection method based on code attribute graph

    CN115438709A