Heterogeneous Graph Embedding Method Based on Information Completion
By using the information completion method in the heterogeneous graph, nodes with missing attributes are processed, and through feature mapping, weighted averaging, similarity learning and adaptive weight training, the attribute missing and long-tail problems in the heterogeneous graph are solved, and the model performance and node representation quality are improved.
Patent Information
- Application Number
- CN202411021558.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-07-29
AI Technical Summary
When dealing with nodes with missing attributes in heterogeneous graphs, the simple filling method adopted by the prior art is prone to introduce noise, resulting in loss of model performance, and cannot effectively solve the long-tail problem, resulting in a degradation of model performance.
Using a heterogeneous graph embedding method based on information completion, the node features are mapped to the shared semantic space by determining the target node set, filtering the first k neighbor nodes for weighted average, obtaining a preliminary complementary feature vector, learning the similarity between nodes, generating a relational adjacency graph, aggregating an adjacency matrix, performing adaptive weight matrix training, performing feature embedding and neighborhood transfer, and finally obtaining the node's complete complementary feature vector.
It effectively solves the problem of missing attributes in heterogeneous graphs, reduces noise information, and improves model performance. Especially when dealing with long-tail problems, it can better learn node representations and improves the performance of downstream tasks.
Smart Images

Figure CN118981554B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graph data mining, and in particular to a heterogeneous graph embedding method based on information completion. Background Art
[0002] As a type of graph data structure, heterogeneous graphs have diverse node and edge types, and carry complex and rich information, which poses great challenges to data analysis. At the same time, it also makes the analysis and mining of heterogeneous graph information a hot topic in the academic community. Among many graph mining methods, graph embedding methods stand out. They can project data in a high-dimensional network into a low-dimensional continuous space while capturing the intrinsic information of the network. Compared with other graph mining methods, they have certain performance advantages.
[0003] When using graph embedding methods to capture the neighborhood information of a network, rich semantic information in node attributes will be used to significantly enhance the representation ability of nodes, thus achieving excellent performance in various graph tasks. However, in real-world scenarios, the nodes of a heterogeneous graph can be divided into two categories: one is nodes with complete attributes, which have a complete set of attributes; the other is nodes with missing attributes. Since models based on attribute feature aggregation are highly sensitive to node attributes, if a graph embedding method is directly applied to a heterogeneous graph with nodes having missing attributes, it will seriously affect the performance of the model based on attribute feature aggregation in the graph embedding method, resulting in the model based on attribute feature aggregation being unable to accurately capture the features of nodes. To solve this problem, the existing technology simply uses one-hot encoding or calculates and fills with other node information when dealing with such nodes with missing attributes. The node features obtained in this way often cannot accurately reflect the true characteristics of nodes and are prone to introducing noise information, affecting the performance of the model.
[0004] Currently, there has been a very mature research on the long-tail problem in homogeneous graphs. However, the solutions to the long-tail problem in homogeneous graphs cannot be directly applied to heterogeneous graphs. In the existing technology, the imbalance in the number of relationship connections of different nodes in heterogeneous graphs has been noticed. Thus, there is relatively little research on the deficiency of the node representations learned by the model for nodes with fewer relationship connections in heterogeneous graphs. In heterogeneous graphs, for nodes with more relationship connections (usually called head nodes), their neighborhood information is relatively rich, and more features can be effectively aggregated. While for nodes with fewer relationship connections (usually called tail nodes), their neighborhood information is relatively scarce, resulting in the model being unable to fully learn the node representations, thereby degrading the performance of downstream tasks. To solve the above problems, the existing technology uses the adjacency matrix for relationship completion, which alleviates the long-tail problem in heterogeneous graphs to a certain extent. However, simply increasing the connections of tail nodes may introduce small neighborhoods with biases or insufficient representativeness, still leading to a decline in model performance, making the model unable to learn the feature representations of nodes well, and thus affecting the performance of downstream tasks. For example, it may lead to inaccurate classification in node classification, poor node information aggregation effect in node clustering tasks, and inability to accurately generate node representations in link prediction. Summary of the Invention
[0005] Therefore, the technical problem to be solved by the present invention is to overcome the problem that in the prior art, only simple filling operations are taken for nodes with missing attributes, which easily introduces noise and causes loss of model performance; the method of using the adjacency matrix for relationship completion cannot well solve the long-tail problem existing in heterogeneous graphs, still leading to a decline in model performance, and further degrading the performance of downstream tasks.
[0006] To solve the above technical problems, the present invention provides a heterogeneous graph embedding method based on information completion, including:
[0007] A heterogeneous graph embedding method based on information completion, characterized by including:
[0008] Based on the heterogeneous graph, determine and construct a target node set; map the initial feature vectors of all nodes in the heterogeneous graph to a unified shared semantic space to obtain the projected feature vectors of all nodes.
[0009] Sort according to the magnitude of the similarity of attribute features between the current target node and its respective neighbor nodes, and select the top k neighbor nodes of the current target node; perform weighted averaging on the projected feature vectors of the current target node and its top k neighbor nodes to obtain the preliminary completed feature vector of the current target node.
[0010] Learn the similarity between nodes in the heterogeneous graph based on the preliminary complemented feature vectors of each target node and the projected feature vectors of each neighbor node of each target node, and obtain the relational adjacency graph of each relationship in the heterogeneous graph; Aggregate the initial adjacency matrix of each relationship with the relational adjacency graph to obtain the target adjacency matrix of each relationship, so as to obtain the target adjacency matrix of the current target node in the current relationship;
[0011] Based on all nodes in the heterogeneous graph, set the adaptive weight matrix and the number of training times corresponding to each node in the initial round. Based on the adaptive feature vector of the current node in the previous round and the adaptive weight matrix corresponding to the current node in the current round, obtain the adaptive feature vector of the current node in the current round; Until the preset number of training times is reached, obtain the target feature vectors of each node; Among them, when the training round is 1, the adaptive feature vector of the current target node in the previous round is its preliminary complemented feature vector, and the adaptive feature vector of each neighbor node of the current target node in the previous round is its projected feature vector;
[0012] Based on the target adjacency matrix of the current target node in the current relationship, perform a normalization operation on the target feature vectors of each neighbor node of the current target node in the current relationship to obtain the feature embedding vectors of each neighbor node of the current target node in the current relationship;
[0013] Perform a weighted average on the target feature vector of the current target node and the feature embedding vectors of each neighbor node of the current target node in the current relationship to obtain the neighborhood embedding vector of the current target node in the current relationship as the actual neighborhood of the current target node in the current relationship; Set the neighborhood transfer vector of the current target node in the current relationship, define the ideal neighborhood of the current target node in the current relationship, and determine the difference between the ideal neighborhood and the actual neighborhood of the current target node in the current relationship; Use scaling and transfer transformations to minimize the difference between the ideal neighborhood and the actual neighborhood of the current target node in the current relationship to obtain the neighborhood transfer vector of the current target node in the current relationship, and sequentially obtain the neighborhood transfer vectors of the current target node in each relationship; Based on the target feature vectors of the current target node and its each neighbor node, and the neighborhood transfer vectors of the current target node in each relationship, obtain the complete complemented feature vector of the current target node;
[0014] Based on the complete complemented feature vectors of each target node and the target feature vectors of each neighbor node of each target node, calculate the relationship strength between nodes in the heterogeneous graph, and aggregate it using the attention mechanism to obtain the context representation of each target node, and perform a residual connection to obtain the final node representation of each target node.
[0015] Preferably, based on all nodes in the heterogeneous graph, set the adaptive weight matrix and the number of training times corresponding to each node in the initial round. Based on the adaptive feature vector of the current node in the previous round and the adaptive weight matrix corresponding to the current node in the current round, obtain the adaptive feature vector of the current node in the current round; until the preset number of training times is reached, obtain the target feature vectors of each node; where, when the training round is 1, the adaptive feature vector of the current target node in the previous round is its preliminary completion feature vector, and the adaptive feature vector of each neighbor node of the current target node in the previous round is its projection feature vector, including:
[0016] Based on the adaptive feature vector of the target node v in the (p - 1)th round and the adaptive weight matrix corresponding to the target node v in the pth round Obtain the adaptive feature vector of the target node v in the pth round, and its expression is:
[0017]
[0018] Where, represents the adaptive feature vector of the target node v in the pth round; represents the adaptive feature vector of the target node v in the (p - 1)th round. In the initial round, that is, when p = 1, represents the adaptive weight matrix corresponding to the target node v in the pth round;
[0019] Based on the adaptive feature vector of any neighbor node u of the target node v in the (p - 1)th round and the adaptive weight matrix corresponding to the neighbor node u of the target node v in the pth round Obtain the adaptive feature vector of the neighbor node u of the target node v in the pth round, and its expression is:
[0020]
[0021] Where, represents the adaptive feature vector of the neighbor node u of the target node v in the pth round; represents the adaptive feature vector of the neighbor node u of the target node v in the (p - 1)th round. In the initial round, that is, when p = 1, represents the adaptive weight matrix corresponding to the neighbor node u of the target node v in the pth round;
[0022] Until the preset number of training times is reached, obtain the target feature vector h′ of the target node v v , the target feature vector h′ of the neighbor node u of the target node v u .
[0023] Preferably, normalizing the target feature vectors of each neighbor node of the current target node in the current relationship based on the target adjacency matrix of the current target node in the current relationship to obtain the feature embedding vectors of each neighbor node of the current target node in the current relationship includes:
[0024] Based on the target adjacency matrix A' of the target node v in the relationship r v,r , normalizing the target feature vector h' of the neighbor node u of the target node v in the relationship r u to obtain the feature embedding vector of the neighbor node u of the target node v in the relationship r The expression is:
[0025]
[0026] where h' u represents the target feature vector of the neighbor node u of the target node v; represents the normalization coefficient of the target node v in the relationship r, and its expression is: D r represents the degree matrix of A' v,r ;
[0027] Preferably, performing a weighted average on the target feature vector of the current target node and the feature embedding vectors of each neighbor node of the current target node in the current relationship to obtain the neighborhood embedding vector of the current target node in the current relationship as the actual neighborhood of the current target node in the current relationship; setting the neighborhood transfer vector of the current target node in the current relationship, defining the ideal neighborhood of the current target node in the current relationship, and determining the difference between the ideal neighborhood and the actual neighborhood of the current target node in the current relationship includes:
[0028] Performing a weighted average on the target feature vector h' of the target node v v and the feature embedding vectors of each neighbor node of the target node v in the relationship r to obtain the neighborhood embedding vector of the target node v in the relationship r, and the expression is:
[0029]
[0030] where represents the neighborhood embedding vector of the target node v in the relationship r and serves as the actual neighborhood of the target node v in the relationship r; N v,r represents the set of neighbor nodes of the target node v in the relationship r; |N v,r | represents the number of neighbor nodes in N v,r ; represents the feature embedding vector of the neighbor node u of the target node v in the relationship r;
[0031] Set the neighborhood transfer vector of the target node v in the relationship r, and define the ideal neighborhood of the target node v in the relationship r Its expression is:
[0032]
[0033] Among them, z v,r represents the neighborhood transfer vector of the target node v in the relationship r; h′ v represents the target feature vector of the target node v;
[0034] The ideal neighborhood of the target node v in the relationship r and the actual neighborhood The difference C v,r , and its expression is:
[0035]
[0036] Among them, C v,r represents the difference between the ideal neighborhood and the actual neighborhood of the target node v in the relationship r.
[0037] Preferably, by using scaling and transfer transformations, minimizing the difference between the ideal neighborhood and the actual neighborhood of the current target node in the current relationship, obtaining the neighborhood transfer vector of the current target node in the current relationship, and successively obtaining the neighborhood transfer vectors of the current target node in each relationship includes:
[0038] In the l-th layer network, use scaling and transfer transformations to obtain the neighborhood transfer vector of the target node v in the relationship r in the l-th layer network Its expression is:
[0039]
[0040] Among them, represents the scaling vector of the target node v in the relationship r in the l-th layer network; represents the transfer vector of the target node v in the relationship r in the l-th layer network; Set The dimension of is the same as that of ; τ represents the unit vector; represents the sum of the neighborhood transfer vectors of all target nodes in the relationship r in the l-th layer network; represents the learnable weight of the l-th layer network;
[0041] According to the neighborhood transfer vectors of the target node v in the relationship r in each layer network, minimize the ideal neighborhood of the target node v in the relationship r and the actual neighborhood The difference C v,r, obtain the neighborhood transfer vector z of the target node v in relation r v,r .
[0042] Preferably, the expression of the complete complementary feature vector F[v] of the target node v is:
[0043]
[0044] where h′ v represents the target feature vector of the target node v; z v,r represents the neighborhood transfer vector of the target node v in relation r; M v represents the set of relationship types corresponding to the target node; hu represents the target feature vector of the neighbor node u of the target node v in relation r; N v,r represents the set of neighbor nodes of the target node v in relation r; represents the sum of the target feature vectors of each neighbor node in the set of neighbor nodes of the target node v in relation r; represents the difference between the neighborhood transfer vector of the target node v in relation r and the sum of the target feature vectors of each neighbor node in the set of neighbor nodes.
[0045] Preferably, the screening of the top k neighbor nodes of the current target node according to the magnitude order of the attribute feature similarities between the current target node and its respective neighbor nodes includes:
[0046] In the neighborhood of the target node v, determine the set of neighbor nodes N v of the target node v, and based on the cosine similarity calculation formula, determine the attribute feature similarity Γ u,v between the target node v and any one of its neighbor nodes u, and its expression is:
[0047]
[0048] where f u represents the projection feature vector of the neighbor node u; f v represents the projection feature vector of the target node v; e represents the edge between the neighbor node u and the target node v;
[0049] Successively obtain the attribute feature similarities between the target node v and each neighbor node in its set of neighbor nodes; sort the attribute feature similarities between the target node v and each neighbor node in its set of neighbor nodes from largest to smallest, select the neighbor nodes corresponding to the top k attribute feature similarities of the target node v as the top k neighbor nodes of the target node v, and construct the set of the top k neighbor nodes N v,k ={u1, u2,..., u k}.
[0050] Preferably, the weighted average of the projection feature vectors of the current target node and its first k neighbor nodes to obtain the preliminary complemented feature vector of the current target node includes:
[0051] Set the weights corresponding to each neighbor node for each neighbor node in the set of the first k neighbor nodes of the target node v; perform a weighted average on the projection feature vectors of the first k neighbor nodes of the target node v and their corresponding weights to obtain the aggregated feature embedding value of the current target node Its expression is:
[0052]
[0053] where u a represents the a-th neighbor node in the set N v,k of the first k neighbor nodes of the target node v; represents the weight corresponding to the a-th neighbor node in the set N v,k of the first k neighbor nodes of the target node v; represents the projection feature vector of the a-th neighbor node in the set N v,k of the first k neighbor nodes of the target node v; k represents the number of neighbor nodes in the set N v,k of the first k neighbor nodes of the target node v;
[0054] Perform a weighted average on the aggregated feature embedding value of the target node v and its projection feature vector f v to obtain the preliminary complemented feature vector of the target node v Its expression is:
[0055]
[0056] where represents the aggregated feature embedding value of the target node v; f v represents the projection feature vector of the target node v.
[0057] Preferably, learning the similarity between nodes in the heterogeneous graph based on the preliminary complemented feature vectors of each target node and the projection feature vectors of each neighbor node of each target node, and obtaining the relational adjacency graph of each relationship in the heterogeneous graph; aggregating the initial adjacency matrix of each relationship with the relational adjacency graph to obtain the target adjacency matrix of each relationship, so as to obtain the target adjacency matrix of the current target node in the current relationship includes:
[0058] In S31 - S35, the tail nodes t m and t n are target nodes; use the preliminary complemented feature vectors of each target node as f* denoted as the preliminary complemented feature vector of the target node v denoted by f v ;
[0059] S31: For a relationship r in the heterogeneous graph, through metric learning, obtain the first-order feature similarity between node features, and its expression is:
[0060]
[0061] where denotes the first-order feature similarity between node and node η; Y represents the number of heads; denotes the corresponding weight when the number of heads is y under the relationship r; denotes the projected feature vector / preliminary complemented feature vector of node ; f η denotes the projected feature vector / preliminary complemented feature vector of node η; if node is the target node, then denotes the preliminary complemented feature vector of node ; if node ζ is not the target node, then denotes the projected feature vector of node ; ⊙ represents the Hadamard product;
[0062] According to the first-order feature similarity obtain the similarity graph whose expression is:
[0063]
[0064] where ε F ∈[0, 1] denotes the sparsity parameter that controls the similarity graph ;
[0065] S32: Based on the relationship r, calculate the feature similarity graph between different types of nodes, that is, the similarity graph j between the source node s m and the tail node t whose expression is:
[0066]
[0067] where denotes the first-order feature similarity between the source node s j and the tail node t m ; ε FB denotes the sparsity parameter that controls the similarity graph ;
[0068] Based on the relationship r, calculate the feature similarity graph between nodes of the same type, and its expression is:
[0069]
[0070] Among them, represents the feature similarity graph between the source node s i and s j ; represents the feature similarity graph between the tail node t m and t n ; represents the first-order feature similarity between the source node s i and s j ; represents the first-order feature similarity between the tail node t m and t n ; f i and f j respectively represent the projected feature vectors of the source node s i and s j ; f m and f n respectively represent the preliminary complemented feature vectors of the tail node t m and t n ; ε FS and ε FT respectively represent the sparsity parameters of the control similarity graph ;
[0071] S33: Through the topological structure propagation of the heterogeneous graph, based on the feature similarity graph between the source node s i and s j and the feature similarity graph between the tail node t m and t n , generate a potential propagation graph, and its expression is:
[0072]
[0073] Among them, respectively represent corresponding propagation graph; A r represents the initial adjacency matrix of the relationship r;
[0074] S34: Adopt channel attention to fuse the feature similarity graph to generate the relationship adjacency graph of the relationship r, and its expression is:
[0075]
[0076] Among them, represents the weight corresponding to the relationship r;
[0077] S35: Aggregate the relational adjacency graph of relationship r with the initial adjacency matrix A of relationship r r to obtain the target adjacency matrix of relationship r, and its expression is:
[0078]
[0079] where softmax(·) represents normalizing the w ψ,r weights; thus, the target adjacency matrix of target node v in relationship r can be obtained as A′ v,r .
[0080] Preferably, calculating the relationship strength between nodes in the heterogeneous graph based on the complete complement feature vectors of each target node and the target feature vectors of each neighbor node of each target node, and aggregating using the attention mechanism to obtain the context representation of each target node, and performing residual connection to obtain the final node representation of each target node includes:
[0081] In S51 - S54, the target feature vectors h′ * of each neighbor node of each target node are all represented by F[*], that is, the target feature vector h′ u of neighbor node u of target node v is represented by F[u];
[0082] S51: For the relationship triple where the target node is the tail node t, and the neighbor node of the tail node t is the source node s, calculate the relationship strength between the source node s and the tail node t to obtain the attention value ATT - head(s, e, t) of the relationship, and its expression is:
[0083] K(s) = K - Linearφ(s)(F[s])
[0084] Q(t) = Q - Linearφ(t)(F[t])
[0085]
[0086] where φ(s) represents the type of source node s; φ(t) represents the type of tail node t, and the tail node t is the target node; represents the type of connection relationship e between the source node s and the tail node t; K(s) represents the relationship strength between the source nodes s; Q(t) represents the relationship strength between the tail nodes t; F[s] represents the target feature vector of the source node s; F[t] represents the complete complement feature vector of the tail node t; Linear φ (*) represents linear projection, projecting the dimension of the source node s or the tail node t from R dbecome d represents the vector dimension, δ represents the number of attention heads, and d / δ represents the vector dimension of each head; represents the parameter for adjusting the attention value; represents the weight corresponding to ATT-head(s, e, t);
[0087] S52: Connect the δ attention heads together to obtain the attention vector for each node pair, and its expression is:
[0088]
[0089] wherein, the attention distribution of the source node s adjacent to the tail node t is normalized using the Softmax function; s ∈ N(t) means that the source node s belongs to the neighbor node set of the tail node t;
[0090] S53: Combine the attention vector of each node pair with the message vector from adjacent nodes to obtain the context representation of each target node, and its expression is:
[0091]
[0092] wherein, the tail node t is the target node; represents the context representation of the tail node t in the ξ-th layer of relation stacking; represents the message vector from adjacent nodes; represents MSG-head γ (s, e, t) corresponding weight;
[0093] S54: Map the context representation of each target node to the distribution of the node type corresponding to each target node and perform a residual connection to obtain the final node representation of each target node, and its expression is:
[0094]
[0095] wherein, the tail node t is the target node; H (ξ) [t] represents the final node representation of the tail node t in the ξ-th layer of relation stacking; H (ξ-1) [t] represents the final node representation of the tail node t in the (ξ - 1)-th layer of relation stacking; ξ represents the number of layers of relation stacking; A-Linear φ(t) represents a linear projection.
[0096] The above technical solutions of the present invention have the following beneficial effects compared with the prior art:
[0097] (1) A heterogeneous graph embedding method based on information completion according to the present invention selects the top k neighbor nodes corresponding to each target node based on the attribute similarity between each target node and each neighbor node within the neighborhood of each target node, and then aggregates the attribute information of the top k neighbor nodes of each target node to complete the feature of each target node with missing attributes, obtaining the preliminary completed feature vectors of each target node. Since in a heterogeneous graph, nodes usually carry rich feature information that can be used to learn the missing attributes of other nodes, it is possible to use the nodes most similar to each target node with missing attributes to improve the quality of each node's data, thus initially solving the problem of node attribute missing and laying a foundation for the subsequent model to better learn the feature representation of nodes.
[0098] (2) A heterogeneous graph embedding method based on information completion according to the present invention multiplies the preliminary completed feature vectors of each target node, the projected feature vectors of each neighbor node of each target node, and their corresponding adaptive weights based on an adaptive weight mechanism to obtain the adaptive feature vectors of each target node and the adaptive feature vectors of each neighbor node of each target node, enabling the model to better learn from the node data. At the same time, based on the target adjacency matrix of various relationships corresponding to each target node, a normalization operation is performed on the adaptive feature vectors of each neighbor node of each target node to obtain the feature embedding vectors of each neighbor node of the target node in various relationships. The target adjacency matrix analyzes the connection information between nodes, mines the feature similarity, and applies a channel attention mechanism, which helps the model to more comprehensively understand the relationship between nodes and improve the accuracy of obtaining the neighborhood information of each target node in the future. In addition, this method also considers neighborhood transfer and improves the neighborhood information of the target node by minimizing the difference between the ideal neighborhood and the actual neighborhood of the target node. Since the neighborhood information of the target head node is relatively rich and not easily affected by information loss and information error, the target head node can be used as the learning object of the target tail node. Therefore, by learning the neighborhood transfer vector of the target head node, the ideal neighborhood information of the target tail node can be obtained, realizing the complete completion operation of the target node, solving the long-tail problem, enabling the target node with fewer relationship connections to learn more information, and thus obtaining a more complete node feature representation, further improving the performance of downstream tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to specific embodiments of the present invention in conjunction with the drawings, where:
[0100] Figure 1 is a flowchart of a heterogeneous graph embedding method based on information completion provided by the present invention;
[0101] Figure 2 It is the schematic diagram of the relationship graph generator;
[0102] Figure 3 It is the schematic diagram of the relationship-level aggregator;
[0103] Figure 4 It is the diagram of the HGNN-I C model;
[0104] Figure 5 It is the diagram of the link prediction result;
[0105] Figure 6 It is the diagram of the visualization result; among which, Figure 6 (a) in it is the result diagram of selecting 1500 nodes on the ACM dataset, Figure 6 (b) in it is the result diagram of selecting 500 nodes on the IMDB dataset, Figure 6 (c) in it is the result diagram of selecting 1500 nodes on the DBLP dataset;
[0106] Figure 7 It is the result diagram of the ablation experiment; among which, Figure 7 (a) in it is the comparison result of the Macro-F1 values on the three datasets of ACM, IMDB, and DBLP, Figure 7 (b) in it is the comparison result of the Micro-F1 values on the three datasets of ACM, IMDB, and DBLP. Specific implementation manner
[0107] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited do not limit the present invention.
[0108] Referring to Figure 1 as shown, Figure 1 it is the flowchart of a heterogeneous graph embedding method based on information completion provided by the present invention; specifically including:
[0109] S1: Based on the heterogeneous graph, determine and construct the target node set; map the initial feature vectors of all nodes in the heterogeneous graph into a unified shared semantic space to obtain the projected feature vectors of all nodes; in this process, the node feature projection module is used to unify the node feature vectors to ensure that nodes of various types have a consistent representation;
[0110] Among them, mapping the initial feature vector of node θ into a unified shared semantic space to obtain the projected feature vector of node u, and its expression is:
[0111]
[0112] Among them, Denote the projected feature vector of node θ; Denote the initial feature vector of node θ; φ(θ) denotes the type of node θ; σ(·) denotes the non-linear activation function; Denote the projection matrix corresponding to the nodes of type φ(θ); b φ (θ) ∈ R 1×d Denote the bias parameter corresponding to the nodes of type φ(θ);
[0113] S2: Sort according to the magnitude of the attribute feature similarity between the current target node and its respective neighbor nodes, and select the top k neighbor nodes of the current target node; perform a weighted average on the projected feature vectors of the current target node and its top k neighbor nodes to obtain the preliminary complemented feature vector of the current target node, including:
[0114] S21: Based on each neighbor node within the neighborhood of the current target node, combine the cosine similarity calculation formula to obtain the attribute feature similarity between the current target node and its respective neighbor nodes; sort the respective attribute feature similarities corresponding to the current target node from largest to smallest, select the top k attribute feature similarities of the current target node, and obtain the top k neighbor nodes of the current target node, that is:
[0115] Within the neighborhood of the target node v, determine the set of neighbor nodes N of the target node v v , and based on the cosine similarity calculation formula, determine the attribute feature similarity Γ between the target node v and any one of its neighbor nodes u u,v , and its expression is:
[0116]
[0117] where, f u Denote the projected feature vector of neighbor node u; f v Denote the projected feature vector of target node v; e represents the edge between neighbor node u and target node v;
[0118] Successively obtain the attribute feature similarities between the target node v and each neighbor node in its neighbor node set; sort the attribute feature similarities between the target node v and each neighbor node in its neighbor node set from largest to smallest, select the neighbor nodes corresponding to the top k attribute feature similarities of the target node v as the top k neighbor nodes of the target node v, and construct the set N of the top k neighbor nodes of the target node v v,k ={u1, u2,..., u k};
[0119] S22: Perform a weighted average on the projection feature vectors of the top k neighbor nodes of the current target node and their corresponding weights to obtain the aggregated feature embedding value of the current target node, and then perform a weighted average with the projection feature vector of the current target node to obtain the preliminary complemented feature vector of the current target node, that is:
[0120] Set the weights corresponding to each neighbor node for each neighbor node in the set of the top k neighbor nodes of the target node v; perform a weighted average on the projection feature vectors of the top k neighbor nodes of the target node v and their corresponding weights to obtain the aggregated feature embedding value of the current target node Its expression is:
[0121]
[0122] where, u a represents the a-th neighbor node in the set N of the top k neighbor nodes of the target node v v,k ; represents the weight corresponding to the a-th neighbor node in the set N of the top k neighbor nodes of the target node v v,k ; represents the projection feature vector of the a-th neighbor node in the set N of the top k neighbor nodes of the target node v; k represents the number of neighbor nodes in the set N of the top k neighbor nodes of the target node v v,k ; v,k ;
[0123] Perform a weighted average on the aggregated feature embedding value of the target node v and its projection feature vector f v to obtain the preliminary complemented feature vector of the target node v Its expression is:
[0124]
[0125] where, represents the aggregated feature embedding value of the target node v; f v represents the projection feature vector of the target node v;
[0126] In a heterogeneous graph, nodes usually carry rich feature information, which can be used to learn the missing attributes of other nodes. However, a heterogeneous graph usually contains a large number of nodes. Directly using all nodes to predict the missing attributes of the target node is likely to lead to excessive computational costs and cause excessive consumption of memory and time. To avoid these problems, in S2 of the present invention, through weighted aggregation of the projected node features by the k-nearest neighbor (KNN method), a preliminary completion operation of the attribute features of the target node with missing attributes is implemented. This process can effectively select the nodes most similar to the target node, thereby generating the feature embedding of the target node. The feature aggregation operation effectively aggregates the attribute information of the similar neighbor nodes of the target node, achieving a preliminary completion of the attribute features of the target node v.
[0127] S3: Based on the preliminary completion feature vectors of each target node and the projected feature vectors of each neighbor node of each target node, learn the similarity between each node and obtain the relationship adjacency graph of various relationships; aggregate the initial adjacency matrix of various relationships with the relationship adjacency graph to obtain the target adjacency matrix of various relationships, thereby obtaining the target adjacency matrix corresponding to the current target node for the current relationship, including:
[0128] In a heterogeneous graph, the more similar two node features are in the feature vector space, the closer the nodes are to each other in the embedding space, the higher the probability of belonging to the same type, and the higher the probability of there being an edge between these two similar nodes. Based on the topological structure of the graph, the present invention introduces a relationship graph generator and uses the relationships between nodes to learn the similarity between node features. The principle is as Figure 2 shown; Figure 2 contains only one kind of relationship r represents a type of relationship between nodes, and this relationship covers the following three relationship triples <s i , e1, t m >, <s j , e2, t m >, <s j , e3, t n >;
[0129] In S31 - S35, the tail nodes t m and t n are target nodes; represent the preliminary completion feature vectors of each target node all with f * , that is, the preliminary completion feature vector of the target node v is represented by f v ;
[0130] S31: For a relationship r in the heterogeneous graph, through metric learning, obtain the first-order feature similarity between node features, and its expression is:
[0131]
[0132] Among them, represents the first-order feature similarity between node and node η; Y represents the number of heads; represents the weight corresponding to the number of heads y under the relationship r; represents the projected feature vector / preliminary completion feature vector of node ; f η represents the projected feature vector / preliminary completion feature vector of node η; if node is the target node, then represents the preliminary completion feature vector of node ; if node ζ is not the target node, then represents the projected feature vector of node ; ⊙ represents the Hadamard product;
[0133] According to the first-order feature similarity to obtain the similarity graph Its expression is:
[0134]
[0135] Among them, ε F ∈ [0, 1] represents the sparsity parameter that controls the similarity graph ; the closer the value of ε F is to 1, the sparser the heterogeneous graph, which means that very similar nodes are retained; if ε F is small, the heterogeneous graph will contain more non-zero elements, reflecting more similar nodes;
[0136] S32: As Figure 2 shown, according to the above steps, based on the relationship r, replacing node node η with the source node s j , the tail node t m , respectively, the similarity graph between the source node s j and the tail node t m can be obtained
[0137] For the same type of relationship calculate the feature similarity graph between different types of nodes, that is, the similarity graph between the source node s j and the tail node t m Its expression is: Its expression is:
[0138]
[0139] Among them, Denote the source node as s j and the tail node as t m The first-order feature similarity between them; ε FB Denote the sparsity parameter of the similarity graph ;
[0140] Two nodes with similar features may have similar neighbors. Therefore, for the same type of relationship it is necessary to study its higher-order similarity, that is, to find the neighbors of nodes with similar features through the structure of the heterogeneous graph; among them, nodes of the same type can be the source node or the tail node;
[0141] Based on the relationship r, calculate the feature similarity graph between nodes of the same type, and its expression is:
[0142]
[0143] where is calculated in a similar way to ; Denote the feature similarity graph between the source node s i and s j ; Denote the feature similarity graph between the tail node t m and t n ; Denote the first-order feature similarity between the source node s i and s j ; Denote the first-order feature similarity between the tail node t m and t n ; f i 、f j Denote the projected feature vectors of the source node s i 、s j respectively; f m 、f n Denote the preliminary complemented feature vectors of the tail node t m 、t n respectively; ε FS 、ε FT Denote the sparsity parameters of the similarity graph respectively;
[0144] S33: Through the topological structure propagation of the heterogeneous graph, based on the feature similarity graph between the source node s i and s j and the feature similarity graph between the tail node t m and t n generate the potential propagation graph, and its expression is:
[0145]
[0146] Among them, respectively represent the corresponding propagation graph; A r represents the initial adjacency matrix of relationship r;
[0147] S34: Use channel attention to fuse the feature similarity graph to generate the relationship adjacency graph of relationship r whose expression is:
[0148]
[0149] Among them, represents the weight corresponding to relationship r;
[0150] S35: To more comprehensively consider the influence of relationships on the graph and further improve the model performance, the relationship adjacency graph of relationship r and the initial adjacency matrix A of relationship r r are aggregated to obtain the target adjacency matrix of relationship r, whose expression is:
[0151]
[0152] Among them, softmax(·) represents normalizing the w ψ,r weights;
[0153] In summary, the target adjacency matrix of the target node v corresponding to relationship r is A′ v,r ;
[0154] In S3, by analyzing the connection information between nodes, mining feature similarities, and applying the channel attention mechanism, a new relationship adjacency matrix is learned; the complete graph structure information helps to more comprehensively understand the relationships between nodes, including the complex information between different node types;
[0155] S4: Based on the new relationship graph learned in S3, perform neighborhood transfer to complete the operation of fully complementing the attribute features of the target node, including:
[0156] S41: Based on all the nodes in the heterogeneous graph, set the adaptive weight matrix and the number of training times corresponding to each node in the initial round. Based on the adaptive feature vector of the current node in the previous round and the adaptive weight matrix corresponding to the current node in the current round, obtain the adaptive feature vector of the current node in the current round; until the preset number of training times is reached, obtain the target feature vectors of each node; where, when the training round is 1, the adaptive feature vector of the current target node in the previous round is its preliminary completion feature vector, and the adaptive feature vector of each neighbor node of the current target node in the previous round is its projection feature vector, including:
[0157] Based on the adaptive feature vector of the target node v in the (p - 1)th round and the adaptive weight matrix corresponding to the target node v in the pth round Obtain the adaptive feature vector of the target node v in the pth round, and its expression is:
[0158]
[0159] Where represents the adaptive feature vector of the target node v in the pth round; represents the adaptive feature vector of the target node v in the (p - 1)th round. In the initial round, that is, when p = 1, represents the adaptive weight matrix corresponding to the target node v in the pth round;
[0160] Based on the adaptive feature vector of any neighbor node u of the target node v in the (p - 1)th round and the adaptive weight matrix corresponding to the neighbor node u of the target node v in the pth round Obtain the adaptive feature vector of the neighbor node u of the target node v in the pth round, and its expression is:
[0161]
[0162] Where represents the adaptive feature vector of the neighbor node u of the target node v in the pth round; represents the adaptive feature vector of the neighbor node u of the target node v in the (p - 1)th round. In the initial round, that is, when p = 1, represents the adaptive weight matrix corresponding to the neighbor node u of the target node v in the pth round;
[0163] Until the preset number of training times is reached, obtain the target feature vector h′ of the target node v v , the target feature vector h′ of the neighbor node u of the target node v u ;
[0164] S42: Based on the target adjacency matrix of the current target node in the current relationship, perform a normalization operation on the target feature vectors of each neighbor node of the current target node in the current relationship to obtain the feature embedding vectors of each neighbor node of the current target node in the current relationship, including:
[0165] Based on the target adjacency matrix A′ of the target node v in the relationship r v,r , perform a normalization operation on the target feature vector hu of the neighbor node u of the target node v in the relationship r to obtain the feature embedding vector of the neighbor node u of the target node v in the relationship r Its expression is:
[0166]
[0167] Among them, h′ u represents the target feature vector of the neighbor node u of the target node v; represents the normalization coefficient of the target node v in the relationship r, and its expression is: D r represents the degree matrix of A′ v,r ;
[0168] Among them, by adopting an adaptive weight allocation mechanism, the model can better learn from node data;
[0169] S43: Perform a weighted average on the target feature vector of the current target node and the feature embedding vectors of each neighbor node of the current target node in the current relationship to obtain the neighborhood embedding vector of the current target node in the current relationship as the actual neighborhood of the current target node in the current relationship; set the neighborhood transfer vector of the current target node in the current relationship, define the ideal neighborhood of the current target node in the current relationship, and determine the difference between the ideal neighborhood and the actual neighborhood of the current target node in the current relationship, including:
[0170] Perform a weighted average on the target feature vector hj of the target node v and the feature embedding vectors of each neighbor node of the target node v in the relationship r to obtain the neighborhood embedding vector of the target node v in the relationship r, and its expression is:
[0171]
[0172] Among them, represents the feature embedding vector of the target node v in the relationship r and serves as the actual neighborhood of the target node v in the relationship r; N v,r represents the set of neighbor nodes of the target node v in the relationship r; |N v,r | represents the number of neighbor nodes in N v,r ; Represents the feature embedding vector of neighbor node u of target node v in relation r;
[0173] Based on the relation transformation model, use the neighborhood transfer operation to model the relationship between target node v and its neighborhood, and its expression is:
[0174]
[0175] where, z v,r Represents the neighborhood transfer vector of target node v in relation r, used to model the missing information in the neighborhood; h′ v Represents the target feature vector of target node v;
[0176] Set the neighborhood transfer vector of target node v in relation r, and define the ideal neighborhood of target node v in relation r Its expression is:
[0177]
[0178] where, Represents the neighborhood transfer vector of target node v in relation r; h′ v Represents the target feature vector of target node v;
[0179] The ideal neighborhood of target node v in relation r and the actual neighborhood The difference C v,r , and its expression is:
[0180]
[0181] where, C v,r Represents the difference between the ideal neighborhood and the actual neighborhood of target node v in relation r;
[0182] If the target node is the head node, since their neighborhood information in the heterogeneous graph is relatively complete, the difference C v,r should be approximately 0; and if the target node is the tail node, their difference is:
[0183] The neighborhood information of the head node is relatively rich and is not easily affected by information loss and information errors; although their neighbor structure is not the optimal neighborhood, it can be used as an object for the tail node to learn; therefore, by learning the neighborhood transfer vector of the head node, the ideal neighborhood information of the tail node can be obtained;
[0184] S44: Using scaling and transfer transformations, minimize the difference between the ideal neighborhood and the actual neighborhood of the current target node in the current relationship to obtain the neighborhood transfer vector of the current target node in the current relationship, and sequentially obtain the neighborhood transfer vectors of the current target node in each relationship, including:
[0185] In the l-th layer network, use scaling and transfer transformations to obtain the neighborhood transfer vector of the target node v in the l-th layer network in relationship r Its expression is:
[0186]
[0187] Where, represents the scaling vector of the target node v in the l-th layer network in relationship r; represents the transfer vector of the target node v in the l-th layer network in relationship r; set to have the same dimension as for the convenience of implementing scaling and transfer in an element-wise manner; τ represents the unit vector; represents the sum of the neighborhood transfer vectors of all target nodes in the l-th layer network in relationship r; represents the learnable weight of the l-th layer network;
[0188] According to the neighborhood transfer vectors of the target node v in relationship r in each layer network, minimize the difference C between the ideal neighborhood and the actual neighborhood v,r of the target node v in relationship r to obtain the neighborhood transfer vector z v,r of the target node v in relationship r;
[0189] S45: The expression of the complete complement feature vector F[v] of the target node v is:
[0190]
[0191] Where, h′ v represents the target feature vector of the target node v; z v,r represents the neighborhood transfer vector of the target node v in relationship r; M v represents the set of relationship types corresponding to the target node v; hu represents the target feature vector of the neighbor node u of the target node v in relationship r; N v,r represents the set of neighbor nodes of the target node v in relationship r; represents the sum of the target feature vectors of each neighbor node in the set of neighbor nodes of the target node v in relationship r; represents the difference between the neighborhood transfer vector of the target node v in relationship r and the sum of the target feature vectors of each neighbor node in the set of neighbor nodes;
[0192] Among them, in order to improve the efficiency of the model, a neighborhood transfer vector is set, which can represent the simple interaction relationship between nodes; in a heterogeneous graph, different types of nodes have different relationship structures, and the neighborhood transfer vector needs to reflect the different context information of nodes, that is, each node itself and unique neighbor information; through calculation, the neighborhood transfer vector is obtained, the neighborhood information of the tail node is obtained, the information gap between the head node and the tail node is balanced, so as to obtain the feature representation of the tail node, and then the feature representations of each node can be obtained;
[0193] S5: Based on the complete complemented feature vectors of each target node and the target feature vectors of each neighbor node of each target node, calculate the relationship strength between nodes in the heterogeneous graph, and use the attention mechanism to aggregate to obtain the context representation of each target node, and perform residual connection to obtain the final node representation of each target node, including:
[0194] For a target node, information from different types of neighbor objects may have different effects; therefore, in order to effectively utilize different relationships, such as Figure 3 as shown, a new node feature representation F is obtained by neighborhood transfer, the importance weights of different types of neighbor objects are learned, and the information of the source node is aggregated using the attention mechanism to obtain the context representation of the target node; in Figure 3 there are two different relationships, which are respectively and the relationship triples <s, e1, t>, <h, e2, t> correspond to them respectively;
[0195] In S51 - S54, the target feature vectors h' * of each neighbor node of each target node are all represented by F[*], that is, the target feature vector h' u of the neighbor node u of the target node v is represented by F[u];
[0196] S51: For the relationship triple The target node is the tail node t, and the neighbor node of the tail node t is the source node s. Calculate the relationship strength between the source node s and the tail node t to obtain the attention value ATT-head(s, e, t) of the relationship, and its expression is:
[0197] K(s) = K-Linear φ(s) (F[s])
[0198] Q(t) = Q-Linear φ(t) (F[t])
[0199]
[0200] Among them, φ(s) represents the type of the source node s; φ(t) represents the type of the tail node t, and the tail node t is the target node; represents the type e of the connection relationship between the source node s and the tail node t; K(s) represents the relationship strength between the source nodes s; Q(t) represents the relationship strength between the tail nodes t; F[s] represents the target feature vector of the source node s; F[t] represents the complete complementary feature vector of the tail node t; Linear φ(*) represents a linear projection that changes the dimension of the source node s or the tail node t from R d to d represents the vector dimension, δ represents the number of attention heads, and d / δ represents the vector dimension of each head; represents a parameter for adjusting the attention value to adapt to different relationships; represents the weight corresponding to ATT-head(s, e, t); since there are multiple relationships between node pairs in the heterogeneous graph, can adapt to different edges;
[0201] S52: To more comprehensively capture the information of different types of neighbor nodes, connect δ attention heads together to obtain the attention vector of each node pair, and its expression is:
[0202]
[0203] Among them, use the Softmax function to normalize the attention distribution of the source node s adjacent to the tail node t; s ∈ N(t) means that the source node s belongs to the neighbor node set of the tail node t; for each target node, perform Softmax processing on the attention values of all its neighbor nodes to ensure that the attention distribution of the neighbor nodes satisfies normalization, which is convenient for effectively performing weighted averaging on different neighbor nodes;
[0204] S53: Combine the attention vector of each node pair with the message vector from adjacent nodes to obtain the context representation of each target node, and its expression is:
[0205]
[0206] Among them, the tail node t is the target node; represents the context representation of the tail node t in the ξ-th layer relationship stack; represents the message vector from adjacent nodes; represents the weight corresponding to MSG-head γ (s, e, t); the attention vector weights and averages the information of different types of neighbor nodes to ensure more attention to the neighbor nodes that have the greatest impact on the target node;
[0207] S54: Map the context representation of each target node to the distribution of the node type corresponding to each target node, and perform a residual connection to obtain the final node representation of each target node. The expression is as follows:
[0208]
[0209] where the tail node t is the target node; H (ξ) [t] represents the final node representation of the tail node t in the ξ-th layer of relation stacking; H (ξ-1) [t] represents the final node representation of the tail node t in the (ξ - 1)-th layer of relation stacking; ξ represents the number of layers of relation stacking; A-Linear φ(t) represents a linear projection.
[0210] Based on the above steps S1 - S5, in the process of attribute completion for the heterogeneous graph, the target nodes are the nodes in the heterogeneous graph with incomplete attribute data that need to be completed; after completing the attribute completion for the target nodes in the heterogeneous graph, through the multi-layer stacked relation-level aggregation module, the node representations of the entire heterogeneous graph can be obtained, which can then be used for downstream tasks. In the downstream tasks, the target nodes are the nodes that need to perform node classification and clustering experiments;
[0211] To solve the problem in the prior art that for nodes with missing attributes, only simple filling operations are taken, which easily introduce noise and lead to performance loss of the model; the method of using the adjacency matrix for relation completion cannot well solve the long-tail problem existing in the heterogeneous graph, and still leads to a decline in the model performance, thereby reducing the performance of downstream tasks, a feature aggregation and neighborhood transfer module is introduced. Based on steps S1 - S5, a heterogeneous graph embedding model based on information completion (HGNN-IC model) is constructed, as specifically Figure 4 shown; the HGNN-IC model takes the original heterogeneous graph as input. First, through converting and projecting the node features and weighted aggregating the k-nearest neighbors, the preliminary completion operation of the attribute features of the target nodes with missing attributes is realized; second, use the relation graph generator and the topological information between nodes to update the graph; third, on the basis of the learned new graph, perform neighborhood transfer to realize the complete completion operation of the attribute features of the target nodes; finally, the relation-level aggregation module aggregates the previous features to obtain the node representations of the entire graph; in order to demonstrate the excellent performance of the HGNN-IC model provided by the present invention, extensive experiments are carried out on three typical heterogeneous information network datasets and detailed comparisons are made with multiple baseline models. The HGNN-IC model is significantly better than the state-of-the-art models in downstream tasks such as node classification and clustering.
[0212] Experiments were conducted using the datasets ACM, IMDB, and DBLP. Table 1 summarizes the specific data of the three datasets. To facilitate the study of this method, nodes were randomly selected at the same ratio, and their attributes were all set to empty to construct nodes with missing attributes, so as to more accurately evaluate the performance of this method.
[0213] Table 1 Datasets
[0214]
[0215] The ACM dataset comes from citation network data and covers three object types: author (A), paper (P), and subject (S). Papers (P) are divided into three research fields: database, data mining, and wireless communication, which are used as labels for the papers. The features of the papers are represented by a bag of words composed of their keywords, while the feature vectors of the authors are the average values of the features of the papers they are connected to. The features of the subjects use one-hot encoding.
[0216] IMDB has been preprocessed and contains three object types: movie (M), actor (A), and director (D). Movies (M) include three types of labels: action movies, comedies, and dramas. The features of the movies are represented by a bag of words composed of their movie plot keywords, while the features of the actors and directors are the average values of the corresponding movie features.
[0217] The DBLP dataset contains four object types: paper (P), author (A), conference (C), and term (T). Authors (A) are divided into four categories according to their research fields: database, data mining, artificial intelligence, and information retrieval. The features of the papers are represented by a bag of words of the paper titles, the features of the authors are represented by a bag of words of the keywords extracted from the papers they published, the features of the terms use pre-trained word vectors, and the features of the conferences use one-hot vectors.
[0218] Experiment 1: Node Classification
[0219] This invention focuses on the paper nodes in the ACM dataset, the author nodes in the DBLP dataset, and the movie nodes in the IMDB dataset to evaluate the performance of the HGNN-IC model in the node classification task. In the experiment, this invention adopted different training ratios, specifically 20%, 40%, 60%, and 80%. The target nodes were divided according to these ratios, and the remaining nodes were used as the validation set and the test set. This invention adopted Macro-F1 and Micro-F1 as evaluation metrics to evaluate the performance of different models in the node classification task. The results of the node classification experiment are shown in Table 2.
[0220] Table 2 Results of Node Classification Experiment
[0221]
[0222] Overall, the F1 scores of the HGNN-IC model on the three datasets are higher than those of other models, and its performance has been significantly improved. The HGNN-AC model is optimized for the problem of missing attributes. Compared with the Metapath2vec model, its performance on the ACM, IMDB, and DBLP datasets has been improved by 22%, 26%, and 8.8% respectively, showing excellent performance. Although the Tail-GNN model makes up for the differences between the head and tail nodes under the guidance of the ideal neighborhood of the target head node, compared with other heterogeneous graph embedding models, the Tail-GNN model fails to utilize the information of high-order neighbor nodes, and its embedding effect is limited. Especially on the ACM dataset, its performance shows a downward trend compared with the HGNN-IC model, with a decrease of 5.2% and 5.1% respectively. Compared with the HRMNN, the HGNN-IC model effectively improves the performance in graph embedding, especially on the IMDB dataset, with an increase of 1.7% and 1.6% respectively.
[0223] Experiment 2: Node Clustering
[0224] The present invention uses the K-Means algorithm for clustering. For the ACM dataset, the paper nodes are divided into a training set (808), a validation set (401), and a test set (2816). For the DBLP dataset, the author nodes are divided into a training set (400), a validation set (400), and a test set (3257). For the IMDB dataset, the movie nodes are divided into a training set (400), a validation set (400), and a test set (3478). The present invention adopts two evaluation metrics, the Normalized Mutual Information (NMI) and the Adjusted Rand Index (ARI), to evaluate the performance of different algorithms in the node clustering task. The results of node clustering are shown in Table 3.
[0225] Table 3 Results of Node Clustering
[0226]
[0227] As can be observed from Table 3, the HGNN-IC model demonstrates superior performance on multiple datasets. Specifically, on the ACM dataset, the NMI value of HGNN-IC reaches 0.68 and the ARI is 0.66. On the IMDB dataset, its NMI and ARI are 0.12 and 0.13 respectively. On the DBLP dataset, they are 0.61 and 0.63 respectively, with the overall performance being better than other models. As homogeneous graph models, GCN and Node2vec do not perform well on heterogeneous graphs, especially when compared with the Metapath2vec model specifically designed for handling heterogeneous graphs. On the ACM dataset, the performance of GCN and Node2vec drops by 14% and 8.9% respectively. On the IMDB dataset, the performance drops reach 30% and 25.5%. Although the Metapath2vec model designs meta-paths to guide random walks for heterogeneous graphs, due to the defects in the definition of its meta-paths and the dependence of random walk behavior on node structures, its performance is still limited to a certain extent. In addition, the existence of tail nodes also affects the performance of Metapath2vec. Other models, such as HGT and HAN, although they aggregate information of nodes and consider the information carried by neighboring nodes and connections, compared with Metapath2vec, their performance has increased by 33.1% and 28%. However, compared with the HGNN-IC model, the performance of HGT and HAN drops by 7.63% and 10.49% respectively, still showing a certain gap.
[0228] Experiment 3: Link Prediction
[0229] When conducting the link prediction experiment, the present invention uses the edges between author nodes and paper nodes in the DBLP dataset to demonstrate the performance of the HGNN-IC model. The present invention regards the connected paper-author pairs as positive node pairs and all unconnected paper-author pairs as negative node pairs. The training set of the present invention contains 13,751 positive node pairs and the same number of randomly sampled negative node pairs are added. In the validation set and the test set, the numbers of positive node pairs and negative node pairs are also kept consistent. In this experiment, two evaluation metrics are used, namely the area under the curve (AUC) and the average precision (AP), and the specific results are as Figure 5 shown.
[0230] From Figure 5It can be found from the experimental results that HRMNN uses a relational graph generator to complete the adjacency matrix, showing obvious advantages in the link prediction experiment, and the performance of this model is also better than other baseline models; HGNN-IC further pays attention to the neighborhood missing situation of the tail node on this basis, uses the neighborhood information of the head node to generate a transfer vector for completion, and realizes improvements of 2.9% and 4.7% respectively in performance; although both the HGNN-IC model and the Tail-GNN model are committed to solving the problem of completing the connection information of the tail node, Tail-GNN only aggregates the information of neighbor nodes when generating neighborhood information, while HGNN-IC adopts a more complex strategy, not only considering the information of neighbor nodes, but also using an attention mechanism to judge the importance of node information, and passing information through edges, so as to generate node representations more accurately; therefore, the HGNN-IC model has a performance improvement of 14.8% and 17.1% compared with the Tail-GNN model.
[0231] Experiment 4: Visualization
[0232] To more intuitively display the performance of the HGNN-IC model, we use the t-SNE algorithm to project the generated embeddings into a two-dimensional space to realize the visualization of node representations; the experimental results are as Figure 6 shown;
[0233] Generally speaking, the HGNN-IC model shows good performance and can generate node representations with a certain degree of discrimination; on the DBLP dataset, the HGNN-IC model performs the best; it can be clearly observed from Figure 6 that there are clear boundary divisions between different types of nodes, and nodes of the same type are closely clustered together to form a clear cluster structure; in the ACM dataset, although there is an overlapping phenomenon between some circular nodes and square nodes, generally speaking, the node distribution is relatively concentrated, and the distance between different groups is also relatively obvious; in the IMDB dataset, the performance of the HGNN-IC model is relatively poor. It can be seen from the figure that triangular nodes cannot be effectively distinguished, and the distance between different category nodes is closer than that in the ACM dataset; this indicates that on the IMDB dataset, the node representations learned by the model still lack sufficient discrimination.
[0234] Experiment 5: Ablation Experiment
[0235] To verify the effectiveness of each component in the HGNN-IC model, we use its variants to conduct ablation studies and use the same parameters for node classification experiments. The results are as Figure 7 shown; the information of the variant models is as follows:
[0236] The HGNN-I CA variant mainly focuses on the completion of attribute information; it uses a feature aggregation module to capture the similarity between nodes and accordingly completes the missing attribute information; subsequently, these complete node attribute information are input into the HRMNN model to form the final node representation.
[0237] The HGNN-I CT variant mainly focuses on the completion of the missing information of the tail node; it combines with the HRMNN model to perform a neighborhood transfer operation to obtain the node embedding value containing the neighborhood information of the tail node.
[0238] The F1 score of the HGNN-I C model is the highest value, indicating that its performance is better than the other two variants; this advantage mainly comes from the effective combination of the feature aggregation module and the neighborhood transfer module, jointly improving the overall performance of the model; by comparing the performance of HGNN-I CT and HGNN-I C, we can find the effectiveness of the similarity strategy considering nodes in the feature aggregation module of the HGNN-I C model, which can reduce the noise propagation caused by missing attributes; on the ACM dataset, the performance of HGNN-I CA is reduced by 3.1% and 3.5% respectively compared with the HGNN-I CT model, which may be due to the fact that there are many tail nodes in the ACM dataset, and simply relying on the completion of attribute information is not sufficient to fully capture the characteristics of nodes, while HGNN-I CT completes the neighborhood information of the tail node through the neighborhood transfer module, thus achieving better performance on the ACM dataset, which further verifies the effectiveness of the neighborhood transfer module in dealing with the tail node problem.
[0239] In the specific embodiments of the present invention, there is also provided a heterogeneous graph embedding system based on information completion, including:
[0240] Projection module: Based on the heterogeneous graph, determine and construct a target node set; map the initial feature vectors of all nodes in the heterogeneous graph to a unified shared semantic space to obtain the projected feature vectors of all nodes;
[0241] Feature aggregation module: Sort according to the size of the attribute feature similarity between the current target node and its respective neighbor nodes, and select the top k neighbor nodes of the current target node; perform a weighted average on the projected feature vectors of the current target node and its top k neighbor nodes to obtain the preliminary completed feature vector of the current target node;
[0242] Relationship graph generator: Based on the preliminary completed feature vectors of each target node and the projected feature vectors of each neighbor node of each target node, learn the similarity between nodes in the heterogeneous graph and obtain the relationship adjacency graph of each relationship in the heterogeneous graph; aggregate the initial adjacency matrix of each relationship with the relationship adjacency graph to obtain the target adjacency matrix of each relationship, so as to obtain the target adjacency matrix of the current target node in the current relationship;
[0243] Neighborhood Transfer Module: Based on all nodes in the heterogeneous graph, set the adaptive weight matrix and the number of training times corresponding to each node in the initial round. Based on the adaptive feature vector of the current node in the previous round and the adaptive weight matrix corresponding to the current node in the current round, obtain the adaptive feature vector of the current node in the current round; until the preset number of training times is reached, obtain the target feature vectors of each node; among them, when the training round is 1, the adaptive feature vector of the current target node in the previous round is its preliminary completion feature vector, and the adaptive feature vector of each neighbor node of the current target node in the previous round is its projection feature vector;
[0244] Based on the target adjacency matrix of the current target node in the current relationship, perform a normalization operation on the target feature vectors of each neighbor node of the current target node in the current relationship to obtain the feature embedding vectors of each neighbor node of the current target node in the current relationship;
[0245] Perform a weighted average on the target feature vector of the current target node and the feature embedding vectors of each neighbor node of the current target node in the current relationship to obtain the neighborhood embedding vector of the current target node in the current relationship as the actual neighborhood of the current target node in the current relationship; set the neighborhood transfer vector of the current target node in the current relationship, define the ideal neighborhood of the current target node in the current relationship, and determine the difference between the ideal neighborhood and the actual neighborhood of the current target node in the current relationship; use scaling and transfer transformations to minimize the difference between the ideal neighborhood and the actual neighborhood of the current target node in the current relationship to obtain the neighborhood transfer vector of the current target node in the current relationship, and sequentially obtain the neighborhood transfer vectors of the current target node in each relationship; based on the target feature vectors of the current target node and its neighbor nodes, and the neighborhood transfer vectors of the current target node in each relationship, obtain the complete completion feature vector of the current target node;
[0246] Relationship-level Aggregator: Based on the complete completion feature vectors of each target node and the target feature vectors of each neighbor node of each target node, calculate the relationship strength between nodes in the heterogeneous graph, and use the attention mechanism to aggregate to obtain the context representation of each target node, and perform a residual connection to obtain the final node representation of each target node.
[0247] Obviously, the above embodiments are only examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the relevant art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. A heterogeneous graph embedding method based on information completion, characterized in that: include: Based on the heterogeneous graph, determine and construct the target node set; Mapping the initial feature vectors of all nodes in the heterogeneous graph into a unified shared semantic space to obtain the projected feature vectors of all nodes; the nodes in the heterogeneous graph are taken from the DBLP dataset; According to the order of the similarity of the attribute features between the current target node and its neighbor nodes, the first k neighbor nodes of the current target node are selected; the projected feature vectors of the current target node and its first k neighbor nodes are weighted averaged to obtain the preliminary completed feature vector of the current target node; Based on the preliminary completed feature vectors of each target node and the projected feature vectors of each neighbor node of each target node, the similarity between nodes in the heterogeneous graph is learned, and the relational adjacency graph of each relationship in the heterogeneous graph is obtained; the initial adjacency matrix of each relationship is aggregated with the relational adjacency graph to obtain the target adjacency matrix of each relationship, thereby obtaining the target adjacency matrix of the current target node in the current relationship; Based on all nodes in the heterogeneous graph, the adaptive weight matrix corresponding to each node in the initial round and the number of training times are set, and the adaptive feature vector of the current node in the current round is obtained based on the adaptive feature vector of the current node in the previous round and the adaptive weight matrix corresponding to the current node in the current round; until the preset number of training times is reached, the target feature vector of each node is obtained; wherein, when the training round is 1, the adaptive feature vector of the current target node in the previous round is its preliminary completion feature vector, and the adaptive feature vector of each neighbor node of the current target node in the previous round is its projected feature vector; Based on the target adjacency matrix of the current target node in the current relationship, the target feature vectors of each neighbor node of the current target node in the current relationship are normalized to obtain the feature embedding vectors of each neighbor node of the current target node in the current relationship; The target feature vector of the current target node and the feature embedding vectors of each neighbor node of the current target node in the current relationship are weighted averaged to obtain the neighborhood embedding vector of the current target node in the current relationship as the actual neighborhood of the current target node in the current relationship; the neighborhood transfer vector of the current target node in the current relationship is set, the ideal neighborhood of the current target node in the current relationship is defined, and the difference between the ideal neighborhood and the actual neighborhood of the current target node in the current relationship is determined; the difference between the ideal neighborhood and the actual neighborhood of the current target node in the current relationship is minimized by scaling and transfer transformation to obtain the neighborhood transfer vector of the current target node in the current relationship, and the neighborhood transfer vector of the current target node in each relationship is obtained in turn; based on the target feature vectors of the current target node and its each neighbor node and the neighborhood transfer vector of the current target node in each relationship, a complete completion feature vector of the current target node is obtained; Based on the complete completed feature vector of each target node and the target feature vector of each neighboring node of each target node, the relationship strength between nodes in the heterogeneous graph is calculated, and the attention mechanism is used for aggregation to obtain the contextual representation of each target node, and residual connection is performed to obtain the final node representation of each target node for link prediction.
2. According to the method for embedding heterogeneous graphs based on information completion in claim 1, it is characterized in that: Based on all nodes in the heterogeneous graph, the adaptive weight matrix corresponding to each node in the initial round and the number of training times are set, and based on the adaptive feature vector of the current node in the previous round and the adaptive weight matrix corresponding to the current node in the current round, the adaptive feature vector of the current node in the current round is obtained; until the preset number of training times is reached, the target feature vector of each node is obtained; wherein, when the training round is 1, the adaptive feature vector of the current target node in the previous round is its preliminary completion feature vector, and the adaptive feature vector of each neighbor node of the current target node in the previous round is its projected feature vector, including: Based on the adaptive feature vector of the target node v in the (p-1) round and the adaptive weight matrix corresponding to the target node v in the p round The adaptive feature vector of the target node v in round p is obtained, and its expression is: in, Represents the adaptive feature vector of the target node v in round p; represents the adaptive feature vector of the target node v in round (p-1). In the initial round, i.e. when p=1, Represents the adaptive weight matrix corresponding to the target node v in round p; Based on the adaptive feature vector of any neighbor node u of the target node v in round (p-1) and the adaptive weight matrix corresponding to the neighbor node u of the target node v in round p The adaptive feature vector of the neighbor node u of the target node v in round p is obtained, and its expression is: in, Represents the adaptive feature vector of neighbor node u of target node v in round p; represents the adaptive feature vector of the neighbor node u of the target node v in round (p-1). In the initial round, i.e. when p=1, Represents the adaptive weight matrix of neighbor node u of target node v in round p; Until the preset number of training times is reached, the target feature vector h' of the target node v is obtained v , the target feature vector h' of the neighbor node u of the target node v u .
3. The heterogeneous graph embedding method based on information completion according to claim 1, characterized in that: The target feature vectors of each neighbor node of the current target node in the current relationship are normalized based on the target adjacency matrix of the current target node in the current relationship to obtain the feature embedding vectors of each neighbor node of the current target node in the current relationship, including: Based on the target adjacency matrix A' of the target node v in relation r v,r , the target feature vector h' of the neighbor node u of the target node v in the relationship r u Perform normalization to obtain the feature embedding vector of the neighbor node u of the target node v in the relationship r Its expression is: Among them, h' u Represents the target feature vector of the neighbor node u of the target node v; Represents the normalized coefficient of the target node v in the relation r, and its expression is: D r A' v,r The degree matrix of 4. The heterogeneous graph embedding method based on information completion according to claim 1, characterized in that: The target feature vector of the current target node and the feature embedding vectors of each neighbor node of the current target node in the current relationship are weighted averaged to obtain the neighborhood embedding vector of the current target node in the current relationship as the actual neighborhood of the current target node in the current relationship; Setting the neighborhood transfer vector of the current target node in the current relationship, defining the ideal neighborhood of the current target node in the current relationship, and determining the difference between the ideal neighborhood and the actual neighborhood of the current target node in the current relationship include: The target feature vector h' for the target node v v , the feature embedding vectors of each neighboring node of the target node v in the relation r are weighted averaged to obtain the neighborhood embedding vector of the target node v in the relation r, which is expressed as: in, Represents the neighborhood embedding vector of the target node v in the relation r and serves as the actual neighborhood of the target node v in the relation r; N v,r Represents the set of neighbor nodes of the target node v in the relation r; |N v,r | indicates N v,r The number of neighbor nodes in ; Represents the feature embedding vector of the neighbor node u of the target node v in the relation r; Set the neighborhood transfer vector of the target node v in the relation r and define the ideal neighborhood of the target node v in the relation r Its expression is: Among them, z v,r represents the neighborhood transfer vector of the target node v in the relation r; h' v represents the target feature vector of the target node v; The ideal neighborhood of the target node v in the relation r The actual neighborhood The difference C v,r , whose expression is: Among them, C v,r Represents the difference between the ideal neighborhood and the actual neighborhood of the target node v in the relation r.
5. The heterogeneous graph embedding method based on information completion according to claim 1, characterized in that: The scaling and transfer transformation is used to minimize the difference between the ideal neighborhood and the actual neighborhood of the current target node in the current relationship, and the neighborhood transfer vector of the current target node in the current relationship is obtained. The neighborhood transfer vector of the current target node in each relationship is obtained in turn, including: In the l-th layer network, use the scaling and transfer transformation to obtain the neighborhood transfer vector of the target node v in the l-th layer network in the relationship r Its expression is: in, Represents the scaling vector of the target node v in the l-th layer network in the relation r; Represents the transfer vector of the target node v in the l-th layer network in the relationship r; Set The dimensions and The dimensions are the same; τ represents a unit vector; Represents the sum of neighborhood transfer vectors of all target nodes in the l-th layer network in relation r; Represents the learnable weights of the l-th layer network; According to the neighborhood transfer vector of the target node v in each layer of the network in the relationship r, minimize the ideal neighborhood of the target node v in the relationship r The actual neighborhood The difference C v,r , get the neighborhood transfer vector z of the target node v in the relationship r v,r .
6. The heterogeneous graph embedding method based on information completion according to claim 1, characterized in that: The expression of the complete completion feature vector F[v] of the target node v is: Among them, h' v represents the target feature vector of the target node v; z v,r represents the neighborhood transfer vector of the target node v in the relation r; M v Indicates the set of relationship types corresponding to the target node; h' u Represents the target feature vector of the neighbor node u of the target node v in the relation r; N v,r Represents the set of neighbor nodes of the target node v in the relation r; Represents the sum of target feature vectors of each neighbor node in the neighbor node set of the target node v in the relation r; It represents the difference between the neighborhood transfer vector of the target node v in the relation r and the sum of the target feature vectors of each neighbor node in the neighbor node set.
7. The heterogeneous graph embedding method based on information completion according to claim 1, characterized in that: The method of sorting the top k neighbor nodes of the current target node based on the similarity of the attribute features between the current target node and each of its neighbor nodes includes: In the neighborhood of the target node v, determine the neighbor node set N of the target node v v , based on the cosine similarity calculation formula, determine the attribute feature similarity Γ between the target node v and any of its neighbor nodes u u,v , whose expression is: Among them, f u represents the projected feature vector of neighbor node u; f v represents the projected feature vector of the target node v; e represents the edge between the neighbor node u and the target node v; Obtain the attribute feature similarity of the target node v and each neighbor node in its neighbor node set in turn; sort the attribute feature similarity of the target node v and each neighbor node in its neighbor node set from large to small, select the neighbor nodes corresponding to the first k attribute feature similarities of the target node v as the first k neighbor nodes of the target node v, and construct the first k neighbor node set N of the target node v v,k ={u1,u2,...,u k }.
8. The heterogeneous graph embedding method based on information completion according to claim 1, characterized in that: The weighted average of the projected feature vectors of the current target node and its first k neighbor nodes to obtain the preliminary completed feature vector of the current target node includes: For each neighbor node in the set of the first k neighbor nodes of the target node v, set the corresponding weight of each neighbor node; perform weighted average on the projected feature vectors of the first k neighbor nodes of the target node v and their corresponding weights to obtain the aggregated feature embedding value of the current target node Its expression is: Among them, u a Represents the first k neighbor node set N of the target node v v,k The ath neighbor node in ; Represents the first k neighbor node set N of the target node v v,k The weight corresponding to the ath neighbor node in ; Represents the first k neighbor node set N of the target node v v,k The projection feature vector of the ath neighbor node in; k represents the first k neighbor node set N of the target node v v,k The number of neighbor nodes in ; Aggregate feature embedding value for target node v With its projected eigenvector f v Perform weighted averaging to obtain the initial completed feature vector of the target node v Its expression is: in, represents the aggregate feature embedding value of the target node v; f v Represents the projected feature vector of the target node v.
9. The heterogeneous graph embedding method based on information completion according to claim 1, characterized in that: The method of learning the similarity between nodes in the heterogeneous graph based on the preliminary completed feature vectors of each target node and the projected feature vectors of each neighbor node of each target node, and obtaining a relational adjacency graph of each relation in the heterogeneous graph; aggregating the initial adjacency matrix of each relation with the relational adjacency graph to obtain a target adjacency matrix of each relation, thereby obtaining a target adjacency matrix of the current target node in the current relation includes: In S31-S35, the tail node t m and t n is the target node; the initial completion feature vector of each target node All use f * Represents, that is, the initial completion feature vector of the target node v Use f v express; S31: For a relation r in a heterogeneous graph, the first-order feature similarity between node features is obtained through metric learning, and its expression is: in, Indicates that the node under relationship r and the first-order feature similarity between node η; Y represents the number of heads; It represents the corresponding weight when the number of heads is y under the relation r; Representation Node Projected feature vector / preliminary completed feature vector; f η Represents the projected feature vector / preliminary completed feature vector of node η; if node is the target node, then Representation Node The initial completion feature vector of If it is not the target node, Representation Node The projected eigenvector of ; ⊙ represents the Hadamard product; According to the first-order feature similarity Get the similarity graph Its expression is: Among them, ε F ∈[0,1] represents the control similarity graph The sparsity parameter of S32: Based on the relationship r, calculate the feature similarity graph between different types of nodes, that is, the source node s j and the tail node t m Similarity graph between Its expression is: in, Represents the source node s j and the tail node t m The first-order feature similarity between FB Represents the control similarity graph The sparsity parameter of Based on the relationship r, the feature similarity graph between nodes of the same type is calculated, and its expression is: in, Represents the source node s i and j Feature similarity graph between ; Represents the tail node t m and t n Feature similarity graph between ; Represents the source node s i and j The first-order feature similarity between ; Represents the tail node t m and t n The first-order feature similarity between i 、f j Represents the source node s i 、s j The projected eigenvector of m 、f n Respectively represent the tail node t m ,t n The initial completion feature vector of FS , ε FT Represent the control similarity graph The sparsity parameter of S33: Topological structure propagation through heterogeneous graphs, based on source node s i and j The feature similarity graph between the tail node t m and t n The feature similarity graph between them generates a potential propagation graph, which is expressed as: in, Respectively The corresponding propagation diagram; A r Represents the initial adjacency matrix of relation r; S34: Using channel attention to feature similarity map Fusion is performed to generate a relational adjacency graph of relation r Its expression is: in, Represents the weight corresponding to the relationship r; S35: Connect the relation adjacency graph of relation r The initial adjacency matrix A of the relationship r r Aggregate and get the target adjacency matrix of relation r, which is expressed as: Among them, softmax(·) means that w ψ,r The weights are normalized; thus, the target adjacency matrix of the target node v in the relationship r can be obtained as A' v,r .
10. The heterogeneous graph embedding method based on information completion according to claim 1, characterized in that: The method calculates the relationship strength between nodes in the heterogeneous graph based on the complete completion feature vector of each target node and the target feature vector of each neighboring node of each target node, and adopts the attention mechanism to aggregate, obtains the context representation of each target node, and performs residual connection to obtain the final node representation of each target node, including: In S51-S54, the target feature vector h' of each neighboring node of each target node is * They are all represented by F[*], that is, the target feature vector h' of the neighbor node u of the target node v u Denoted by F[u]; S51: For relation triples The target node is the tail node t, and the neighbor node of the tail node t is the source node s. The relationship strength between the source node s and the tail node t is calculated to obtain the attention value ATT-head(s,e,t) of the relationship, which is expressed as: K(s)=K-Linear φ(s) (F[s]) Q(t)=Q-Linear φ(t) (F[t]) Among them, φ(s) represents the type of source node s; φ(t) represents the type of tail node t, and tail node t is the target node; represents the type of connection relationship e between the source node s and the tail node t; K(s) represents the strength of the relationship between the source node s; Q(t) represents the strength of the relationship between the tail node t; F[s] represents the target feature vector of the source node s; F[t] represents the complete complement feature vector of the tail node t; Linear φ(*) Represents a linear projection, which changes the dimension of the source node s or the tail node t from R d become d represents the vector dimension, δ represents the number of attention heads, and d / δ represents the vector dimension of each head; Represents the parameter for adjusting the attention value; Indicates the weight corresponding to ATT-head(s,e,t); S52: Connect the δ attention heads together to obtain the attention vector of each node pair, which is expressed as: Among them, the Softmax function is used to normalize the attention distribution of the source node s adjacent to the tail node t; s∈N(t) indicates that the source node s belongs to the set of neighbor nodes of the tail node t; S53: Combine the attention vector of each node pair with the message vector from the adjacent nodes to obtain the context representation of each target node, which is expressed as: Among them, the tail node t is the target node; represents the context representation of the tail node t in the ξ-th layer relation stack; represents the message vector from the neighboring nodes; Indicates MSG-head γ The weight corresponding to (s, e, t); S54: Map the context representation of each target node to the distribution of the node type corresponding to each target node, and perform residual connection to obtain the final node representation of each target node, which is expressed as: Among them, the tail node t is the target node; H (ξ) [t] represents the final node representation of the tail node t in the ξ-th layer relation stack; H (ξ-1) [t] represents the final node representation of the tail node t in the (ξ-1)th layer of the relationship stack; ξ represents the number of layers of the relationship stack; A-Linear φ(t) Represents a linear projection.