Heterogeneous graph embedding learning method based on attention mechanism
By introducing a three-layer attention mechanism in heterogeneous graph embedding learning, the shortcomings of traditional models in dealing with complex heterogeneous graph relationships are solved, and a higher node classification accuracy is achieved.
Patent Information
- Application Number
- CN202110479995.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-30
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-04-30
AI Technical Summary
Traditional isomorphic deep networks cannot effectively handle complex, cross-pattern interactions in heterogeneous graphs, and it is difficult to correctly handle complex relationships of different types of nodes and edges in heterogeneous infographics.
A heterogeneous graph embedding learning method based on attention mechanism is designed, and three layers of attention mechanisms are adopted: type-level attention, node-level attention and semantic attention. The nodes are converted to a unified feature space through the type conversion matrix, and the self-attention mechanism is used for weighted aggregation to obtain the final node embedding.
In the node classification experiments on the ACM dataset and IMDB dataset, this method significantly improved the classification accuracy of micro-f1 and macro-f1 compared with the HAN model, and improved 1.9% and 1.85% on the ACM dataset respectively.
Smart Images

Figure CN113095439B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a heterogeneous graph embedding learning method based on an attention mechanism, and belongs to the field of graph neural networks and artificial intelligence. Background Art
[0002] In recent years, a special type of graph called heterogeneous information graph (HIN) has become a hot topic in network mining research. Heterogeneous graphs with mixed edges and nodes represent complex relationships, such as recommendation systems, paper citation networks, etc. The most important feature of HIN is the meta-path, which reflects the semantic relationship in the node-edge tuple. Taking the paper citation network as an example, the relationship between two papers can be taken as Paper-Author-Paper illustration (co-author relationship) and Paper-Subject-Paper (same-topic relationship). From the tuples listed above, we can see that in a heterogeneous graph, different connection patterns contain different relationships, and traditional homogeneous graph deep networks cannot handle complex, cross-modal interactions in heterogeneous networks. Attention mechanism is widely used in deep learning. It can handle data of variable size and make the model tend to focus on the most salient parts of the data. Graph attention network (GAT) is a neural network model specially designed to process graph structured data, but it only processes non-heterogeneous graphs with a single type of nodes or edges. Compared with homogeneous graph embedding learning, heterogeneous graphs are more complicated. Summary of the invention
[0003] The purpose of this invention is to provide a heterogeneous graph embedding learning method based on attention mechanism. Based on self-attention and multi-head attention, combined with hierarchical attention structure, a new heterogeneous graph embedding learning model based on three-layer attention mechanism is designed.
[0004] To achieve the above object, the present invention adopts the following technical solution:
[0005] A heterogeneous graph embedding learning method based on attention mechanism includes the following steps:
[0006] Step 1: Convert all nodes in the heterogeneous graph to a unified feature space through a type conversion matrix;
[0007] Step 2: Design type-level attention to learn the attention weights of a given node for neighbors of different categories;
[0008] Step 3: Design node-level attention learning to learn the attention weights of neighbor nodes based on meta-paths, and perform weighted aggregation based on the attention weights to obtain node embeddings based on specific meta-paths;
[0009] Step 4: Design semantic-level attention to learn the attention weights of different meta-paths, and perform weighted aggregation on the node embeddings based on different meta-paths according to the attention weights to obtain the final node embeddings;
[0010] Step 5: Perform prediction training on node labels;
[0011] Step 6: Design the loss function and use the back propagation algorithm to perform model optimization training.
[0012] In step 1, since the heterogeneous graph contains different types of nodes, all nodes are converted to a unified feature space through the type conversion matrix:
[0013] h′ i =M τ ·h i
[0014] Among them, M τ is the projection feature, h i and h′ i They are the node features before and after projection respectively.
[0015] The step 2 comprises the following steps:
[0016] Step 21, given node v i , its neighborhood embedding representation of type τ is defined as That is, for node v i The sum of the neighbor node features of type τ, where is the normalized adjacency matrix, v j Indicates v i The neighbor nodes of is node v i The set of nodes of type τ among the neighbor nodes of ;
[0017] Step 22, based on node v i The projected node feature h′ i and its neighborhood embedding representation of type τ Compute Node v i Attention scores about τ-type neighborhoods:
[0018]
[0019] Among them, || represents the connection operation, is a τ-type attention vector shared by all nodes; σ(·) represents the activation function; the superscript T represents transposition;
[0020] Step 23, normalize the attention scores of all types through the softmax function to obtain the type-level attention weights:
[0021]
[0022] The step 3 comprises the following steps:
[0023] Step 31, design node-level attention coefficient based on self-attention: given a node pair (v i ,v j ), by concatenating the representations of node pairs and using attention vectors to learn the importance between nodes and their neighbors;
[0024]
[0025]
[0026] Among them, att node Represents a deep neural network that computes node-level attention, where the node pair (v i ,v j ) It shows that j v i The importance of is asymmetric, that is, node v i For node v j The importance and node v j For node v i The importance of can vary greatly; is a node pair (v i ,v j ), σ represents the activation function, || represents the concatenation operation, is the attention vector (parameter), Indicates that for type τ j The node v j Type-level attention weights;
[0027] Step 32, learning a node embedding representation with specific semantics through node-level aggregation operations, where the embedding representation of each node is obtained by weighted aggregation of its neighbor representations;
[0028]
[0029] in, It is the embedded representation of a node under a certain meta-path. Given a certain meta-path, node-level attention learns the representation of the node under a certain semantics.
[0030] Step 33, repeatedly calculate the node-level attention R times, and concatenate the learned embeddings to learn the embedding representation based on the meta-path Φ
[0031]
[0032] Where r = 1, 2…R;
[0033] Given a set of meta-paths {Φ 0 ,Φ 1 ,…,Φ P}, which contains P different meta-paths. After performing node-level attention calculation, we get node embedding representations with P groups of specific semantics.
[0034] The step 4 comprises the following steps:
[0035] Step 41, given a set of meta-paths, node-level attention is used to learn node representations under different semantics, and semantic-level attention is used to learn the importance of semantics and fuse node representations under multiple semantics. The formal description of semantic-level attention is as follows:
[0036]
[0037] in, is the attention weight of each meta-path, att sem represents a deep neural network that performs semantic-level attention, using a single-layer neural network and semantic-level attention vectors to learn the importance of each semantic and normalize it through softmax;
[0038] Step 42, given a set of meta-paths {Φ 0 ,Φ 1 ,…,Φ P}, for each node, we can get the node embedding representation of P groups of specific semantics
[0039] Step 43, using the self-attention mechanism to aggregate the different embeddings of each node on different meta-paths, first define three matrices Query Matrix Q, Key Matrix K and Value Matrix V:
[0040]
[0041]
[0042]
[0043] The importance of node embedding based on this meta-path is calculated based on the self-attention model:
[0044]
[0045] Where |V| represents the number of nodes on the meta-path φ, d refers to the square root of the dimension, and q is the attention vector of the semantic relationship;
[0046] After obtaining the importance score of each meta-path, it is normalized by the softmax function, and the meta-path φ p The weight coefficient is:
[0047]
[0048] All the above parameters are shared for all meta-paths and semantic-specific embeddings. can be interpreted as the metapath Φ p Contribution to a specific task, obviously, The higher the meta-path Φ p The more important it is; by weighted fusion of multiple semantics, the final node representation is obtained to obtain the final node embedding Z:
[0049]
[0050] In step 5, a multi-layer perceptron or a softmax function is used to predict the node label.
[0051] The step 6 comprises the following steps:
[0052] In step 61, the final node embedding is aggregated by all the embeddings with specific semantics, and the final node embedding is applied to specific tasks, and different loss functions are designed; for semi-supervised node classification, the cross entropy between the predicted class label distribution and the true class label of the class-labeled node is minimized:
[0053]
[0054] Among them, C is the parameter of the classifier, y L is the set of labeled node indices, Y l and Z l are the labels and embedding representations of labeled nodes;
[0055] Step 62, under the guidance of the labeled data, the model parameters are optimized through the back propagation algorithm and the learned node embedding.
[0056] Beneficial effects: The present invention proposes a heterogeneous graph embedding learning algorithm based on the attention mechanism. Based on self-attention and multi-head attention, combined with the hierarchical attention structure, a new heterogeneous graph embedding learning model based on the three-layer attention mechanism is designed. The model follows a hierarchical attention structure, including type-level attention, node-level attention and semantic-level attention. First, all nodes are converted to a unified feature space through the type conversion matrix, and then enter the hierarchical attention module, which is type-level attention, node-level attention and semantic-level attention, so as to obtain node embedding containing specific semantics for specific tasks. Finally, the label of the node is predicted by the MLP layer. The present invention uses a heterogeneous graph embedding model based on a hierarchical attention mechanism to conduct node classification experiments on the ACM dataset and the IMDB dataset. The results show that compared with the model HAN based on the deep graph neural network, the present model has achieved better classification accuracy, and the micro-f1 and macro-f1 on the ACM dataset have increased by 1.9% and 1.85% respectively. In addition, the work of the present invention provides a new research idea for how to apply the attention mechanism to heterogeneous graphs. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 The overall architecture of the heterogeneous graph embedding model based on the attention mechanism. DETAILED DESCRIPTION
[0058] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation methods.
[0059] Figure 1 The overall architecture of the present invention is shown. First, all nodes are converted to a unified feature space through a type conversion matrix, and then enter the hierarchical attention module. After obtaining the node embedding containing specific semantics for a specific task, the node label is predicted through the MLP layer.
[0060] First, the symbols used in the present invention are summarized in Table 1:
[0061] Table 1 Symbols and corresponding explanations
[0062]
[0063] The present invention provides a heterogeneous graph embedding learning method based on an attention mechanism, comprising the following steps:
[0064] (1) All nodes in the heterogeneous graph are converted to a unified feature space through a type conversion matrix;
[0065] Since the heterogeneous graph contains different types of nodes, firstly, all nodes are converted to a unified feature space through the type conversion matrix:
[0066] h′ i=M τ ·h i
[0067] Among them, M τ is the projection feature, h i and h′ i They are the node features before and after projection respectively.
[0068] (2) Design type-level attention to learn the attention weights of a given node for its neighbors of different categories;
[0069] Generally, given a specific node, neighboring nodes of different types may have different effects on it. For example, neighboring nodes of the same type can carry more useful information. In addition, neighboring nodes of the same type may also have different importance. To this end, a new three-layer attention mechanism is designed.
[0070] Type-level attention can learn the weight coefficients of neighbors of different categories. First, given a node v i , its neighborhood embedding representation of type τ is defined as That is, for node v i The sum of the neighbor node features of type τ, where is the normalized adjacency matrix, v j Indicates v i The neighbor nodes of is node v i The set of nodes of type τ among the neighbor nodes of .
[0071] Afterwards, based on node v i The embedding representation vector (i.e. the node feature after projection) h i and type embedding representation Compute Node v i Attention scores about τ-type neighborhoods:
[0072]
[0073] Among them, || represents the connection operation, is an attention vector (parameter) of type τ, which is shared by all nodes. The superscript T represents the transpose, and σ(·) represents the activation function, such as Leaky ReLU.
[0074] Afterwards, the type-level attention weights are obtained by normalizing the attention scores of all types using the softmax function:
[0075]
[0076] (3) Design node-level attention to learn the attention weights of neighbor nodes based on meta-paths, and perform weighted aggregation based on the attention weights to obtain node embeddings based on specific meta-paths;
[0077] Before aggregating the meta-path neighbor information of each node, it can be noted that the meta-path-based neighbors of each node play different roles and show different importance in the node embedding learning of specific tasks. Node-level attention is introduced here, so that the importance of meta-path-based neighbors for each node in the heterogeneous graph can be learned, and the representations of these meaningful neighbors are aggregated to form node embedding.
[0078] In a specific task, the neighbor nodes of a node on the meta-path have different importance. Node-level attention can learn the representation of a node's neighbor nodes based on the meta-path as the embedding of the node. The node-level attention coefficient is designed based on self-attention. Given a node pair (v i ,v j ), by concatenating the representations of node pairs and utilizing an attention vector to learn the importance between a node and its neighbors.
[0079]
[0080]
[0081] Among them, att node Represents a deep neural network that computes node-level attention, where the node pair (v i ,v j ) It shows that j v i The importance of is asymmetric, that is, node v i For node v j The importance and node v j For node v i The importance of can vary greatly; is a node pair (v i ,v j ), σ represents the activation function, || represents the concatenation operation, is the attention vector (parameter), Indicates that for type τ j The node v j The type-level attention weights of .
[0082] The results show that node-level attention can maintain asymmetry, which is an important property of heterogeneous graphs. Finally, node-level aggregation operations are used to learn semantically specific node embedding representations. The embedding representation of each node is obtained by weighted aggregation of its neighbors’ representations.
[0083]
[0084] in, It is the embedded representation of a node under a certain meta-path. Given a certain meta-path, node-level attention can learn the representation of the node under a certain semantics. However, in actual heterogeneous graphs, there are often multiple meta-paths with different semantics, and a single meta-path can only reflect one aspect of the node. In order to fully describe the node, it is necessary to fuse the semantic information of multiple meta-paths.
[0085] Since heterogeneous graphs are scale-free, the variance of graph data is large. To solve the above problem, the node-level attention is expanded to use a multi-head attention mechanism to make the training process more stable. Specifically, the node-level attention is repeatedly calculated R times, and the learned embeddings are concatenated to learn the embedding representation based on the meta-path Φ
[0086]
[0087] Where r = 1, 2…R;
[0088] Given a set of meta-paths {Φ 0 ,Φ 1 ,…,Φ P}, after performing node-level attention calculation, we can get the node embedding representation of the specific semantics of group P
[0089] (4) Design semantic-level attention to learn the attention weights of different meta-paths, and perform weighted aggregation on the node embeddings based on different meta-paths according to the attention weights to obtain the final node embeddings;
[0090] Given a set of meta-paths, node-level attention is used to learn node representations under different semantics. Furthermore, semantic-level attention can be used to learn the importance of semantics and fuse node representations under multiple semantics. The formal description of semantic-level attention is as follows:
[0091]
[0092] in, is the attention weight of each meta-path. semRepresents a deep neural network that performs semantic-level attention. Specifically, it uses a single-layer neural network and a semantic-level attention vector to learn the importance of each semantic (meta-path) and normalizes it through softmax.
[0093] Given a set of meta-paths {Φ 0 ,Φ 1 ,…,Φ P}, for each node, we can get the node embedding representation of P groups of specific semantics (meta-path)
[0094] Considering that each node has different embedded representations on different meta-paths, we consider using the self-attention mechanism to aggregate them. First, we define three matrices: Query Matrix Q, Key Matrix K, and Value Matrix V.
[0095]
[0096]
[0097]
[0098] Based on the self-attention model
[77] Compute the importance of node embeddings based on this meta-path:
[0099]
[0100] Here |V| represents the number of nodes on the meta-path φ, d refers to the square root of the dimension, and q is the attention vector of the semantic relationship. After obtaining the importance score of each meta-path, it is normalized by the softmax function. Then the meta-path φ p The weight coefficient is:
[0101]
[0102] Note that, for meaningful comparison, all the above parameters are shared for all meta-paths and semantic-specific embeddings. can be interpreted as the metapath Φ p Contribution to a specific task. Obviously, The higher the meta-path Φ p The more important it is. By weighted fusion of multiple semantics (meta-paths), the final node representation can be obtained. Note that the meta-path weights here are optimized for specific tasks. For different tasks, the meta-path Φ pThere may be different weights. These semantically specific embeddings are fused using the learned weights as coefficients to obtain the final embedding Z, as shown below:
[0103]
[0104] (5) Use a multi-layer perceptron or softmax function to predict node labels;
[0105] (6) Design loss function and use back propagation algorithm to optimize model training;
[0106] The final embedding is aggregated by all the embeddings of specific semantics. The final embedding is then applied to specific tasks and different loss functions are designed. For semi-supervised node classification, the cross entropy between the predicted class label distribution and the true class label of the class-labeled node is minimized:
[0107]
[0108] Where C is the parameter of the classifier, y L is the set of labeled node indices, Y l and Z l It is the label and embedding representation of labeled nodes. Guided by the labeled data, the model parameters are optimized through the back-propagation algorithm and the learned node embeddings.
[0109] The pseudo code of the heterogeneous graph embedding learning algorithm based on the attention mechanism is shown in Table 2.
[0110] Table 2 Heterogeneous graph embedding model algorithm flow
[0111]
[0112]
[0113] The present invention mainly uses the ACM dataset and the IMDB dataset to test the classification accuracy of the node classification test model.
[0114] ACM: Extract papers published on KDD, SIGMOD, SIGCOMM, MobiCOMM, and VLDB and classify them into three categories (database, wireless communication, and data mining). Then construct a heterogeneous graph containing 3025 articles (P), 5835 authors (A), and 56 topics (S), as well as two types of meta-paths: PAP (article---author---article) and PSP (article---topic---article). Paper features are elements in the bag of words corresponding to the keyword representation, and the conferences where the papers are published are used as their labels.
[0115] IMDB: A subset of IMDB is extracted for experiments. The dataset contains 4780 movies (M), 5841 actors (A), and 2269 directors (D), as well as three types of meta-paths: MAM (movie-actor-movie) and MDM (movie-director-movie). Among them, only movie nodes have labels, and movie nodes have three different types: action movies, comedies, and dramas. Movie features are elements in a set of word bags corresponding to plot keywords.
[0116] Here, the dataset is divided into three parts: training set, validation set, and test set. For the ACM dataset, 600 papers are randomly selected as the training set, 300 papers are used as the validation set, and the rest are used as the test set. For the IBDM dataset, 300 movies are randomly selected as the training set, 300 movies are used as the validation set, and the rest are used as the test set.
[0117] For model settings, we first randomly initialize the learning parameters and use the Adam algorithm to optimize the model. In addition, for some hyperparameter settings: the learning rate is set to 0.005, the L2 regularization coefficient is 0.001, the semantic level attention vector q dimension is 128, the number of multi-heads in the multi-head attention mechanism is 8, the dropout loss rate is 0.6, and the early stopping mechanism is added (if the accuracy of the validation set remains unchanged for 30 epochs, the training is stopped). For fairness, the embedding dimension of all methods is set to 64. Since the variance of graph structure data may be large, the test is repeated 20 times and the averaged Macro-F1 and Micro-F1 indicators are reported.
[0118] This completes the heterogeneous graph embedding learning algorithm based on the attention mechanism.
[0119] ACM dataset and IMDB dataset are used for semi-supervised node classification experiments. First, the learning parameters are randomly initialized and the Adam algorithm is used to optimize the model. In addition, for some hyperparameter settings: the learning rate is set to 0.005, the L2 regularization coefficient is 0.001, the semantic level attention vector q dimension is 128, the number of multi-heads in the multi-head attention mechanism is 8, the dropout loss rate is 0.6, and the early stopping mechanism is added (if the accuracy of the validation set remains unchanged for 100 epochs, the training is stopped). For fairness, the embedding dimension of all methods is set to 64. Since the variance of graph structure data may be large, the test is repeated 20 times and the averaged Macro-F1 and Micro-F1 indicators are reported.
[0120] The test results based on the ACM dataset are shown in Table 3.
[0121] Table 3 Test results based on ACM dataset
[0122]
[0123] The test results based on the IMDB dataset are shown in Table 4.
[0124] Table 4 Test results based on IMDB dataset
[0125]
[0126] So far, the node classification test based on the ACM dataset and IMDB dataset has been completed.
[0127] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A heterogeneous graph embedding learning method based on attention mechanism, Features: The steps include: Step 1: Convert all nodes in the heterogeneous graph to a unified feature space through a type conversion matrix; Step 2: Design type-level attention to learn the attention weights of a given node for neighbors of different categories; Step 3: Design node-level attention learning to learn the attention weights of neighbor nodes based on meta-paths, and perform weighted aggregation based on the attention weights to obtain node embeddings based on specific meta-paths; Step 4: Design semantic-level attention to learn the attention weights of different meta-paths, and perform weighted aggregation on the node embeddings based on different meta-paths according to the attention weights to obtain the final node embeddings; Step 5: Perform prediction training on node labels; Step 6: Design a loss function and use the back propagation algorithm to perform model optimization training; Node classification experiments are conducted on the ACM dataset and IMDB dataset using a heterogeneous graph embedding model based on a hierarchical attention mechanism.
2. The heterogeneous graph embedding learning method based on the attention mechanism according to claim 1, Features: In step 1, since the heterogeneous graph contains different types of nodes, all nodes are converted to a unified feature space through the type conversion matrix: h' i =M τ ·h i Among them, M τ is the projection feature, h i and h' i They are the node features before and after projection respectively.
3. The heterogeneous graph embedding learning method based on the attention mechanism according to claim 1, Features: The step 2 comprises the following steps: Step 21, given node v i , its neighborhood embedding representation of type τ is defined as That is, for node v i The sum of the neighbor node features of type τ, where is the normalized adjacency matrix, v j Indicates v i The neighbor nodes of is node v i The set of nodes of type τ among the neighbor nodes of ; Step 22, based on node v i The projected node feature h' i and its neighborhood embedding representation of type τ Compute Node v i Attention scores about τ-type neighborhoods: Among them, || represents the connection operation, is a τ-type attention vector shared by all nodes; σ(·) represents the activation function; the superscript T represents transposition; Step 23, normalize the attention scores of all types through the softmax function to obtain the type-level attention weights:
4. The heterogeneous graph embedding learning method based on the attention mechanism according to claim 1, Features: The step 3 comprises the following steps: Step 31, design node-level attention coefficient based on self-attention: given a node pair (v i ,v j ), by concatenating the representations of node pairs and using attention vectors to learn the importance between nodes and their neighbors; Among them, att node Represents a deep neural network that computes node-level attention, where the node pair (v i ,v j ) It shows that j v i The importance of is asymmetric, that is, node v i For node v j The importance and node v j For node v i The importance of can vary greatly; is a node pair (v i ,v j ), σ represents the activation function, || represents the concatenation operation, is the attention vector (parameter), Indicates that for type τ j The node v j Type-level attention weights; Step 32, learning a node embedding representation with specific semantics through node-level aggregation operations, where the embedding representation of each node is obtained by weighted aggregation of its neighbor representations; in, It is the embedded representation of a node under a certain meta-path. Given a certain meta-path, node-level attention learns the representation of the node under a certain semantics. Step 33, repeatedly calculate the node-level attention R times, and concatenate the learned embeddings to learn the embedding representation based on the meta-path Φ Where r = 1, 2…R; Given a set of meta-paths {Φ 0 ,Φ 1 ,…,Φ P }, which contains P different meta-paths. After performing node-level attention calculation, we get node embedding representations with P groups of specific semantics.
5. The heterogeneous graph embedding learning method based on the attention mechanism according to claim 1, Features: The step 4 comprises the following steps: Step 41, given a set of meta-paths, node-level attention is used to learn node representations under different semantics, and semantic-level attention is used to learn the importance of semantics and fuse node representations under multiple semantics. The formal description of semantic-level attention is as follows: in, is the attention weight of each meta-path, att sem represents a deep neural network that performs semantic-level attention, using a single-layer neural network and semantic-level attention vectors to learn the importance of each semantic and normalize it through softmax; Step 42, given a set of meta-paths {Φ 0 ,Φ 1 ,…,Φ P }, for each node, we can get the node embedding representation of P groups of specific semantics Step 43, using the self-attention mechanism to aggregate the different embeddings of each node on different meta-paths, first define three matrices Query Matrix Q, Key Matrix K and Value Matrix V: The importance of node embedding based on this meta-path is calculated based on the self-attention model: Where |V| represents the number of nodes on the meta-path φ, d refers to the square root of the dimension, and q is the attention vector of the semantic relationship; After obtaining the importance score of each meta-path, it is normalized by the softmax function, and the meta-path φ p The weight coefficient is: All the above parameters are shared for all meta-paths and semantic-specific embeddings. can be interpreted as the metapath Φ p Contribution to a specific task, obviously, The higher the meta-path Φ p The more important it is; by weighted fusion of multiple semantics, the final node representation is obtained to obtain the final node embedding Z:
6. The heterogeneous graph embedding learning method based on the attention mechanism according to claim 1, Features: In step 5, a multi-layer perceptron or a softmax function is used to predict the node label.
7. The heterogeneous graph embedding learning method based on the attention mechanism according to claim 1, Features: The step 6 comprises the following steps: In step 61, the final node embedding is aggregated by all the embeddings with specific semantics, and the final node embedding is applied to specific tasks, and different loss functions are designed; for semi-supervised node classification, the cross entropy between the predicted class label distribution and the true class label of the class-labeled node is minimized: Among them, C is the parameter of the classifier, y L is the set of labeled node indices, Y l and Z l are the labels and embedding representations of labeled nodes; Step 62, under the guidance of the labeled data, the model parameters are optimized through the back propagation algorithm and the learned node embedding.
Citation Information
Patent Citations
Heterogeneous graph neural network generation method and device, electronic device and storage medium
CN110046698A
Method for searching Top-k similarity on semantically enhanced heterogeneous information network
CN111222049A