Dynamic heterogeneous graph representation learning method based on mask auto-encoder
By using mask autoencoder and attention mechanism in dynamic heterograph representation learning, combined with occlusion mechanism and joint optimization, the problems of processing difficulties and label dependence on dynamic heterographs in the prior art are solved, and better robustness and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510252018.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
AI Technical Summary
Existing graph representation learning methods are difficult to effectively deal with dynamic heterogeneous graphs, and rely on label information, and cannot fully utilize the advantages of self-supervised learning.
A dynamic heterogeneous graph representation learning method based on mask autoencoder is adopted, and the space-time information is captured through two layers of attention mechanisms, and the occlusion mechanism is used to force the model to restore the obscured part through context information, combining reconstruction errors and prediction errors for joint optimization.
It realizes learning accurate dynamic heterogeneous graph embedding without relying on label information, which improves the robustness and generalization capabilities of the model, and can better capture the spatio-temporal information of the graph structure.
Smart Images

Figure CN120181129A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of graph neural networks, and specifically relates to a dynamic heterogeneous graph representation learning method based on a masked autoencoder. Background Art
[0002] A graph is a data structure composed of nodes and edges between nodes, which can effectively reveal the relationships between entities and has powerful expressive ability. Therefore, many real-world data abstraction models adopt the form of graphs. The process of mapping sparse high-dimensional node information into a low-dimensional vector space through a certain model is called graph representation learning. This method has been successfully applied in various fields such as recommendation systems, biology, and transportation networks.
[0003] However, many graph structures transformed from real-world data often have heterogeneity and dynamics. Spatially, both nodes and relationships have different categories; temporally, the nodes and connections in the graph will change over time, and this graph structure is called a temporal heterogeneous graph (THG). Many existing graph representation learning methods either only focus on static homogeneous graphs or extend the methods to heterogeneous graphs but do not consider the dynamics of the graph structure. Moreover, many existing graph neural network models mainly adopt the supervised learning paradigm, and the learning process is completely guided by real labels, which is costly to obtain and cannot be widely used. Self-supervised learning (SSL) that does not use or uses very few labels becomes even more important.
[0004] Among the SSL learning methods on graphs, contrastive learning has been studied more. This method often relies on high-quality data augmentation or carefully designed optimization algorithms. And graph augmentation methods are generally related to the graph structure itself, and the performance varies greatly in different graphs and data. In addition, contrastive learning also involves complex negative sampling design, making it difficult to guarantee the effect. Generative SSL methods can avoid the above problems. Generative methods focus on reconstructing the data of the input graph without considering complex algorithm design. Masked autoencoder (MAE) is a generative SSL method that has been widely applied in computer vision and natural language processing in recent years. VGAE uses the idea of autoencoders, and GraphMAE was the first to apply the masked autoencoder to the field of graph representation learning, but this method ignores the heterogeneity and dynamics of real networks. HGMAE extends the masked autoencoder to the heterogeneous graph field, but this method does not consider the dynamic characteristics of the graph. Therefore, existing graph representation learning methods either do not consider both temporal and spatial information at the same time or rely on label information and do not utilize the advantages of self-supervised learning. Summary of the Invention
[0005] The present invention aims to solve the problems of the lack of self-supervised methods and incomplete coverage of spatio-temporal information in dynamic heterogeneous graph representation learning, and proposes a dynamic heterogeneous graph representation learning method based on a masked autoencoder. Specifically, on the one hand, a two-layer attention mechanism is used to aggregate the information of nodes with the same relationship and the semantic information between nodes with different relationships within a single time snapshot. Then, a temporal attention mechanism is used to fuse the information in different time snapshots to obtain the overall dynamic heterogeneous graph embedding, and then a decoder is used for future link prediction. On the other hand, in order not to rely on label information, random masking operations are respectively performed on the node features and topological structure of the original dynamic heterogeneous graph. The masked embedding information is obtained using the same encoder as above, and the original input information is reconstructed through a decoder with the same structure as the encoder. This can force the model to recover and predict the masked part through the remaining context information, thereby improving the robustness and generalization ability of the model. Finally, the model is jointly optimized using the reconstruction error and prediction error to obtain the final embedding representation.
[0006] The dynamic heterogeneous graph representation learning method based on a masked autoencoder of the present invention is implemented by the following steps:
[0007] Step 1: Read the dynamic heterogeneous graph in the dataset and construct it into a dynamic heterogeneous graph sequence in the form of snapshots.
[0008] Step 2: Input the original dynamic heterogeneous graph sequence into the dynamic heterogeneous graph encoder to obtain an embedding vector, and use the embedding vector to perform a dynamic link prediction task through a predictor to obtain a prediction loss.
[0009] The dynamic heterogeneous graph encoder uses a two-layer attention mechanism to capture spatial heterogeneous features and a temporal attention mechanism to capture the temporal feature information between snapshots.
[0010] Step 3: Perform a masking operation on the original dynamic heterogeneous graph: randomly mask the edges, and randomly select a part of the node features and replace the original features with mask tokens. During the masking process, follow the retention replacement rule: a part of the selected features remains unchanged, and the other part is replaced with mask tokens, and finally the masked dynamic heterogeneous graph is obtained.
[0011] Step 4: Input the masked dynamic heterogeneous graph into the dynamic heterogeneous graph encoder to obtain a masked embedding vector, and perform a re-masking operation on the embedding vector: replace some of the features in the embedding vector with a new mask token to replace the original feature information.
[0012] Step 5: Input the re-masked embedding vector into a decoder with the same structure as the dynamic heterogeneous graph encoder to obtain a feature vector reconstructed and restored by the decoder.
[0013] Step 6: Calculate the error between the reconstructed feature vectors and the feature vectors of the original dynamic heterogeneous graph to obtain the reconstruction loss.
[0014] Step 7: Combine the reconstruction loss in Step 6 with the prediction loss obtained in Step 2 as the final training loss.
[0015] Step 8: Repeat Steps 2 to 7, and use gradient descent and backpropagation to update the loss until the model converges.
[0016] Preferably, publicly available datasets such as ogbn, Aminer, DBLP, and Yelp are used to verify the performance of dynamic heterogeneous graph representation learning through dynamic link prediction tasks.
[0017] Preferably, the open-source tool library DeepGraphLibrary is used as the tool for dynamic heterogeneous graph training.
[0018] Beneficial effects of the present invention:
[0019] The present invention proposes a dynamic heterogeneous graph representation learning model based on a masked autoencoder. The masked autoencoder architecture is introduced into the dynamic heterogeneous graph representation learning model, alleviating the dependence on label information. The masking mechanism is completed by edge masking and node feature masking. The node feature masking can force the model to learn the missing information from the surrounding neighbor nodes, thereby improving the performance of the model. The edge masking randomly masks some edges before training to disrupt the connectivity of the graph, prompting the model to better capture the semantic information of the original graph structure. At the same time, the encoder adopts a two-layer attention mechanism to capture the intra-relationship and inter-relationship information of the single-snapshot heterogeneous graph respectively, and uses temporal attention to fuse the time information of multiple snapshots to learn accurate dynamic heterogeneous graph embeddings. Description of the Drawings
[0020] Figure 1 is the overall framework diagram of the present invention. Detailed Embodiments
[0021] The following further illustrates the present invention in conjunction with the drawings and specific embodiments.
[0022] As Figure 1 shown, the dynamic heterogeneous graph representation learning method based on a masked autoencoder specifically includes the following steps:
[0023] Step 1: Read the dynamic heterogeneous graph in the dataset and construct it into a heterogeneous graph sequence in the form of snapshots. The dynamic heterogeneous graph contains both heterogeneous information of different node types and edge types and time-varying information, and can be expressed as where G is the snapshot at the corresponding moment, and T is the size of the time window, that is, the number of snapshots included. Each heterogeneous graph snapshot can be expressed as where is a set of nodes of type , and ε is a set of edges of type . And represents the adjacency matrix, represents the feature matrix of nodes. and are the sets of node types and edge types respectively, and the node set and the edge set satisfy
[0024] Step 2: Input the original dynamic heterogeneous graph sequence into the dynamic heterogeneous graph encoder to obtain the embedding vector, and use the embedding vector to perform the dynamic link prediction task through the predictor to obtain the prediction loss. Put the original dynamic heterogeneous graph adjacency matrix A t and the corresponding node features X t into the encoder f Encoder , so as to obtain the unmasked latent node embedding: H = f Encoder (A t , X t ). The specific learning process of the encoder is divided into node-level aggregation, semantic-level aggregation, and temporal aggregation. Since the feature dimensions of different types of nodes in the dynamic heterogeneous graph may be inconsistent, they are first all mapped to the same dimension: where are the original feature vector of dimension d' and the projected feature vector of dimension d respectively, W φ(i) and b φ(i) represent the trainable projection matrix and bias matrix specific to type φ(i) respectively, and σ(·) represents the non-linear activation function.
[0025] Node-level aggregation: Split the individual snapshots into different subgraphs according to different connection relationships. The present invention uses the self-attention mechanism to aggregate the node feature information of each subgraph. At time t, the self-attention score of a pair of nodes (i, j) with connection relationship r is: where σ(·) represents the non-linear activation function LeakyReLU, is the representation vector of node i at time t, is the linear transformation matrix under relationship r, || represents the vector concatenation operation, represents the set of neighbor nodes connected to node i at time t with relationship r, represents the attention vector related to relationship r. Then, the node embedding of node i at time t under the condition of relationship r is calculated by weighted summation: To strengthen the feature capture ability, the multi-head attention mechanism can also be used.
[0026] Semantic-level aggregation: The present invention also uses the self-attention mechanism to fuse the feature information between different relationships. The importance of different relationships can be expressed as: where is the attention score related to relationship r, R(i) represents the set of all relationship types related to node i, b is the bias vector, is the set of all nodes connected by relationship r at time t, and are both learnable transformation matrices regarding the connection type R. Finally, by performing a weighted sum of different relationships, the final embedding representation of node i at time t can be obtained:
[0027] Temporal aggregation: In order to fuse the feature representations on multiple snapshots, the present invention uses the temporal attention mechanism for information aggregation. First, use to add temporal encoding information to the embedding information of different snapshots, where W T , b T represent the linear transformation matrix and the bias matrix of the temporal aggregation layer respectively, and PE(t) is the temporal encoding function. The calculation method of the k-th bit is: Then use to represent the node embedding matrix of node i from time 1 to T, that is to get Q = H i W q , K = H i W k and V = H i W v , where are all linear transformation matrices. Finally, the embedding representation after temporal aggregation is obtained where is the mask matrix, which can ensure that only the data with a time step less than t is focused on during training. Specifically, M ij is 0 when i ≤ j, and -∞ otherwise. This can ensure that in the softmax function, the corresponding attention scores will approach 0 when i > j. as the final node embedding representation.
[0028] To alleviate the over-smoothing problem that may be caused by stacking multiple layers of graph neural networks, the present invention additionally adopts a gating mechanism. The representation of node i at time t is where and are the trainable weight parameter and transformation matrix respectively.
[0029] Finally, the possibility of an edge existing between node pair (i, j) at time T + 1 can be obtained through the obtained embedding where MLP(·) is a multi - layer perceptron for calculating node pairs (i, j), σ is a non - linear activation function for calculating the edge existence probability, and ∥ represents the vector concatenation operation. Finally, the prediction loss is obtained.
[0030] Step 3: Perform a masking operation on the original dynamic heterogeneous graph: Randomly mask the edges and randomly select a part of the node features and replace the original features with mask tokens. During the masking process, follow the retention - replacement rule: A part of the selected features remains unchanged, and the other part is replaced with mask tokens, finally obtaining the masked dynamic heterogeneous graph.
[0031] Node feature masking: Masking the features of nodes can force the model to learn the missing information from the surrounding neighbor nodes, thereby improving the performance of the model. For each snapshot of the original dynamic heterogeneous graph Before training, select t from G at a set ratio δ as the target nodes for feature masking operation, that is, for each node in , replace its original feature vector with a special mask token , finally obtaining a masked feature matrix where the node feature vector is: Meanwhile, to improve the stability of the model in the face of diverse data, the present invention adopts a variable masking ratio. Specifically, apply a linear masking adjustment function formally defined as where n is the current training epoch, is the maximum number of training epochs. γ(n) will control the masking ratio δ to increase at a fixed step size Δ, that is, δ(n + 1)=δ(n)+Δ. Further set γ(0)=MIN δ , and to ensure that the model can converge, where Δ, MIN δ , MAX δ are hyperparameters that can be adjusted according to different datasets. The present invention also applies the retention - replacement criterion to solve the possible mismatch problem between the training phase and the inference phase. Specifically, first replace a part of the mask tokens with random tokens with a probability of p r , and in addition, select another part of the nodes with a probability of p u , and during the masking process, the features of these nodes will remain unchanged.
[0032] Edge masking: Edge masking disrupts the connectivity of the graph by randomly masking some edges before training, prompting the model to better capture the semantic information of the original graph structure. From a dynamic heterogeneous graph snapshot G t ={V t , E t , A t , X t}, for the edge set E t in 1 ≤ t ≤ T, a set of edges E Mask to be masked is randomly sampled at a certain ratio, and then A Mask is the adjacency matrix of E Mask . Then the final adjacency matrix after edge masking can be expressed as
[0033] Step 4: Input the masked dynamic heterogeneous graph into the dynamic heterogeneous graph encoder to obtain the masked embedding vectors, and perform a re-masking operation on the embedding vectors: replace some of the features in the embedding vectors with a new masking token to replace the original feature information. Put the adjacency matrix of the dynamically heterogeneous graph after graph masking and the node features into the encoder f Encoder mentioned in Step 2 to obtain the masked node embeddings: To ensure that the encoder can learn more meaningful embeddings without relying on the decoder's restoration ability, the present invention performs a re-masking operation before the decoding process. For each node in the masked node set , a special masking token is used to replace the embedding representation of the node. The process of re-masking is expressed as: where is the re-masking representation matrix composed of .
[0034] Step 5: Input the re-masked embedding vectors into a decoder with the same structure as the dynamic heterogeneous graph encoder to obtain the feature vectors reconstructed and restored by the decoder. Put and the adjacency matrix mentioned above into the decoder f Decoder to obtain the final reconstructed feature matrix expressed as
[0035] Step 6: Calculate the error between the reconstructed feature vectors and the feature vectors of the original dynamic heterogeneous graph to obtain the reconstruction loss. Take the error between the original graph and the reconstructed graph feature matrix as the loss function, specifically
[0036] Step 7: Combine the reconstruction loss in Step 6 with the prediction loss obtained in Step 2 as the final training loss. The overall loss of the model of the present invention is defined as where λ and μ are the corresponding balancing weights, which are hyperparameters that do not change during training.
[0037] Step 8: Repeat Steps 2 to 7, and use gradient descent and backpropagation to update the loss until the model converges.
[0038] The present invention conducts experiments on four real-world datasets: ogbn, Aminer, DBLP, and Yelp.
[0039] Compare the experimental results of the disclosed method of the present invention with the results of DGI, HGY, SimpleHGN, HDE, RGCN, Ev-olveGCN, DySat, and HTGNN. The results on the ogbn dataset are shown in Table 1, the results on the Aminer dataset are shown in Table 2, the results on the DBLP dataset are shown in Table 3, and the results on the Yelp dataset are shown in Table 4. It can be seen that there is an improvement on all four real datasets.
[0040] Table 1
[0041]
[0042] Table 2
[0043]
[0044] Table 3
[0045]
[0046] Table 4
[0047]
Claims
1. A dynamic heterogeneous graph representation learning method based on masked autoencoder, characterized in that: The following steps are involved: Step 1: Read the dynamic heterogeneous graph in the data set and construct it into a dynamic heterogeneous graph sequence in the form of snapshots; Step 2: Input the dynamic heterogeneous graph sequence into the dynamic heterogeneous graph encoder to obtain an embedding vector, and use the embedding vector to perform the dynamic link prediction task through the predictor to obtain the prediction loss; Step 3: Perform a masking operation on the dynamic heterogeneous graph, input the masked dynamic heterogeneous graph into the dynamic heterogeneous graph encoder to obtain a masked embedding vector, and perform a re-masking operation on the masked embedding vector; Step 4: Input the masked embedding vector into the decoder with the same structure as the dynamic heterogeneous graph encoder to obtain the feature vector reconstructed by the decoder; Step 5: Calculate the error between the reconstructed feature vector and the embedding vector of the dynamic heterogeneous graph to obtain the reconstruction loss. Combine the reconstruction loss with the prediction loss as the final training loss and train until convergence.
2. The dynamic heterogeneous graph representation learning method based on masked autoencoder according to claim 1 is characterized in that: The dynamic heterogeneous graph encoder adopts a two-layer attention mechanism to capture spatial heterogeneous features and a temporal attention mechanism to capture temporal feature information between snapshots.
3. The dynamic heterogeneous graph representation learning method based on masked autoencoder according to claim 2 is characterized in that: The specific learning process of the dynamic heterogeneous graph encoder is divided into node-level aggregation, semantic-level aggregation and time aggregation; Node-level aggregation: Split individual snapshots into different subgraphs according to different connection relationships, and use the self-attention mechanism to aggregate the node feature information of each subgraph; at time t, the self-attention score of a pair of nodes (i, j) with a connection relationship of r for: Where σ(·) represents the nonlinear activation function LeakyReLU, is the representation vector of node i at time t, is the linear transformation matrix under the relation r, || represents the vector concatenation operation, represents the set of neighbor nodes connected to node i by relationship r at time t, Represents the attention vector related to relation r; then the node embedding of node i at time t under the condition of relation r is calculated by weighted summation: Semantic level aggregation: The self-attention mechanism is used to fuse the feature information between different relations. The importance of different relations is expressed as: in is the attention score associated with relation r, R(i) represents the set of all relation types associated with node i, b is the bias vector, is the set of all nodes connected by relationship r at time t, and They are all learnable transformation matrices about the connection type R; finally, the weighted summation of different relationships is performed to obtain the final embedding representation of node i at time t: Temporal aggregation: Use the temporal attention mechanism to aggregate information. First, use Add time coding information to the embedded information of different snapshots, where W T , b T They represent the linear transformation matrix and bias matrix of the time aggregation layer respectively. PE(t) is the temporal encoding function. The calculation method of the kth bit is: Then use The node embedding matrix representing node i from time 1 to time T is We get Q = H i W q , K=H i W k And V=H i W v ,in They are all linear transformation matrices; finally, we get the embedded representation after time aggregation in is the mask matrix, M ij It is 0 when i≤j, otherwise it is -∞; As the final node its embedding representation; A gating mechanism is adopted; the node i at time t is expressed as in and They are trainable weight parameters and transformation matrices respectively; Finally, the embedding obtained is used to obtain the possibility that the node pair (i, j) has an edge connection at time T+1 Where MLP(·) is a multi-layer perceptron that calculates the node pair (i, j), σ is a nonlinear activation function that calculates the probability of an edge, and ∥ represents a vector concatenation operation; finally, the prediction loss is 4. The dynamic heterogeneous graph representation learning method based on masked autoencoder according to claim 3 is characterized in that: The masking operation on the dynamic heterogeneous graph specifically includes: randomly masking the edges, randomly selecting a part of the node features and replacing the original features with mask marks; The retention replacement rule is followed during the masking process: a part of the selected features remain unchanged, and the other part is replaced with mask marks, and finally a dynamic heterogeneous graph is obtained after masking.
5. The method for learning dynamic heterogeneous graph representation based on masked autoencoder according to claim 4, characterized in that: The re-masking operation of the embedded vector specifically includes: replacing the original feature information of part of the features in the embedded vector with a new mask mark.
6. The method for learning dynamic heterogeneous graph representation based on masked autoencoder according to claim 5, characterized in that: In the masking process, a variable mask ratio is used and a linear mask adjustment function is applied, which is formally defined as Where n is the current training round, is the maximum number of training rounds, γ(n) controls the mask ratio δ to increase according to a fixed step size Δ, that is, δ(n+1)=δ(n)+Δ; set and Ensure that the model converges, where Δ, MIN δ , MAX δ It is a hyperparameter adjusted according to different datasets.