A Temporal Heterogeneous Network Link Prediction Method and System Based on Hierarchical Contrastive Learning
Through the hierarchical comparison learning method, time-sequence heterogeneous networks are mined from the perspectives of node-level, edge-level and time-level, which solves the problem of failing to effectively utilize fine-grained differential relationships in the existing methods, achieves more efficient link prediction, and improves the prediction accuracy in social networks, traffic management, and medical health fields.
Patent Information
- Application Number
- CN202411431193.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-10-14
AI Technical Summary
The existing timing heterogeneous network link prediction methods fail to effectively utilize fine-grained differential relationships, bridge spatial and temporal heterogeneity, resulting in poor link prediction performance.
A hierarchical contrast learning method is used to mine and model the timing heterogeneous network from three micro perspectives: node-level, edge-level and time-level. The heterogeneous structure and timing dependencies are captured through graph attention networks and recurrent neural networks, and link prediction is optimized using contrast learning loss function.
The performance of link prediction has been significantly improved, especially in application scenarios such as social networking, traffic management and medical health, improving the accuracy and reliability of prediction.
Smart Images

Figure CN119583368B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of temporal heterogeneous network prediction, and specifically relates to a method and system for predicting links in a temporal heterogeneous network. Background Art
[0002] Contemporary information networks (such as social networks and biological systems) are developing and iterating rapidly and becoming increasingly complex. These networks are usually composed of various types of nodes and link relationships, and continuously evolve over time. Traditional static homogeneous networks can no longer comprehensively abstract these systems, and the emergence of temporal heterogeneous networks can more accurately simulate complex real-world systems. Link prediction in these temporal heterogeneous complex networks is a long-term challenge.
[0003] The key challenges in link prediction research for temporal heterogeneous networks lie in modeling heterogeneous entity relationships and dynamic snapshot change patterns, which we call "spatial complexity" and "time complexity". Link prediction methods in existing literature mainly divide dynamic temporal networks into static snapshot sequences according to time points in chronological order, and characterize the complex entity relationships existing in each snapshot. Among them, (1) "spatial complexity" is mainly reflected in the complex heterogeneous relationships between multi-type nodes, and is mainly modeled by heterogeneous graph embedding methods to model various co-occurrence paradigms in complex networks. It can be generally divided into methods based on meta-paths and methods based on node attributes. The method based on meta-paths constructs meta-paths within each snapshot to mine heterogeneous information. The method based on node attributes aims to make full use of rich multiple attributes, fuse adjacent attributes together, and then input them into the node embedding process. (2) "Time complexity" mainly utilizes the dynamic distribution changes in the snapshot sequence, and temporal methods focus on capturing continuous evolution processes, which can be roughly divided into sequence methods and graph methods. The sequence method obtains inputs from a chronologically arranged sequence and captures the evolutionary dependencies between different snapshots according to recurrent neural networks (RNNs) and attention mechanisms. The graph method aggregates the embeddings of each dynamic node and continuously encodes the features that appear or disappear over time in the network according to graph neural networks (GNNs). Although these methods perform well in terms of efficiency, a significant drawback is that these methods coarsely model dynamic heterogeneous representations, only focusing on representation paradigms, but ignoring the fine-grained differential relationships widely distributed in temporal heterogeneous networks, which leads to poor performance in link prediction tasks.
[0004] For example, the authorized invention document discloses a method for predicting temporal network links using GraphSAGE (patent number ZL202010746960.8, authorized announcement number CN111756587B). It uses the time slicing method to divide the temporal network into a series of network snapshots, and then preprocesses the data on the number of connections and connection duration information of node pairs within the network snapshots; uses the preprocessed data as the input of the GraphSAGE algorithm and learns and trains to obtain a node embedding generation model; constructs a node similarity index by combining the embedding similarity of nodes and the topological structure similarity of nodes, so as to predict the future connection status of the corresponding node pairs.
[0005] In summary, among the existing related network link prediction methods, no one has proposed how to utilize this differential relationship to bridge the heterogeneity in space and time to comprehensively depict detailed dynamic and diverse features to improve the performance of link prediction. No one has proposed that the fine-grained differential relationships (space) between different nodes and edges and the differences in evolution paradigms (time) will affect representation learning and the performance of link prediction. Therefore, it is very necessary to provide a temporal heterogeneous network link prediction method based on hierarchical contrast learning. Summary of the Invention
[0006] The technical problem to be solved by the present invention is:
[0007] The purpose of the present invention is to provide a temporal heterogeneous network link prediction method (CLP) and system based on hierarchical contrast learning, which mines and models the spatial complexity and temporal complexity in the network from three different micro perspectives of node level, edge level and time level, so as to realize the prediction of node connection behaviors in real-world temporal heterogeneous networks.
[0008] The technical solution adopted by the present invention to solve the above technical problem is: A temporal heterogeneous network link prediction method based on hierarchical contrast learning, and the formal definition of temporal heterogeneous network link prediction:
[0009] Heterogeneous network: Given is an undirected graph, where V = {v1, v2,..., v N}, E = {e1, e2,..., e M}, A v and A e represent the node set, edge set, node type and edge type respectively; each node v i ∈V and edge e j ∈E both have a corresponding node type and edge type And |A v | + |A e | > 2, then is a heterogeneous graph;
[0010] Temporal heterogeneous network: Given an undirected heterogeneous graph consisting of a sequence of heterogeneous network snapshots at different times, that is where is the static snapshot graph of the network at time t, where and represent the node set and edge set at time t respectively, T represents the total number of snapshots, and for any node a ∈ V, its node representation is a feature vector of a fixed size then X = {x a} a∈V represents the feature matrix of all nodes;
[0011] The method is used to predict the probability that there is a target link e = (a, b) between two nodes a, b ∈ V in the temporal heterogeneous network at the future T + 1 moment, and obtain the user representation by training the learning model CLP and to obtain the probability that e exists in the snapshot graph ;
[0012] Propose an end-to-end model CLP based on contrastive learning, and obtain the deep representation of the temporal heterogeneous network by learning the heterogeneous information and dynamic dependencies of different types of nodes and edges, so as to express the spatial difference and temporal difference, and predict the existence of the target link through the similarity of the two target user vector representations;
[0013] CLP mainly includes three parts: (1) Spatial feature learning, (2) Temporal information modeling, (3) Output layer;
[0014] (1) Spatial feature learning: First, set the Graph Attention Network (GAT) to retain the heterogeneous structure information of the temporal heterogeneous network at the node level and edge level, so as to obtain the node-level representation vector of the heterogeneous structure and the edge-level representation vector At the same time, regard it as a homogeneous network, deploy GNN to capture its inherent structure information and generate the corresponding node-level representation vector and edge-level representation vector For and and perform double-layer contrastive learning at the node level and edge level respectively to enhance the acquisition of network structure differences;
[0015] (2) Temporal Information Modeling: Secondly, GRU and LSTM are respectively used to save the hidden states of each static snapshot graph, so as to obtain the temporal-level representation vectors. and perform contrastive learning on the node representations and to enhance the capture of temporal differences in the temporal heterogeneous network;
[0016] (3) Output Layer: Finally, the similarity between node a and node b is calculated to represent the target link e=(a, b), and it is fed into an overall loss function to map the probability of the existence of the target link.
[0017] The present invention has the following beneficial technical effects:
[0018] The present invention proposes a new temporal heterogeneous network link prediction model (CLP) based on hierarchical contrastive learning, which solves the technical problems in the background art, performs multi-perspective differential representation on nodes, and models the heterogeneous semantic relationships and dynamic dependencies in the temporal heterogeneous network, so as to predict whether two nodes will generate a target link. It can formulate a scientific and reasonable decision-making reference scheme for applications such as social network friend recommendation, precise targeted advertising placement, traffic route optimization, and drug effect monitoring. The main contributions of the present invention are: (1) A three-layer hierarchical contrastive differential relationship extraction module is proposed, which realizes the function of eliminating multi-perspective differences and provides a differential bridging effect for spatial and temporal heterogeneity in the link prediction scenario respectively. (2) A heterogeneous temporal graph network is designed to model the structural and temporal distribution paradigms and overall eliminate the differences from different perspectives. Specifically, we learn the node-level and edge-level distribution paradigms through the graph attention network and propose a two-channel temporal network to capture the temporal dependencies between different snapshots at the time level. (3) Extensive experiments are carried out on four real temporal heterogeneous information networks, and the performance of the CLP model is compared with the existing advanced temporal heterogeneous network link prediction models. The experiments prove that the present model is significantly superior to the existing models in two metrics, AUC and AP.
[0019] The present invention believes that if the differential relationship (the fine-grained differential relationship between different nodes and edges) can be utilized to bridge the heterogeneity in space and time, it can better depict comprehensive and detailed dynamic diversity characteristics, thereby improving the performance of link prediction. Specifically, focusing on the aforementioned "temporal complexity" and "spatial complexity", the fine-grained differential relationship (space) between different nodes and edges and the difference in the evolution paradigm (time) directly affect the representation learning, and can largely affect the performance of link prediction. A method for predicting the links of a temporal heterogeneous network based on hierarchical contrast learning (CLP) provided by the present invention mines and models the spatial complexity and temporal complexity in the network from three different microscopic perspectives of the node level, edge level, and time level, so as to predict the node connection behavior in the temporal heterogeneous network in the real world. Description of the Drawings
[0020] Figure 1 is the overall framework of the CLP model (the overall framework of the method for predicting the links of a temporal heterogeneous network based on hierarchical contrast learning);
[0021] Figure 2 is the curve graph of the parameter selection experiment for the node embedding dimension d. In the graph: Figure 2 (a) shows the influence of the node embedding dimension d on AUC, Figure 2 (b) shows the influence of the node embedding dimension d on AP;
[0022] Figure 3 is the curve graph of the parameter selection experiment for the number of multi-head attention heads (h). In the graph: Figure 3 (a) shows the influence of the number of attention heads h on AUC, Figure 3 (b) shows the influence of the number of attention heads h on AP;
[0023] Figure 4 is the curve graph of the parameter selection experiment for the weight adjustment factors (λ1, λ2, λ3). In the graph: Figure 4 (a) shows the influence of the weight coefficient λ1 on AUC, Figure 4 (b) shows the influence of the weight coefficient λ1 on AP, Figure 4 (c) shows the influence of the weight coefficient λ2 on AUC, Figure 4 (d) shows the influence of the weight coefficient λ2 on AP, Figure 4 (e) shows the influence of the weight coefficient λ3 on AUC, Figure 4 (f) shows the influence of the weight coefficient λ3 on AP.
[0024] Figure 5 is the curve graph of the parameter selection experiment for the temperature coefficient (τ). In the graph: Figure 5 (a) shows the influence of the temperature coefficient τ on AUC, Figure 5 (b) shows the influence of the temperature coefficient τ on AP. Detailed Embodiment
[0025] 1. The proposed method for predicting links in a temporal heterogeneous network based on hierarchical contrastive learning aims to predict the probability of future connections between any two nodes in a temporal heterogeneous network, i.e., the link prediction task. By storing the heterogeneous structural information of node representation vectors through the proposed link prediction method, the temporal evolution process of the heterogeneous network is captured. At the same time, the topological dependencies between heterogeneous snapshots are captured to characterize the distribution patterns in complex temporal heterogeneous networks, thereby predicting the dynamic and complex connection relationships between entities. This greatly promotes the development of many real-life applications, including social recommendation, traffic management, healthcare, and network biology.
[0026] For example, in a social network platform, recommending people one knows, optimizing route management in traffic planning, completing and restoring case knowledge graphs in healthcare, and predicting protein interactions in network biology, link prediction helps to understand the complexity and functions of the above applications and provides new ideas and methods for them.
[0027] 2. Combined with the attached Figure 1 , the implementation of the technical solution of the proposed method for predicting links in a temporal heterogeneous network based on hierarchical contrastive learning is elaborated as follows:
[0028] 2.1 Overall method
[0029] Formal definition of link prediction in a temporal heterogeneous network:
[0030] Heterogeneous network: Given as an undirected graph, where V = {v1, v2,..., v N}, E = {e1, e2,..., e M}, A v and A e represent the node set, edge set, node type, and edge type respectively. Each node v i ∈ V and edge e j ∈ E has a corresponding node type and edge type and |A v | + |A e | > 2. Then is a heterogeneous graph.
[0031] Temporal heterogeneous network: Given an undirected heterogeneous graph consisting of a sequence of heterogeneous network snapshots at different times, i.e., where is the static snapshot graph of the network at time t, where, and represent the node set and edge set at time t respectively, T represents the total number of snapshots, and For any node a ∈ V, its node representation is a feature vector of a fixed size Then X = {x a} a∈V represents the feature matrix of all nodes.
[0032] The task of the present invention is to predict the probability of the existence of a target link e = (a, b) between two nodes a, b ∈ V at the future T+1 moment in a temporal heterogeneous network, and obtain the user representation by training a learning model CLP and the similarity between to obtain the probability of the existence of e in the snapshot graph
[0033] The present invention proposes an end-to-end model CLP based on contrastive learning. The main idea is to obtain the deep representation of a temporal heterogeneous network by learning the heterogeneous information and dynamic dependencies of different types of nodes and edges, so as to express the spatial and temporal differences, and predict the existence of the target link through the similarity between two target user vector representations. As Figure 1 shown, CLP is mainly divided into three parts: (1) spatial feature learning, (2) temporal information modeling, and (3) output layer.
[0034] (1) Spatial feature learning: First, we set up a Graph Attention Network (GAT) to retain the heterogeneous structural information of the temporal heterogeneous network at the node level and the edge level, so as to obtain the node-level representation vector of the heterogeneous structure and the edge-level representation vector of the heterogeneous structure. At the same time, regard it as a homogeneous network, and deploy a GNN to capture its inherent structural information to generate the corresponding node-level representation vector and the edge-level representation vector Perform double-layer contrastive learning on and and respectively at the node level and the edge level to enhance the acquisition of network structure differences.
[0035] (2) Temporal information modeling: Secondly, we use GRU and LSTM respectively to save the hidden state of each static snapshot graph to obtain the temporal-level representation vector and Perform contrastive learning on the node representations and to enhance the capture of the temporal differences of the temporal heterogeneous network.
[0036] (3) Output layer: Finally, we calculate the similarity between nodes a and b to represent the target link e=(a, b), and feed it into an overall loss function to map the probability of the existence of the target link.
[0037] 2.2 Spatial Feature Learning
[0038] To capture heterogeneous structural features, we first divide the static snapshots into specific sub-networks according to different edge types. Subsequently, we use node-level and edge-level GATs to represent different types of edges and nodes in the temporal heterogeneous graph. Through the modeling of the two-level GAT, the different importance degrees of different types of nodes and edges are effectively captured, promoting the processing of heterogeneous semantics and structures in the temporal heterogeneous graph. Finally, we design node-level and edge-level contrastive learning to enhance the representation of nodes from the node view and edge view respectively.
[0039] 2.2.1 Node-level Heterogeneous Feature Learning
[0040] For each snapshot graph it is divided into |R| subgraphs according to the edge type r ∈ R. Then the attention score between any two nodes (node a and b) belonging to the r-type subgraph of the static snapshot at time t is expressed as:
[0041]
[0042] where x a and x b are used to initialize nodes a and b. A rt and W rt are the attention weight vector and mapping matrix corresponding to the r-type subgraph of , which are updated during the training process. σ represents the activation function, and ∥ is the concatenation operation. Then the weight between nodes a and b in the r-type subgraph of is:
[0043]
[0044] where represents the r-th type neighbor of node a in graph . Then the representation of node is obtained by weighted summation:
[0045]
[0046] In addition, to enhance the representation of each node, we adopt a node-level contrastive learning method to ensure that multiple augmented representations of the same node are intrinsically similar while being different from the representations of other nodes. Specifically, we ignore the categories of nodes and edges and use a GNN to linearly aggregate all unclassified node embeddings:
[0047]
[0048] To ensure that the two representations and of the same node a aggregated by GAT and GNN respectively are similar, while being dissimilar to the representation of a different node b aggregated by GAT, we use a node-level InfoNCE loss function, considering as a positive sample pair and as a negative sample pair, as shown in formulas (5) and (6):
[0049]
[0050] where τ is the temperature coefficient. sim() is represented by a dot product operation and is used to measure the similarity between two vectors, that is In addition, |V t |, T, and R represent the graph of, the number of snapshots, and the number of edge types respectively.
[0051] 2.2.2 Edge-level Heterogeneous Feature Learning
[0052] The node-level feature learning module can capture specific information of specific edge types. However, heterogeneous networks usually contain multiple edge types. To effectively obtain information from all edge types of each node, we design an edge-level feature learning module to obtain the importance weights of different edge types. Specifically, we aggregate these specific edge information of different types to generate node embeddings with heterogeneous edge information. Specifically, the embedding vector of each node is generated through a non-linear mapping. In the t-th snapshot graph, the attention weight between node a and edge type r is output through the softmax activation function, as shown in formulas (7) and (8):
[0053]
[0054] where z, W t and b represent the trainable attention weight vector, weight matrix, and bias vector respectively. σ is a non-linear activation function. Then, its representation is obtained by weighted summing the nodes of specific edge types in :
[0055]
[0056] In addition, to enhance the representation of each specific type of edge node, we adopt an edge-level contrastive learning method to ensure that different augmented representations of the same subgraph have intra-similarity, while having inter-dissimilarity in comparison with nodes of other types of subgraphs. Specifically, we ignore the node and edge types in graph G t and regard it as an isomorphic network, and use GNN to linearly aggregate the node representations in this isomorphic network:
[0057]
[0058] where denotes the neighbor set of node a with edge type r in graph .
[0059] To ensure the similarity between the specific edge type embedding of node a and its embedding in the isomorphic graph , while making it significantly different from the aggregated representation of node b, we set an edge-level InfoNCE loss function. This loss function regards the heterogeneous and isomorphic aggregated representations of the same node as positive sample pairs and the aggregated representations from different nodes in the heterogeneous graph as negative sample pairs as shown in formulas (11) and (12):
[0060]
[0061] where denotes all adjacent nodes of node a in graph , covering all edge types.
[0062] 2.3 Temporal Information Modeling
[0063] The time information modeling layer aims to address the heterogeneity problem of time series information in different sequence modeling scenarios, and different modeling techniques can capture different sequence patterns. Specifically, we adopt a two-channel architecture to learn different sequence dependence paradigms: (1) In the long-term channel, we deploy LSTM to explore the inherent interdependencies in the long-term time evolution process; (2) In the short-term channel, we apply GRU to analyze the interactions between adjacent snapshots in the short-term evolution process. However, existing research has overlooked the heterogeneity between long-term and short-term dependencies. To this end, we propose a contrastive learning method to emphasize the differences between different sequence learning paradigms, thereby focusing on modeling time heterogeneity. We use LSTM and GRU to represent the spatio-temporal patterns of learning long-term and short-term dependencies respectively, denoted as and
[0064]
[0065] where, represents the embedding of node a in all snapshots at all time points, generated from the embedding vectors of node a in all snapshots. To address the temporal inhomogeneity between different semantic spaces, we design a time-level contrastive learning method to bring the hidden vector representations obtained by LSTM and GRU closer.
[0066]
[0067] where, V T is the set of nodes in the last snapshot graph .
[0068] 2.4 Output Layer
[0069] We use the embedding vector of node a in the last snapshot to perform the link prediction task, aiming to infer the potential connections between node i and other nodes. Therefore, this problem is transformed into evaluating the similarity between node i and the neighbor nodes in the last snapshot . We adopt binary cross-entropy minimization as the objective function, which is defined as shown in formula (17):
[0070]
[0071] where, for any node a in the final snapshot graph , its embedding representation is obtained by mean pooling of and , that is Meanwhile, in the graph Among them, neighbor i of node a is regarded as a positive example, while randomly selected non-neighbor node j is regarded as a negative example, thus forming a triple (a, i, j). The set is defined as the set of all possible triples, that is Then is the set of positive example sample pairs, is the set of negative example sample pairs.
[0072] Then, by weighting the cross-entropy loss node-level contrastive learning loss edge-level contrastive learning loss and time-level contrastive learning loss the total loss
[0073]
[0074] is obtained. Among them, λ1, λ2, and λ3 are learnable weight adjustment factors used to coordinate the three losses.
[0075] The method for predicting links in a temporal heterogeneous network based on hierarchical contrastive learning according to the present invention is applied to social recommendation or traffic management scenarios. A social network is a dynamic platform with diverse interaction characteristics among users. In the rapidly evolving pattern of social networks, accurately predicting potential connections between users is crucial for improving user engagement and satisfaction. The present invention studies the technology of predicting links in a temporal heterogeneous network. By utilizing the heterogeneity and temporal dependence information of social interactions, the accuracy of recommending potential friends in a social network can be significantly improved. Specifically, the social network at time T is regarded as a temporal heterogeneous graph to form a sequence of T static heterogeneous snapshots. (1) In terms of heterogeneity, a social network is essentially heterogeneous, including various types of node sets V t (such as users, posts) and edge sets E t (such as likes, comments, follows). The present invention uses a spatial feature learning layer to learn and reflect rich node embeddings with diversity, capturing complex interdependencies within the network. (2) In terms of temporal dynamics, time information is crucial for understanding the evolution of social relationships. The present invention fully utilizes the recent and historical interaction data of the social network through a temporal information modeling layer, capturing the long-term and short-term interaction characteristics between social network snapshots, and capturing the temporal evolution of user behavior, thereby generating more timely and relevant friend recommendations. (3) Finally, through the overall loss function in the output layer of the present invention Train the embedding vectors of social network nodes to optimize the probability that there is a link e between node pairs (a, b) formed by any two users (user a and user b) in the social network. By accurately modeling heterogeneous information and temporal information in the social network, the present invention can provide more personalized and timely friend recommendations for the social network, ultimately improving user satisfaction and engagement, and bringing great potential and opportunities in promoting the expansion of interpersonal relationships, knowledge sharing, and offline activity organization, etc.
[0076] 2.5 Algorithm Description
[0077] The pseudo-code of a temporal heterogeneous network link prediction method and system based on hierarchical contrast learning is shown in Algorithm 1, which describes the learning training and convergence process of the CLP model. The node embedding dimension, the number of nodes, and the number of snapshots of the model are denoted as d, N, and T respectively, then the time complexity of the model is O(TNd 2 ).
[0078]
[0079] 3 Verification of the Effect of the Present Invention
[0080] 3.1 Model Implementation Details
[0081] Model parameter settings: During the training process, the model uses a batch size of 1024 and achieves convergence within 5 epochs. The learning rate is set to 1e-4. In addition, the three balance parameters λ1, λ2, and λ3 are all assigned the value of 1e-08. In the three InfoNCE loss functions, τ is set to 0.1. The dimension of node embedding is set to 32, while the number of attention heads in the spatial feature learning module is set to 4. All experiments are carried out on the Ubuntu 18.04 operating system and the NVIDIA RTX A2000 12G graphics card.
[0082] 3.2 Datasets
[0083] To verify the efficiency and generalization of the model, experiments are carried out in four different heterogeneous temporal scenarios in this paper, namely Math-overflow, Taobao, OGBN-MAG, and COVID-19. The details of the datasets are shown in Table 3-1.
[0084] Table 3-1 Dataset Statistics
[0085]
[0086] Math-overflow: The Math-overflow dataset was collected by the Math Overflow community on the Stack Exchange website and published on the SNAP platform. It contains three different types of interactions (questions and answers, questions and comments, answers and comments) among users within 2350 days. In the experiment, with a time window of 124 days, this dataset was divided into 11 snapshots in subsequent experiments.
[0087] Taobao: The Taobao dataset consists of records of users clicking on cloud-themed products in the Taobao App from April 1 to May 31, 2008, and contains three types of nodes (users (U), products (I), and themes (T)) and three types of links (U-I, U-T, I-T). In the experiment, with a time window of 12 days, this dataset was divided into 5 snapshots.
[0088] OGBN-MAG: The OGBN-MAG dataset is composed of a subset of the Microsoft Academic Graph (MAG). The dataset contains four types of nodes (papers, authors, institutions, and research fields) and four relationships between them (an author belongs to an institution, an author writes a paper, a paper cites another paper, and the research field of a paper is a certain area). In this experiment, a THN was extracted from OGBN-MAG, and the time span of this THN was from January 1 to 10, 2010. The time slices were divided by day, and 1000 edges of each type were taken in each time slice to form the dataset used in the experiment.
[0089] COVID-19: The COVID-19 dataset comes from the 1point3acres platform and contains daily case reports at the state and county levels (such as confirmed cases, new cases, deaths, and recovered cases). This model uses the daily new COVID-19 cases as the time series data for each state and county. The dataset consists of two types of nodes (states and counties) and four relationships between them, namely two administrative subordination relationships (a state includes counties, a county belongs to a state) and two geospatial relationships (a state is near a state, a county is near a county). The THN constructed in this experiment has a time span from May 1 to 21, 2020, and contains 21 time slices. At most 2000 edges of each type are taken in each time slice.
[0090] 3.3 Experimental Evaluation
[0091] 3.3.1 Experiments and evaluations are conducted on the proposed CLP model using four datasets: Math-overflow, Taobao, OGBN-MAG, and COVID-19. For the problem of predicting the links between two nodes in a temporal heterogeneous network, the present invention uses two metrics, the Area Under Curve (AUC) and Average Precision (AP), to evaluate the prediction effect of this model. AUC refers to the area under the Receiver Operating Characteristic Curve (ROC), with the false positive rate shown on the x-axis and the true positive rate shown on the y-axis. AP refers to the area under the Precision-Recall curve, where the recall rate is the abscissa and the precision is the ordinate. The larger the values of AUC and AP, the closer they are to 1, and the better the prediction effect of this model.
[0092] 3.3.2 Comparison with baseline methods
[0093] To measure the performance of the CLP model, CLP is compared with the following several classic link prediction methods, which are mainly distributed in four categories: (1) static homogeneous networks, (2) static heterogeneous networks, (3) dynamic homogeneous networks, and (4) dynamic heterogeneous networks. Specifically as follows.
[0094] (1) Static homogeneous networks:
[0095] SEAL: Extract local graphs from target links and learn local graph features to map the probability of the existence of target links.
[0096] VGNAE: Combines variational inference, graph convolutional networks, and normalization techniques to effectively learn probabilistic node embeddings from graph-structured data.
[0097] (2) Static heterogeneous networks:
[0098] Metapath2Vec: Inputs the random walk sequences generated under the guidance of metapaths into the skip-gram model to learn the low-dimensional embeddings of nodes in heterogeneous information networks, thereby capturing the structural and semantic relationships in the network.
[0099] GATNE: Proposes a comprehensive framework to learn node representations in multi-attribute networks, addressing the challenges faced in integrating multiple types of relationships and node attributes into a unified representation.
[0100] (3) Dynamic homogeneous networks:
[0101] TGAT: A method for inductive representation learning on temporal graphs is proposed, which combines temporal encoding and G-graph neural networks to capture the dynamic characteristics of such graphs. This method can be generalized to new nodes and future graph snapshots.
[0102] TDGNN: A method for continuous-time link prediction in dynamic graphs is designed, which uses time encoding techniques and attention mechanisms to enhance the prediction ability of graph neural networks.
[0103] (4) Dynamic heterogeneous networks
[0104] THAN: Conducts representation learning on graphs with temporal heterogeneity, using memory mechanisms and transformer architectures to capture complex temporal and structural information in temporally heterogeneous graphs.
[0105] THAT: Proposes a comprehensive dynamic heterogeneous graph representation learning framework, which combines neighborhood type modeling, neighborhood information aggregation, and temporal dynamics integration to achieve accurate and effective node representations.
[0106] Table 3-2 Experimental results of comparison with the Baseline method on four datasets
[0107]
[0108] Table 3-2 summarizes the performance of the CLP model and eight baseline models on four datasets: Math-overflow, Taobao, OGBN-MAG, and COVID-1. Among them, the optimal results are in bold, and the sub-optimal results are underlined. The following conclusions can be drawn from the table: (1) The static homogeneous network models SEAL and VGNAE perform poorly on all four datasets, mainly because they cannot utilize temporal dynamics and heterogeneous information. (2) The static heterogeneous network models Metapath2Vec and GATNE show better performance than SEAL and VGNAE due to the utilization of heterogeneous information. However, their limitation lies in the neglect of network dynamics. For example, a link that is considered a negative link during training may become a positive link in the test set due to the dynamic characteristics of the original network. The inability to capture these temporal changes leads to node representation vectors conveying opposite meanings, thus reducing the model performance. (3) The dynamic homogeneous network models TGAT and TDGNN have limitations in utilizing heterogeneous information. In the network, different types of nodes and edges have different importance, but TGAT and TDGNN cannot distinguish them. (4) The dynamic heterogeneous network models THAN and THAT effectively integrate and model the temporal dynamics and heterogeneous types in graph data, which is reflected in their superior performance compared to other benchmark models. (5) On the four datasets, our CLP model significantly outperforms all baseline models in terms of AUC and AP. Our CLP model effectively learns heterogeneous features through node-level and edge-level contrastive learning techniques, and at the same time effectively models temporal information through temporal-level contrastive learning methods. Our method successfully captures the dynamic patterns and semantic information in dynamic heterogeneous graph data. Compared with the best baseline (i.e., THGAT), our model shows a significant performance improvement in terms of AUC and AP, with an average increase of 10.10% and 13.44% in AUC and AP, respectively. (6) The results of all models in the COVID-19 dataset generally decline. A possible explanation for this phenomenon is the scarcity of observed nodes and the resulting data instability, which leads to poor link prediction results.
[0109] 3.3.3 Ablation Experiment Settings
[0110] To verify the effectiveness of each module in the CLP model, the present invention sets up ablation experiments:
[0111] CLP -N : Remove the node-level contrastive learning module to evaluate its importance in modeling structural differences.
[0112] CLP -E : Remove the edge-level contrastive learning module to verify its effectiveness in modeling structural differences.
[0113] CLP -T : Remove the temporal contrast learning module to observe the contribution of temporal differences to the improvement of the system's prediction performance.
[0114] Table 3-3 Experimental results of comparison with variant methods on four datasets
[0115]
[0116] Ablation experiments were conducted on four datasets, Math-overflow, Taobao, OGBN-MAG, and COVID-1, to evaluate the performance of three key modules in the CLP model: node-level contrast learning, edge-level contrast learning, and temporal contrast learning. The comparison results are shown in Table 3-3. It can be seen from the table that CLP performs best compared with other variant models, and the prediction performance drops significantly after removing any one of the three modules. In particular, (1) when we remove the node-level contrast learning module (i.e., CLPl -N ), both AUC and AP show obvious decreases, with the decrease range from 2.35% to 25.30%. This further confirms the key role of node-level contrast learning representation in improving the efficiency of link prediction. (2) After removing the edge-level contrast learning module (i.e., CLP -E ), the most significant drop in prediction performance occurs, ranging from 3.36% to 35.65%. This emphasizes the indispensable role of edge-level contrast learning in our model. (3) Removing the temporal contrast learning module (i.e., CLP -T ) also causes a decrease in AUC and AP, ranging from 4.28% to 26.58%. In summary, node-level, edge-level, and temporal contrast learning losses constitute important elements of CLP, which fundamentally improve the accuracy of predicting time-heterogeneous links.
[0117] 3.3.4 Parameter comparison experiments
[0118] The present invention conducts a comparison experiment on the selection of key parameters. We select the node representation vector dimension (d), the number of multi-head attention heads (h), the loss function weight adjustment factors (λ1, λ2, λ3), and the temperature coefficient (τ) for comparative analysis. Figures 2 - 5 Shows the changes in the prediction metrics AUC and AP with these parameters. (1) For the node representation vector dimension (d). We set d = 8, 16, 32, 64, 128. As the vector dimension increases, the model performance shows an upward trend because the training ability of learnable parameters is enhanced under a larger training scale. However, too many dimensions will lead to a decline in model performance, which may be due to overfitting and an increase in noise in the representation vector. As Figure 2As shown, when d = 32, the optimal model performance can be obtained. (2) Number of multi-head attention heads (h). The h-head attention mechanism in the node-level and edge-level feature learning modules partitions the sub-semantic space, enabling our model to direct attention to different heterogeneous information dimensions. As Figure 3 shown, when h = 4, the expressive ability and attention allocation ability of our model are significantly improved. (3) Loss function weight adjustment factors (λ1, λ2, λ3). Adjusting the weights of the node-level, edge-level, and time-level contrast losses (denoted by λ1, λ2, and λ3 respectively) can achieve a good optimization balance. As Figure 4 shown, when λ1, λ2, and λ3 are all set to e-08, the best prediction performance can be obtained. (4) Temperature coefficient (τ). τ affects the training process and the performance of the final model by adjusting the sensitivity and attention to sample similarity. As Figure 5 shown, when τ is 0.1, our model effectively completes the contrastive learning training.
[0119] In summary, the method proposed in the present invention, as verified by experiments and practical applications, validates the claimed technical effects and practicality of the present invention.
[0120] The algorithm (method) proposed in the present invention is the underlying technical core of the present invention. Based on this algorithm, various products can be derived to apply to predicting dynamic and complex connection relationships between entities in related fields. It greatly promotes the development of many applications in real life, including social recommendation, traffic management, medical health, and network biology. For example, recommending people you know in a social network platform, optimizing route management in traffic planning, complementing and restoring case knowledge graphs in medical health, predicting protein interactions in network biology, and link prediction helps to understand the complexity and functions of the above applications.
[0121] Based on the algorithm (method) proposed in the present invention, a link prediction system for temporal heterogeneous networks based on hierarchical contrastive learning is developed using programming languages. The system has program modules corresponding to the steps of the above technical solution and executes the steps in the above link prediction method for temporal heterogeneous networks based on hierarchical contrastive learning when running.
[0122] The computer program of the developed system (software) is stored on a computer-readable storage medium. The computer program is configured to implement the steps of the above link prediction method for temporal heterogeneous networks based on hierarchical contrastive learning when called by a processor. That is, the present invention is materialized on a carrier to become a computer program product.
[0123] The present invention also provides a temporal heterogeneous network link prediction device based on hierarchical contrast learning. The device includes at least one processor and a memory communicatively connected to the at least one processor. Wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the temporal heterogeneous network link prediction method based on hierarchical contrast learning, so as to realize the prediction of the dynamic and complex connection relationships between entities in the temporal heterogeneous network.
[0124] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, dedicated ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0125] The computing programs (also referred to as programs, software, software applications, or code) in the present invention include machine instructions of a programmable processor, and these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used in the present invention, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., magnetic disks, optical disks, memories, programmable logic devices PLD) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0126] It should be understood that the various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved, all within the protection scope of the present invention.
Claims
1. A method for predicting temporal heterogeneous network links based on hierarchical contrastive learning, and the formal definition of temporal heterogeneous network link prediction: Heterogeneous network: Given is an undirected graph, where V = {v1, v2, …, v N}, E = {e1, e2, …, e M}, A v and A e represent the node set, edge set, node type, and edge type respectively; each node v i ∈V and edge e j ∈E has a corresponding node type and edge type associated with it, and if |A v | + |A e | > 2, then is a heterogeneous graph; Temporal heterogeneous network: Given an undirected heterogeneous graph It consists of a sequence of heterogeneous network snapshots at different times, that is where is the static snapshot graph of the network at time t, where and represent the node set and edge set at time t respectively, T represents the total number of snapshots, and For any node a ∈ V, its node representation is a feature vector of a fixed size Then X = {x a} a∈V represents the feature matrix of all nodes; The method is used to predict the probability that there exists a target link e=(a, b) between two nodes a, b∈V in a time-series heterogeneous network at the future time T+1, and the user representation is obtained by training a learning model CLP and the similarity between to obtain the probability that e exists in the snapshot graph It is characterized in that an end-to-end model CLP based on contrastive learning is proposed, which obtains the deep representation of the temporal heterogeneous network by learning the heterogeneous information and dynamic dependencies of different types of nodes and edges, so as to express the spatial and temporal differences, and predicts the existence of the target link through the similarity between two target user vector representations; CLP mainly includes three parts: (1) Spatial feature learning, (2) Temporal information modeling, (3) Output layer; (1) Spatial feature learning: First, set the Graph Attention Network (GAT) to retain the heterogeneous structural information of the temporal heterogeneous network at the node level and edge level, so as to obtain the node-level representation vectors of the heterogeneous structure and the edge-level representation vectors At the same time, regard it as a homogeneous network, deploy GNN to capture its inherent structural information and generate the corresponding node-level representation vectors and edge-level representation vectors Perform double-layer contrast learning at the node level and edge level on and and and respectively to enhance the acquisition of network structure differences; (2) Temporal information modeling: Secondly, GRU and LSTM are respectively used to save the hidden states of each static snapshot graph to obtain the temporal-level representation vectors and perform contrastive learning on the node representations and to enhance the capture of temporal differences in the temporal heterogeneous network; (3) Output layer: Finally, calculate the similarity between node a and node b to represent the target link e=(a, b), and feed it into an overall loss function to map the probability of the existence of the target link.
2. A method for predicting temporal heterogeneous network links based on hierarchical contrastive learning according to claim 1, characterized in that Spatial feature learning is specifically as follows: To capture heterogeneous structural features, first, static snapshots are divided into specific sub-networks according to different edge types; subsequently, node-level and edge-level GATs are used to represent different types of edges and nodes in the temporal heterogeneous graph; through the modeling of the two-level GAT, the importance corresponding to different types of nodes and edges is effectively captured, promoting the processing of heterogeneous semantics and structures in the temporal heterogeneous graph; finally, node-level and edge-level contrastive learning are designed to enhance the representation of nodes from the node view and the edge view respectively.
3. The method for predicting time-series heterogeneous network links based on hierarchical contrastive learning according to claim 2, wherein The node-level heterogeneous feature learning is specifically as follows: For each snapshot graph it is divided into |R| subgraphs according to the edge type r ∈ R; then the attention score between any two nodes a and b of the r-type subgraph belonging to the static snapshot at time t is expressed as: where x a and x b are used to initialize nodes a and b; A rt and W rt are the attention weight vector and mapping matrix corresponding to the subgraph of type r, which are updated during the training process; σ represents the activation function, and || is the concatenation operation; then the weight between nodes a and b in the subgraph of type r is calculated as: Among them, represents the r-th type neighbors of node a in the figure, then the representation of node is obtained by weighted summation: The node-level contrastive learning method is used to enhance the representation of each node, ensuring that multiple enhanced representations of the same node have intrinsic similarity, while being different from the representations of other nodes; ignoring the categories of nodes and edges, use GNN to linearly aggregate all unclassified node embeddings:
4. A method for predicting time-series heterogeneous network links based on hierarchical contrastive learning according to claim 3, characterized in that, During the node-level heterogeneous feature learning process, to ensure that the two representations of the same node a aggregated by GAT and GNN respectively and are similar, while the representation of a different node b aggregated by GAT is not similar, the node-level InfoNCE loss function is used to regard as positive sample pairs, as negative sample pairs, as shown in formulas (5) and (6): where τ is a hyperparameter named temperature; sim() is represented by a dot product operation and is used to measure the similarity between two vectors, i.e., |V t |, T, and R represent the graph, number of snapshots, and number of edge types, respectively.
5. A method for predicting temporal heterogeneous network links based on hierarchical contrastive learning according to claim 4, characterized in that, The edge-level heterogeneous feature learning is specifically as follows: Construct an edge-level feature learning module to obtain information from all edge types of each node to obtain the importance weights of different edge types; Aggregate different types of specific edge information to generate node embeddings with heterogeneous edge information. The embedding vector of each node is generated through a non-linear mapping. In the t-th snapshot graph, the attention weight between node a and edge type r is output through the softmax activation function, as shown in formulas (7) and (8): Among them, z, W t and b represent a trainable attention weight vector, a weight matrix, and a bias vector respectively; σ is a non-linear activation function; then, by nodes of specific edge types in performing a weighted sum to obtain its representation: Adopt the edge-level contrastive learning method to enhance the representation of each edge node of a specific type, ensuring intra-similarity for different enhanced representations of the same subgraph and inter-dissimilarity in the comparison with nodes of other types of subgraphs; ignore the node and edge types in graph G t and regard it as an isomorphic network, and use GNN to linearly aggregate the node representations in this isomorphic network: Among them represents the set of neighbors of node a with edge type r in the figure; Set the InfoNCE loss function at the edge level to ensure the embedding of a specific edge type of node a is similar to its embedding in the isomorphic graph while being significantly different from the aggregated representation of node b ; this loss function treats the heterogeneous and isomorphic aggregated representations of the same node as positive sample pairs and the aggregated representations from different nodes in the heterogeneous graph as negative sample pairs as shown in formulas (11) and (12): Among them, represents all adjacent nodes of node a in the figure, covering all edge types.
6. A method for predicting time-series heterogeneous network links based on hierarchical contrastive learning according to claim 5, characterized in that, The temporal information modeling is specifically as follows: Use LSTM and GRU to represent spatio-temporal patterns for learning long-term dependencies and short-term dependencies respectively, denoted as and Among them, represents the embedding of node a in all snapshots at all time points, which is generated from the embedding vectors of node a in all snapshots; design a time-level contrastive learning method to solve the temporal inhomogeneity between different semantic spaces, so as to achieve the proximity of the hidden vector representations obtained by LSTM and GRU. Specifically: Among them, V T is the set of nodes in the last snapshot graph.
7. A method for predicting time-series heterogeneous network links based on hierarchical contrastive learning according to claim 6, characterized in that, The output layer is specifically as follows: Use the last snapshot The embedding vector of node a in is used to perform the link prediction task to infer the potential connections between node i and other nodes; it is transformed into evaluating the similarity between node i and its neighbor nodes in the last snapshot The binary cross-entropy minimization is adopted as the objective function, which is defined as shown in Equation (17): Among them, for the final snapshot graph for any node a in it, its embedding representation is and obtained by mean pooling, that is Meanwhile, in the graph node a's neighbor i is regarded as a positive example, while a randomly selected non-neighbor node j is regarded as a negative example, thus forming a triple (a, i, j); the set is defined as the set of all possible triples, that is Then is the set of positive example sample pairs, is the set of negative example sample pairs; Then, by using the cross-entropy loss node-level contrastive learning loss edge-level contrastive learning loss and the time-level contrastive learning loss weight them to obtain the total loss Among them, λ1, λ2, and λ3 are learnable weight adjustment factors used to coordinate the three losses.
8. A method for predicting time-series heterogeneous network links based on hierarchical contrastive learning according to claim 7, characterized in that, The method is applied to the social recommendation scenario, Regard the social network at time T as a temporal heterogeneous graph Form a sequence of T static heterogeneous snapshots; node set V t Corresponding to users and / or posts, edge set E t Corresponding to likes, comments, and / or follows; make full use of the recent and historical interaction data of the social network through the temporal information modeling layer, capture the long-term and short-term interaction features between social network snapshots, and capture the temporal evolution of user behavior, so as to generate more timely and relevant friend recommendations; through the overall loss function in the output layer of the present invention Train the social network node embedding vectors to optimize the possibility that there is a link e between the node pair (a, b) formed by any two users a and b in the social network.
9. A time-series heterogeneous network link prediction system based on hierarchical contrastive learning, characterized in that: The system has program modules corresponding to the steps of any one of claims 1-8, and executes the steps in the method for predicting temporal heterogeneous network links based on hierarchical contrastive learning when running.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of a method for predicting temporal heterogeneous network links based on hierarchical contrastive learning according to any one of claims 1-8 when called by a processor.
Citation Information
Patent Citations
A method for predicting temporal network links using GraphSAGE
CN111756587B
Link prediction method based on extensible representation of dynamic heterogeneous information network
CN114218446A
Network performance prediction method, performance prediction model training method and related device
CN114900441A