Time sequence knowledge graph balanced reasoning method and system based on evolution contrast learning
By constructing the Balancer model, combining the graph attention network, gated loop unit and comparison learning module, the space-time dependence and temporal characteristics of entities and relationships are captured, and the problem of insufficient prediction capabilities in the timing knowledge graph is solved, and higher prediction accuracy is achieved.
Patent Information
- Application Number
- CN202510512209.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-29
AI Technical Summary
The existing timing knowledge graph model lacks ability to predict brand new events and fails to fully utilize potential temporal features, resulting in limited prediction accuracy in dynamic scenarios.
Build a Balancer model, combine the graph attention network and gated loop unit to capture the space-time dependence of entities and relationships, introduce a comparative learning module to distinguish history from new event features, and capture cycles and cumulative effects through the time information embedding module to improve prediction capabilities.
It significantly improves the prediction accuracy of the timing knowledge graph extrapolation task, especially in the new event prediction, and improves the comprehensiveness and accuracy of the model.
Smart Images

Figure CN120387519A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and particularly relates to a temporal knowledge graph balanced reasoning method and system. Background Art
[0002] As a structured representation tool for real-world knowledge and events, the Knowledge Graph (KG) plays a key role in fields such as recommendation systems, dialogue question answering, and information retrieval. Traditional static knowledge graphs represent facts in the form of triples (s, p, o), where s is the subject entity, o is the object entity, and p is the relationship between entities, which can meet the knowledge representation needs of early static scenarios. However, with the increasing demand for modeling complex dynamic scenarios such as financial transactions, social networks, and public events, static knowledge graphs are difficult to capture the time dependence of facts, and Temporal Knowledge Graphs (TKGs) emerge as the times require. Temporal knowledge graphs represent events in the form of quadruples (s, p, o, t) by introducing the time dimension t. Typical dataset examples include the Global Database of Events, Language, and Tone (GDELT) [1] and the Integrated Crisis Early Warning System (ICEWS) [2], etc. Figure 2 Shows two examples of interrelated historical facts from the ICEWS14 dataset. Different arrows contain the time when the event occurred and the corresponding actions. Temporal knowledge graphs provide dynamic modeling capabilities for tasks such as stock prediction and social relationship evolution analysis.
[0003] The core tasks of temporal knowledge graphs include interpolation (predicting missing facts within a known time interval) and extrapolation (predicting future or historical unknown facts beyond the training time range). Due to its high application value in scenarios such as event prediction and user behavior analysis (such as guiding content distribution strategies), the extrapolation task has become a research hotspot in recent years. Existing extrapolation methods mainly achieve prediction by modeling the structure and time dependence of historical temporal knowledge graph sequences. For example, RE-GCN [3] combines the Graph Convolutional Network (GCN) with the Recurrent Neural Network (RNN) to capture the entity neighborhood structure and time series dynamics; L2TKG [4] mines the implicit associations between nodes through a latent relationship learning module; xERTE [5] predicts target entities based on a subgraph search strategy. These methods rely on the repetitive patterns of historical events and perform excellently in predicting high-frequency historical events, but face the following core challenges:
[0004] (1)Insufficient ability to predict completely new events: In real-world scenarios, user behaviors or entity relationships may give rise to new events that have never occurred in history due to factors such as environmental changes and public opinion influence (e.g., sudden technological innovations, unconventional social behaviors). Existing models overly rely on "repetitive pattern learning" of historical data and lack an explicit modeling mechanism for completely new events. For example, when predicting low-frequency new events such as "a user's first purchase of a certain type of product", due to the lack of corresponding patterns in the training data, the model is prone to overlooking their potential correlations, resulting in prediction biases.
[0005] (2)Insufficient utilization of potential time features: Most existing methods treat timestamps as simple embeddings or intervals and do not effectively capture the periodicity (e.g., annual promotional activities) or cumulativeness (e.g., the evolution of user preferences over time) of events, leading to limited accuracy in long-term trend prediction.
[0006] To address the above two core challenges in the existing technology, the present invention proposes a balanced framework that integrates historical pattern learning and contrastive reasoning. By introducing a contrastive learning module to distinguish historical and completely new event features and combining a time information embedding module to capture periodic and cumulative effects, while ensuring the prediction accuracy of historical events, the ability to discover new events is significantly improved. This solution fills the gap in the inference of unconventional events in dynamic scenarios by existing models and provides a more comprehensive solution for the extrapolation task of temporal knowledge graphs. Summary of the Invention
[0007] In view of the above situation, the purpose of the present invention is to provide a method and system for balanced inference of temporal knowledge graphs based on evolutionary contrastive learning, which solves the problems of insufficient ability of existing models to predict completely new events and insufficient utilization of potential time features, and improves the overall prediction accuracy of the extrapolation task of temporal knowledge graphs.
[0008] The method for balanced inference of temporal knowledge graphs based on evolutionary contrastive learning proposed by the present invention constructs a balanced framework that integrates historical pattern learning and contrastive reasoning. By introducing a contrastive learning module to distinguish historical and completely new event features and combining a time information embedding module to capture periodic and cumulative effects, while ensuring the prediction accuracy of historical events, the ability to discover new events is significantly improved. The specific steps are as follows:
[0009] Step 1: Construct a historical pattern learning module. The historical pattern learning module captures the dependencies of entities and relationships in the spatio-temporal dimension through a Graph Attention Network (GAT) and a Gated Recurrent Unit (GRU);
[0010] Step 2: Construct a contrastive learning module. Use a supervised contrastive learning module to generate contrastive representations of historical events and completely new events, and train a binary classifier to generate auxiliary vectors;
[0011] Step 3: Construct a time information embedding module. The time information embedding module generates periodic time vectors and aperiodic time vectors, and extracts the periodic and cumulative features in the dataset;
[0012] Step 4: Construct a decoder module, the structure of which is as Figure 4 shown. The decoder module uses the embedding vectors of entities, relationships, and time information for training after convolution operations and splicing processing, and generates prediction vectors in the prediction phase;
[0013] Step 5: Combine the prediction vectors generated by the decoder and the auxiliary vectors generated by the contrast learning module to obtain the prediction results of the temporal knowledge graph extrapolation task.
[0014] In the present invention:
[0015] The specific process of Step 1 is as follows:
[0016] Step 1-1: Single-step aggregation. The graph attention network (GAT) is used to capture the influence of the neighbor nodes of entities at each time step, and learn the entity behavior and network structure features;
[0017] Step 1-2: Multi-step aggregation. The gated recurrent unit (GRU) is used to integrate the embedding representations of all time steps, capture the temporal dependencies of entities and relationships, and simulate the time decay effect of historical events;
[0018] Step 1-3: Aggregate all historical information at time step t-1 to generate the final embedding representations of entities and relationships for subsequent prediction;
[0019] The specific process of Step 2 is as follows:
[0020] Step 2-1: Define its historical dataset S through the given query q=(s, p,?, t) t and frequency
[0021] Step 2-2: Encode the embedding vectors for contrast learning through a multi-layer perceptron (MLP) to generate the contrast representations of historical events and new events;
[0022] Step 2-3: Use the contrast representations of historical events and new events to train a binary classifier; generate auxiliary vectors in the prediction phase to assist in correcting the prediction results and enhancing the attention to new events.
[0023] The specific process of Step 3 is as follows:
[0024] Step 3-1: Introduce a time information embedding module, design periodic and aperiodic time vectors to capture the periodic features (such as event repetition period) and cumulative features (such as event probability increasing over time) in historical facts respectively;
[0025] Step 3-2: The periodic and non-periodic time vectors are generated through sine function and linear function respectively to adapt to the time characteristics of different datasets.
[0026] The specific process of Step 4 is as follows:
[0027] Step 4-1: Based on the T-convKB decoder, the embedding vectors of entities, relations and time information are convolved to generate the final prediction score, which is used to infer the network structure at the next time step t;
[0028] The specific process of Step 5 is as follows:
[0029] Step 5-1: Combine the prediction vector generated by the decoder and the auxiliary vector generated by the contrast learning module to obtain the prediction result of the temporal knowledge graph extrapolation task.
[0030] In summary, the present invention proposes a Balancer model, the structure of which is as Figure 3 shown. The model includes a historical pattern learning module, a contrast learning module, a time information embedding module and a decoder module; the historical pattern learning module captures the spatio-temporal dependencies of entities and relations through a graph attention network and a gated recurrent unit; the contrast learning module uses supervised contrast learning to distinguish historical and brand-new events and generates an auxiliary prediction vector; the time information embedding module combines periodic and non-periodic functions to extract time features; the decoder module is trained according to the learned embedding vectors and provides prediction results in the prediction phase.
[0031] The innovation of the present invention lies in: the present invention provides a temporal knowledge graph balanced reasoning method and system based on evolutionary contrast learning, which solves the problems of insufficient prediction ability of existing models for brand-new events and insufficient utilization of potential time features. In particular, a novel Balancer model is proposed, which can not only accurately predict conventional historical events, but also effectively mine potential brand-new events in the future. In order to enhance the ability of the model to distinguish event features, a supervised contrast learning module is innovatively introduced, and a new contrast learning embedding generation method is proposed, so that the embedding representation contains richer information. At the same time, the present invention designs a time information embedding module for capturing periodic features and cumulative features in historical facts. The present invention realizes the optimization in historical event prediction and potential brand-new event discovery, and significantly improves the comprehensiveness and accuracy of temporal knowledge reasoning. Experiments show that the Hits@1 index of the present invention on the YAGO, ICEWS14, and ICEWS18 datasets is improved by 21.12%, 23.21%, and 51.40% respectively, effectively improving the prediction accuracy of the temporal knowledge graph extrapolation task and enhancing the prediction ability for brand-new events. Brief Description of the Drawings
[0032] Figure 1 This is a flowchart of the balanced reasoning method for the temporal knowledge graph of the present invention.
[0033] Figure 2 This is a schematic diagram of historical facts in the temporal knowledge graph dataset ICEWS14, showing events at different time steps and their behavioral relationships.
[0034] Figure 3 This is the overall architecture diagram of the Balancer model.
[0035] Figure 4 This is a schematic diagram of the training process of the T-convKB decoder, showing the calculation process of obtaining the prediction score through the embedding vector. Detailed implementation manners
[0036] The present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0037] The balanced reasoning method for the temporal knowledge graph provided by the present invention has the specific process as Figure 1 shown. The specific steps are as follows:
[0038] Step 1: Construct a historical law learning module, as Figure 3 shown. The historical law learning module captures the dependencies of entities and relationships in the spatio-temporal dimension through the Graph Attention Network (GAT) and the Gated Recurrent Unit (GRU). The specific process is as follows:
[0039] Step 1-1: Single-step aggregation. Referring to Figure 3 shown, the historical law learning module captures the influence of the neighbor nodes of the entity at each time step through the graph attention network, and learns the entity behavior and network structure features. Among them, the attention coefficient α ij of node j on node i can be calculated by formula (1):
[0040]
[0041] where e is the embedding representation of the entity, W is the trainable parameter matrix, T is the transpose operation, is the concatenation operation, a is a learnable vector for determining the attention direction, N i is the set of all neighbor nodes of entity i, is the activation function LeakyReLU, and its calculation method is as formula (2):
[0042]
[0043] Among them, b is a relatively small constant, and its value usually ranges between 0.01 and 0.3. This method uses the multi-head attention mechanism to calculate the attention weights, making the attention mechanism more scalable. Then, the features of each attention head are aggregated and averaged to obtain the final entity features. The finally aggregated entity embedding representation e i Calculated by formula (3):
[0044]
[0045] Among them, W k is a trainable parameter matrix, K is the number of attention heads, σ is the activation function sigmoid, and its calculation method is as formula (4):
[0046]
[0047] Step 1-2: Multi-step aggregation. See Figure 3 As shown in, the historical pattern learning module uses gated recurrent units to integrate the embedding representations of all time steps, capture the temporal dependencies of entities and relationships, and simulate the time decay effect of historical events. At time step t, first calculate the update gate z of GRUs using the following formulas (5) and (6) t and y t :
[0048]
[0049] Among them, W z 、W y 、b1 and b2 are learnable parameters, contains the information of the previous time step t-1, represents the hidden layer state of the entity at the current step. The update gate adds these two pieces of information to the activation function σ, that is, sigmoid. Therefore, the final embedding at time step t contains two parts, namely the current step embedding and the final embedding information of time step t-1 The update gate function determines the proportion of each retention. The hidden layer state of the entity at time step t can be calculated by the following formula (7):
[0050]
[0051] Among them, W e is a learnable parameter, ⊙ is the Hadamard product, tanh is the activation function, and its calculation method is as the following formula (8):
[0052]
[0053] The final embedding at time step t It can be calculated by the following formula (9):
[0054]
[0055] At the same time, a relation-specific GRU is used to update the relation. Since the embedding of the relation at time t is affected by the evolved embeddings of V r,t,s and V r,t,o at time t and their embedding information at time t-1. Among them, V r,t,s and V r,t,o are defined by the following formulas (10) and (11):
[0056] V r,t,s = {i|(i, p, o, t) ∈ K t}, #(10)
[0057] V r,t,o = {i|(s, p, i, t) ∈ K t},#(11)
[0058] Considering that entities usually affect relations in the form of triples, the evolution of the relation embedding [[ID=3,5]]is defined as the following formula (12):
[0059]
[0060] Among them, pooling is the mean pooling operation, and represent the set of head and tail node embedding representations of the triples containing relation r at the time step t-1, and W r is a learnable parameter. and After the mean pooling operation, they are concatenated with r in the form of a triple. represents the relation embedding vector t-1 that aggregates the evolved embedding features of entities at the time step. Finally, the relation embedding matrix is updated across time steps through the gated recurrent unit GRU as follows:
[0061]
[0062] R t = GRU(R t-1 , R t′ ), #(14)
[0063] Among them, R t-1 is the relation embedding matrix t-1 that aggregates and contains the temporal information up to the time step, and R t′ represents the aggregated embedding set of all relation types.
[0064] Step 1-3: Aggregate all historical information at time step t-1 to generate the final embedding representations of entities and relationships for subsequent tasks.
[0065] Step 2: Construct a contrastive learning module, as shown in Figure 3 . Use the supervised contrastive learning module to generate the contrastive representations of historical events and new events, and train a binary classifier to generate auxiliary vectors. The specific process is
[0066] Step 2-1: Define its historical dataset S through the given query q=(s,p,?,t) t and frequency The calculation methods are shown in formulas (15) and (16):
[0067] S t ={o|(s,p,o,j)∈G j≤t}, #(15)
[0068]
[0069] where G j≤t represents the quadruple of the subject s, relationship p, and object o at time step t and before. S t represents the set of objects o included in the corresponding historical event set in this query. is the number of o that meet the conditions so far.
[0070] Step 2-2: As shown in Figure 3 , the contrastive learning module encodes the embedding vectors for contrastive learning through a multi-layer perceptron (MLP), and the encoding method is shown in formula (17):
[0071]
[0072] where the variable t l represents the time difference between the time point when the event last occurred and the current prediction time point. The introduction of t l considers the time decay characteristic of the event influence. represents the frequency of o becoming the tail node. Matrices W m and W n are trainable mapping matrices. represents the encoder. Furthermore, the supervised contrastive learning module can be used to generate the contrastive representations of historical events and new events. The loss function of contrastive learning can be calculated by the following formula (18):
[0073]
[0074] Among them, τ = 0.1 is the temperature parameter selected to obtain the best performance. u is the embedding vector for contrastive learning. M represents the training batch. C(q) represents the set of Boolean values of the batch M excluding q, and its value is the same as I(q), specifically expressed as formula (19):
[0075]
[0076] Step 2-3: Using the contrastive representation of historical events and new events, train a binary classifier to generate auxiliary vectors to assist in correcting the prediction results and enhancing the attention to new events. The loss function for training the binary classifier can be calculated by the following formula (20):
[0077]
[0078] Among them, y i represents the label of the sample, divided into 1 and 0. p i represents the probability that sample i is predicted to be positive.
[0079] Step 3: Construct a time information embedding module, as shown in Figure 3 . The time information embedding module generates periodic time vectors and aperiodic time vectors, and extracts the time periodicity and cumulative characteristics in the dataset.
[0080] Step 3-1: Introduce the time information embedding module. As shown in Figure 3 , the time information embedding module generates periodic and aperiodic time vectors, capturing the periodic characteristics (such as event repetition period) and cumulative characteristics (such as event probability increasing over time) in historical facts respectively. The periodic and aperiodic time vectors are generated by sine function and linear function respectively to adapt to the time characteristics of different datasets. The specific construction method is as follows in formulas (21) and (22):
[0081]
[0082] Among them, and represent the d-dimensional periodic time vector and aperiodic time vector respectively. w p , w np , θ p , θ np are learnable parameters.
[0083] Step 4: Construct a decoder module, as shown in Figure 3 . Figure 4 Then shows the specific structure of the decoder. The decoder module uses the embedding vectors of entities, relationships, and time information for training after convolution operations and splicing, and generates prediction vectors in the prediction phase.
[0084] Step 4-1: Refer to Figure 4 As shown, the decoder based on T-convKB performs convolution processing on the embedding vectors of entities, relationships, and time information to generate the final prediction score, where the scoring function is calculated by the following formula (23):
[0085]
[0086] where Ω and ω are trainable shared parameters, g(x) = |x| is the activation function, * represents the convolution operator, ω is a filter, concat represents multiple concatenation operations, and represent the d-dimensional periodic time vector and the aperiodic time vector respectively, and e s , r, e represent the embedding representations of the head node, relationship, and tail node respectively. Use the Adam optimizer and train T-convKB by minimizing the regularization loss function on the weight vector w of the model, which can be calculated by the following formula (24):
[0087]
[0088] X = (s, p, o) ∈ {G ∪ G'}, #(25)
[0089] where, G ′ is the set of invalid triples generated by destroying the valid triples in G. The decoder can generate the prediction vector for inferring the network structure at the next time step t.
[0090] Step 5: Combine the prediction vector generated by the decoder and the auxiliary vector generated by the contrastive learning module to obtain the prediction result of the temporal knowledge graph extrapolation task.
[0091] Step 5-1: Combine the prediction vector generated by the decoder and the auxiliary vector generated by the contrastive learning module to obtain the prediction result of the temporal knowledge graph extrapolation task. The prediction score is calculated by the following formula (26):
[0092] FinalScore = βScore m + (1 - β)Score t , #(26)
[0093] where, Score m is the auxiliary vector generated by the contrastive learning module, Score t = f score is the prediction result generated by the decoder, and β is the hyperparameter used to adjust the weights of the two.
[0094] The overall performance of this method is verified through Experiment 1 and Experiment 2. Experiment 1 is conducted on three standard temporal knowledge graph datasets, namely YAGO, ICEWS14, and ICEWS18. Experiment 2 is conducted on three brand-new event datasets after preprocessing. Among them, ICEWS14 and ICEWS18 contain political and diplomatic events, and YAGO is a general knowledge graph. All datasets are divided into training sets, validation sets, and test sets in a ratio of 8:1:1 according to time. The specific information of the datasets is shown in Table 1. The construction method of the brand-new event datasets in Experiment 2 is as follows: Filter out all data that appear in the test dataset but not in the training dataset and the validation dataset (referred to as "brand-new events"), and then use these datasets to construct a new test set. The experiment compares static and temporal knowledge graph embedding models, and uses MRR and Hits@1 / 3 / 10 as evaluation metrics. Their calculation methods are shown in the following formulas (27) and (28):
[0095]
[0096] Here, S represents the set of test triples, rank i is the predicted ranking of the i-th triple, and I(·) is the indicator function, which returns 1 if the condition is true and 0 otherwise. These metrics are averaged over five independent runs to ensure the robustness of the results. The larger the values of these two metrics, the better the model performance. At the same time, the dimensions of the embedding vectors of entities and relationships are set to d = 200, the number of GAT layers is 2, the number of attention heads is 4, and the Adam optimizer is selected (learning rate ∈ = 1e -4 ), and the batch size is 256.
[0097] The results of Experiment 1 are shown in Table 2. The results show that this method performs optimally in most tasks. For the Hits@1 metric, compared with the second-best model, the accuracy rates on YAGO, ICEWS14, and ICEWS18 are increased by 21.12%, 23.21%, and 51.40% respectively; for the Hits@3 metric, the accuracy rates of YAGO and ICEWS18 are increased by 16.14% and 16.76% respectively. Although the performance is not outstanding in some Hits@3 and Hits@10 metrics, this method significantly leads in the more valuable Hits@1 metric.
[0098] The results of Experiment 2 are shown in Table 3. The results indicate that in most metrics, this method outperforms other baseline models. For example, in the Hit@1 metric, the performance of this method on three completely new event datasets is 18.10%, 4.26%, and 14.66% higher than that of the second-ranked model respectively. These results show that this method is better at predicting completely new events than other models. The above results verify the effectiveness of the contrastive learning module and the temporal information embedding module of this method, especially the advantage in the key metric Hits@1, indicating that this model has significant practicality in the extrapolation task of temporal knowledge graphs, especially in the prediction task of completely new events.
[0099] Table 1. Specific information and parameters of YAGO, ICEWS14, and ICEWS18 datasets
[0100]
[0101] Table 2. Comparison of Experiment 1 results, showing the performance differences between the present invention and baseline methods on three standard datasets
[0102]
[0103] Table 3. Comparison of Experiment 2 results, showing the performance differences between the present invention and baseline methods on three completely new event datasets.
[0104]
[0105] References
[0106] [1] Kalev Leetaru and Philip A Schrodt. Gdelt: Global data on events, location, and tone, 1979–2012. In ISA Annual Convention, pages 1 49, 2013.
[0107] [2] Alberto Sebastijan and Mathias Niepert. Learning sequence encoders for temporal knowledge graph completion. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4816–4821, 2018.
[0108] [3]Zixuan Li,Xiaolong Jin,Wei Li,Saiping Guan,Jiafeng Guo,HuaweiShen,Yuanzhuo Wang,and Xueqi Cheng.Temporal knowledge graph reasoning basedon evolutional representation learning.In Proceedings of the 44thInternational ACM SIGIR Conference on Research and Development in InformationRetrieval,pages 408–417,2021.
[0109] [4]Mengqi Zhang,Yuwei Xia,Qiang Liu,Shu Wu,and Liang Wang.Learninglatent relations for temporal knowledge graph reasoning.In Proceedings of the61st Annual Meeting of the Association for Computational Linguistics(Volume1:Long Papers),pages 12617 12631,2023.
[0110] [5]Zhen Han,Peng Chen,Yunpu Ma,and Volker Tresp.Explainable subgraphreasoning for forecasting on temporal knowledge graphs.In InternationalConference on Learning Representations,pages 1–24,2020.
[0111] [6] Seyed Mehran Kazemi, Rishab Goel, Sepehr Eghbali, Janahan Ramanan, Jaspreet Sahota, Sanjay Thakur, Stella Wu, Cathal Smyth, Pascal Poupart, and Marcus A. Brubaker. Time2vec: Learning a vector representation of time. CoRR, abs / 1907.05321: 1–16, 2019.
[0112] [7] Elizabeth Boschee, Jennifer Lautenschlager, Sean O’Brien, Steve Shellman, James Starz, and Michael Ward. ICEWS coded event data, 2015。
Claims
1. A temporal knowledge graph balanced reasoning method based on evolutionary contrastive learning, characterized in that, Construct a balanced framework that integrates historical law learning and contrastive reasoning. By introducing a contrastive learning module to distinguish the features of historical and brand-new events, and combining a time information embedding module to capture periodic and cumulative effects, while ensuring the prediction accuracy of historical events, improve the discovery ability of new events. The specific steps are as follows: Step 1: Construct a historical law learning module to capture the dependencies of entities and relationships in the spatio-temporal dimension through a Graph Attention Network (GAT) and a Gated Recurrent Unit (GRU). The specific process is as follows: Step 1-1: Single-step aggregation; capture the influence of the neighbor nodes of the entity at each time step through a Graph Attention Network (GAT) to learn the entity behavior and network structure features. Step 1-2: Multi-step aggregation; Use a Gated Recurrent Unit (GRU) to integrate the embedding representations of all time steps, capture the temporal dependencies of entities and relationships, and simulate the time decay effect of historical events. Step 1-3: Aggregate all historical information at time step t-1 to generate the final embedding representations of entities and relationships for subsequent prediction. Step 2: Construct a contrastive learning module. Use a supervised contrastive learning module to generate contrastive representations of historical events and brand-new events, and train a binary classifier to generate auxiliary vectors. The specific process is as follows: Step 2-1: Define its historical dataset S through the given query q = (s, p,?, t) t and frequency Step 2-2: Encode the embedding vectors used for contrastive learning through a Multi-Layer Perceptron (MLP) to generate contrastive representations of historical events and brand-new events. Step 2-3: Use the contrastive representations of historical events and brand-new events to train a binary classifier; generate auxiliary vectors in the prediction stage to assist in correcting the prediction results and improve the attention to brand-new events. Step 3: Construct a time information embedding module to generate periodic time vectors and aperiodic time vectors, and extract the time periodicity and cumulative features in the dataset. The specific process is as follows: Step 3-1: Introduce a time information embedding module and design periodic and aperiodic time vectors to capture the periodic features and cumulative features in historical facts respectively. Step 3-2: The periodic and aperiodic time vectors are generated through sine functions and linear functions respectively to adapt to the time characteristics of different datasets. Step 4: Construct a decoder module. After convolution operations and concatenation processing on the embedding vectors of entities, relationships, and time information, use them to train the decoder to generate prediction vectors. Specifically: Based on the T-convKB decoder, perform convolution processing on the embedding vectors of entities, relationships, and time information to generate the final prediction scores for inferring the network structure at the next time step t. Step 5: Combine the prediction vectors generated by the decoder and the auxiliary vectors generated by the contrastive learning module to obtain the prediction results of the temporal knowledge graph extrapolation task.
2. The sequential knowledge graph balanced reasoning method according to claim 1, wherein In Step 1: In step 1-1, the attention coefficient α of node j to node i ij is calculated by formula (1): Among them, e is the embedding representation of the entity, W is the trainable parameter matrix, T is the transpose operation, ⊕ is the concatenation operation, a is a learnable vector for determining the attention direction, and N i is the set of all neighbor nodes of entity i, is the activation function LeakyReLU, and its calculation method is as shown in formula (2): Among them, b is a relatively small constant, and its value ranges between 0.01 and 0.3; then, the features of each attention head are aggregated and averaged to obtain the final entity features; the finally aggregated entity embedding representation e i Calculated by formula (3): where W k is a trainable parameter matrix, K is the number of attention heads, and σ is the sigmoid activation function, and its calculation method is as shown in Equation (4): In step 1-2, at time step t, first calculate the update gate z of the GRUs using equations (5) and (6) t and y t : Among them, W z , W y , b1 and b2 are learnable parameters, contains the information t-1 of the previous time step, represents the hidden layer state of the entity at the current step; the update gate adds these two pieces of information to the activation function σ, that is, sigmoid; the final embedding at time step t consists of two parts, namely the current step embedding and the final embedding information at time step t-1 The update gate function determines the retention ratio of each; the hidden layer state of the entity at time step t is calculated by the following formula (7): where, W e is a learnable parameter, ⊙ is the Hadamard product, and tanh is the activation function, and its calculation method is as follows in formula (8): Final embedding at time step t Calculated by the following formula (9): Update the relationship using a relationship-specific GRU at the same time; since the relationship embedding at time t is affected by V r,t,s and V r,t,o evolved embedding at time t and its embedding information at time t - 1; where V r,t,s and V r,t,o are defined by the following formulas (10) and (11): Since the entity affects the relationship in the form of a triple, the evolution of the relationship embedding is defined as follows in Equation (12): Among them, pooling is the mean pooling operation, and represent the set of head and tail node embedding representations of the triples containing the relationship r at the time step t - 1, and W r are learnable parameters; and After the mean pooling operation, they are concatenated with r to form a triple; represents the relationship embedding vector t - 1 that aggregates the embedding features of entity evolution at the time step; finally, the relationship embedding matrix is updated across time steps through a gated recurrent unit (GRU) as follows: Among them, R t-1 is a relational embedding matrix Rt-1 that aggregates and includes temporal information up to and including the time step t′ represents the aggregated embedding set of all relation types.
3. The sequential knowledge graph balanced reasoning method according to claim 2, wherein In Step 2: In step 2-1, the historical data set S t and the frequency are calculated as shown in equations (15) and (16): Among them, G j≤t represents the quadruple of subject s, relation p, and object o at time step t and before; S t represents the set of objects o contained in the corresponding historical event set in this query; is the number of o that meet the conditions up to the current time step; In Step 2-2, encode the embedding vectors used for contrastive learning through a Multi-Layer Perceptron (MLP). The encoding method is as follows: Among them, the variable t l represents the time difference between the time point when the event last occurred and the current prediction time point, represents the frequency at which o becomes the tail node; the matrix W m and W n are trainable mapping matrices; represents the encoder; furthermore, a supervised contrast learning module is adopted to generate the contrast representation of historical events and new events; the loss function of contrast learning is calculated as shown in Equation (18): Among them, τ = 0, 1 are temperature parameters selected to obtain the best performance; u is the embedding vector used for contrastive learning; M represents the training batch; C(q) represents the set of boolean values of the batch M excluding q, and its value is the same as I(q), which is specifically expressed as Equation (19): In Step 2-3, the loss function for training the binary classifier is calculated as follows in Equation (20): where y i represents the label of the sample, which is divided into 1 and 0; p i represents the probability that sample i is predicted as positive.
4. The sequential knowledge graph balanced reasoning method according to claim 3, wherein In step 3, the periodic and aperiodic time vectors are generated by a sine function and a linear function respectively, as shown in the following formulas (21) and (22): wherein, and represent a d-dimensional periodic time vector and an aperiodic time vector respectively, and w p , w np , θ p , θ np are learnable parameters.
5. The sequential knowledge graph balanced reasoning method according to claim 4, wherein In step 4, the final prediction score is generated, and this score is calculated by the following formula (23): Among them, Ω and ω are trainable shared parameters, g(x) = |x| is the activation function, * represents the convolution operator, ω is a filter, and concat represents multiple concatenation operations. and respectively represent the d-dimensional periodic time vector and the aperiodic time vector, and e s , r, and e respectively represent the embedding representations of the head node, the relation, and the tail node; Use the Adam optimizer and train T-convKB by minimizing the regularized loss function on the weight vector w of the model, which is calculated by Equation (24) below: X = (s, p, o) ∈ {G ∪ G'}, #(25) Among them, G' is a set of invalid triples generated by destroying the valid triples in G; the decoder generates a prediction vector for inferring the network structure at the next time step t.
6. The sequential knowledge graph balanced reasoning method according to claim 5, wherein In step 5, the prediction result of the temporal knowledge graph extrapolation task is obtained by combining the prediction vector generated by the decoder and the auxiliary vector generated by the contrast learning module, and the prediction score is calculated by the following formula (26): FinalScore = βScore m +(1 - β)Score t ,#(26) Among them, Score m is an auxiliary vector generated by the contrastive learning module, and Score t = f score is the prediction result generated by the decoder, and β is a hyperparameter used to adjust the weights of the two.
7. An ordered knowledge graph balanced reasoning system based on the ordered knowledge graph balanced reasoning method according to claim 5, characterized in that It includes a historical pattern learning module, a contrast learning module, a time information embedding module, and a decoder module; the historical pattern learning module captures the spatio-temporal dependencies of entities and relationships through a graph attention network and a gated recurrent unit; the contrast learning module uses supervised contrast learning to distinguish historical and brand-new events and generates an auxiliary prediction vector; the time information embedding module extracts time features by combining periodic and aperiodic functions; The decoder module is trained based on the learned embedding vectors and provides prediction results in the prediction phase.