Intelligent completion method based on dual-tense fusion perception
By constructing an intelligent completion method based on dual-temperature fusion perception, the problem of insufficient utilization of time information in the timing knowledge graph is solved, and more efficient and accurate completion effect is achieved, improving the robustness and integrity of the timing knowledge graph.
Patent Information
- Application Number
- CN202510066502.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-07-22
AI Technical Summary
The existing timing knowledge graph completion model fails to effectively utilize the sequence diachrony of time information, resulting in insufficient accuracy and robustness in the completion process.
Using an intelligent completion method based on bitemporal fusion perception, a diachronic and synchronous temporal perception module is constructed, combining single-hot encoding and regularization technology, an intelligent completion model based on quaternion is established, and the Hamilton operator is used to rotate relationships and timestamps, and the model is optimized using scoring functions and loss functions.
It improves the completion accuracy and robustness of the timing knowledge graph, can better capture the relationship attributes and development laws in the real world, and enhances the integrity and prediction capabilities of the model.
Smart Images

Figure CN120354060A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of Internet knowledge graphs, and particularly relates to an intelligent completion method based on dual-temporal fusion perception. Background Art
[0002] A knowledge graph is a technology that describes entities, events in the real world and their interrelationships in a structured form, generally divided into a static knowledge graph and a temporal knowledge graph. Since the real world contains time information, the temporal knowledge graph can more accurately describe events in the real world and has higher applicability and expressiveness.
[0003] The temporal knowledge graph stores facts in the form of quadruples. Generally, the form of a quadruple is: head entity, relation, time, tail entity. Existing knowledge bases (such as ICEWS, NELL, and OpenIE) contain a large amount of complex information and have been successfully applied in various fields, including intelligent prediction, information retrieval, and search engines. However, due to the extensiveness and complexity of real-world facts, the knowledge graph evolving over time cannot completely capture the exponentially growing knowledge. To address these challenges, researchers have widely focused on the knowledge graph completion task and developed various robust models. However, most of these models focus on static knowledge graphs, and less attention has been paid to temporal knowledge graphs that contain time information. At the same time, compared with static knowledge graphs, temporal knowledge graphs are closer to real-world scenarios, making the completion process more challenging.
[0004] For temporal knowledge graph completion, it mainly focuses on completing the case where a quadruple (head entity, relation, time, tail entity) is missing, that is, the completion task for cases such as (head entity, relation, time,?).
[0005] In the process of implementing the present invention, the inventor found that existing research on temporal knowledge graph completion uses time information as a supplement and fails to perceive the fusion characteristics and development trajectories of objective facts from the sequential diachronicity of time itself. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and provide an intelligent completion method based on dual-temporal fusion perception to achieve lightweight and efficient intelligent temporal graph completion.
[0007] The present invention solves its technical problems through the following technical solutions:
[0008] An intelligent completion method based on dual-temporal fusion perception, the steps of the method are as follows:
[0009] S1. Divide the dataset to be completed into a training set and a test set according to a ratio;
[0010] S2. Perform structured data preprocessing on the dataset, and successively perform entity extraction, relationship extraction, time information annotation and standardization, entity unification, and coreference resolution on the dataset;
[0011] S3. Encode the preprocessed structured dataset using one-hot encoding;
[0012] S4. Establish an intelligent completion model based on dual-temporal fusion perception, including a scoring function, regularization, and a loss function;
[0013] S5. Set hyperparameters for the intelligent completion model, including the learning rate, batch size, entity embedding dimension, relationship embedding dimension, time embedding dimension, and maximum number of iterations;
[0014] S6. Use the training set to train the above intelligent completion model, and substitute the data in the training set into the loss function of the intelligent completion model until the loss function of the model converges;
[0015] S7. First, apply the trained model to the test set, score each quadruple in the test set through the scoring function in the model, and the score of each quadruple reflects the prediction possibility of the quadruple in the model or the matching degree with historical data. The model automatically evaluates the rationality of each quadruple by calculating the score;
[0016] S8. Conduct experimental evaluation and verification on the trained intelligent completion model, and comprehensively evaluate the performance of the model on the test set by calculating and comparing the results of evaluation metrics such as MRR, Hit@1, Hit@3, and Hit@10.
[0017] Moreover, the entity extraction in S2 is to identify meaningful entities from the text, including person names, place names, and dates. The relationship extraction is quadruple extraction, that is, a dataset is represented as a set of head entity, relationship, time, and tail entity. The time information annotation and standardization refer to annotating the time information and converting it into a unified format. The entity unification means unifying entities with different representations into a standard expression. The coreference resolution aims to solve the actual entity pointed to by pronouns or other referential words in the text.
[0018] Moreover, the specific operation of step S3 is as follows: Each entity e in the structured dataset i is represented by an f-dimensional one-hot encoded binary vector, making the i-th element of and the i-th element of Set the j-th element to 1 and the rest to 0; for each timestamp, first map it to a one-hot encoded binary vector of a fixed length according to the type of the timestamp, where the index position of the timestamp is 1 and the other positions are set to 0, thus completing the encoding of time information. After the encoding processes of all entities, relationships, and timestamps are completed, a complete representation of the structured data is obtained.
[0019] Moreover, the intelligent completion model based on dual-temporal fusion perception in S4 includes a diachronic temporal perception module and a synchronic temporal perception module;
[0020] The diachronic temporal perception module is responsible for processing the perception of diachronic time, and the specific process is as follows:
[0021] Construct a diachronic timestamp, map each time point to an angle in sequence, and use trigonometric functions to convert it into the scalar component of a quaternion:
[0022]
[0023] Use the Hamilton operator with the diachronic timestamp to rotate the relationship R r by quaternion into a diachronic relationship
[0024]
[0025] Normalize to a unit quaternion. Rotate the head entity Q s by doing a Hamilton operator between and the normalized s :
[0026]
[0027] The synchronic temporal perception module is responsible for processing the perception of synchronic time, and the specific process is as follows:
[0028] Jointly reorganize the synchronic timestamp and the relationship embedding quaternion to obtain two new quaternions:
[0029]
[0030] Use the Hamilton operator to perform a rotation operation on through , where is the normalized representation of
[0031]
[0032] Rotate the head entity Q sand perform a Hamilton operator between them to rotate the head entity Q s :
[0033]
[0034] The scoring function is expressed as:
[0035]
[0036] The regularization and loss function use tensor nuclear norm regularization to enhance the temporal knowledge graph, and the following regularization is proposed:
[0037]
[0038] For the nuclear 4-norm, for time constraints, a time regularizer is usually used to smooth the representations of adjacent timestamps. The following linear time regularizer is used:
[0039]
[0040] where: W b is the linear time constraint bias, randomly initialized and learned during training;
[0041] The loss function is expressed as:
[0042]
[0043] The advantages and beneficial effects of the present invention are:
[0044] 1. The method of fusing and perceiving the diachronic embedding and synchronic embedding of timestamps in the present invention can accurately capture the relationship attributes and development laws of the real world, making the temporal knowledge graph have better robustness and integrity.
[0045] 2. The present invention summarizes the key features of temporalized facts from two aspects: diachrony and synchrony: 1) Diachronic tense: Facts often exhibit different characteristics and development trends in different time periods. For example, assume that in March 2023, Alice and Bob are friends. As time goes by, in May and June, Alice and David start to build a closer relationship, and in July, their relationship smoothly develops into a romantic couple. At the same time, Alice and Bob still maintain an ordinary friendship and there is no significant change. 2) Synchronous tense: In a specific time context, various relationships between entities usually interact and influence each other, thus forming potential semantic connections. For example, in May and June 2023, the relationship between Alice and David has significantly warmed up, showing progressive signs of intimacy. At the same time, the interaction between Alice and Bob is relatively scarce, and there is only one simple communication during this period, and their relationship still remains in the state of ordinary friends. The above two tenses are not completely independent, and they often rely on each other and are mutually based. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is the model architecture diagram of the present invention;
[0047] Figure 2 is the flowchart of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The following will make a detailed description of an intelligent completion method based on dual-temporal fusion perception of the present invention in combination with embodiments and drawings.
[0049] As Figure 1 , 2 shown, an intelligent completion method based on dual-temporal fusion perception of the present invention is innovative in that it includes the following steps:
[0050] 1) Dataset division: Divide the original dataset into a training set (T) and a test set (S) in a ratio of 8:2 for model training and verification;
[0051] 2) Data preprocessing: Use the Language Technology Platform (LTP) to perform structured processing on the dataset, specifically including:
[0052] Entity extraction: Identify entities in the text;
[0053] Relationship extraction: Extract relationships between entities;
[0054] Time annotation and standardization: Annotate and uniformly standardize time information;
[0055] Entity unification: Eliminate the inconsistencies between entities with different representations;
[0056] Anaphora Resolution: Solve anaphora problems such as pronouns
[0057] 3) Encoding Process: When encoding structured data, one-hot encoding is used to convert each feature in the dataset into a digital form
[0058] 4) Model Construction: Establish a dual-temporal fusion perception model, and define parameters for the temporal knowledge graph, including:
[0059] In the present invention, the temporal knowledge graph is embedded into the quaternion space, and the Hamilton operator is used to score the quadruple to enhance knowledge representation. Specifically, the present invention utilizes the following dual temporal perception channels: (a) Synchronous Temporal Perception. The present invention recombines the synchronous timestamp and the relational quaternion into two composite quaternions. Then we use the Hamilton operator to realize their interaction. (b) Diachronic Temporal Perception. The present invention iterates through all timestamps, maps each time point to an angle in turn, and uses trigonometric functions to convert it into the scalar component of the quaternion to construct the diachronic timestamp. Then the Hamilton operator is used to rotate the relationship between it and the diachronic timestamp. Suppose the graph G consists of N entities, M relationships, and T timestamps. The inventor uses the quaternion matrix to represent the embeddings of all entities, to represent the embeddings of all relationships, to represent the embeddings of all timestamps, where each row is the embedding vector of a specific entity, with a dimension of k. Given a quadruple (s, r, o, τ), the head entity s, the relationship r, and the tail entity o correspond to
[0060] and
[0061] Diachronic Temporal Perception (DP): First, iterate through all timestamps, map each time point to an angle in turn, and use trigonometric functions to convert it into the scalar component of the quaternion to construct the diachronic timestamp:
[0062]
[0063] where: (e′, f′, g′) is the unit vector of the rotation axis;
[0064]
[0065] is the time point.
[0066] Use the Hamilton operator with the diachronic timestamp to rotate the relationship R r quaternion into the diachronic relationship
[0067]
[0068] Normalize to a unit quaternion. Then, rotate the head entity Q s by performing the Hamilton operator between the head entity Q and the normalized s :
[0069]
[0070] Synchronous Temporal Perception (SP): To achieve the interactive learning of temporal information and relational information, the synchronous timestamp W sτ = a sτ + e sτ i + f sτ j + g sτ k is jointly recombined with the relational embedding quaternion R r to obtain two new quaternions:
[0071]
[0072] Perform a rotation operation on using the Hamilton operator through , where is the normalized representation of
[0073]
[0074] Then, rotate the head entity Q s and by performing the Hamilton operator between them to rotate the head entity Q s :
[0075]
[0076] Scoring function: The scoring function of the bi-temporal fusion perception model is expressed as
[0077]
[0078] Regularization: The temporal knowledge graph can be regarded as a fourth-order tensor. Therefore, tensor nuclear norm regularization can be used to enhance TKGs. The inventors propose the following regularization:
[0079]
[0080] where: is the nuclear 4-norm
[0081] In addition, for time constraints, time regularizers are usually used to smooth the representation of adjacent timestamps. Therefore, the following linear time regularizer is used:
[0082]
[0083] where: W b is the linear time constraint bias
[0084] Loss function: The loss function can be expressed as:
[0085]
[0086] where: Π′ samples Π from the set of never-observed quadruples - ;
[0087] λ a is the regularization weight;
[0088] λ b is the time regularization weight.
[0089] 5) Hyperparameter settings: Set hyperparameters for the bi-temporal fusion perception model, including:
[0090] Learning rate: The learning rate controls the magnitude of model parameter updates. Too large may lead to unstable optimization, and too small may lead to slow convergence. An appropriate learning rate can accelerate convergence and improve model performance.
[0091] Batch size: The batch size is the number of samples used in each training. Small batches can accelerate convergence but have unstable updates; large batches can stabilize the gradients but may slow down the training speed and consume more memory.
[0092] Entity embedding vector dimension (ent_vec_dim): The entity embedding vector dimension determines the complexity of entity representation. A larger dimension can better represent entities but may lead to overfitting, while a smaller dimension may not capture enough information
[0093] Relation embedding vector dimension (rel_vec_dim): The relation embedding vector dimension determines the representational ability of relations. A larger dimension can improve the model's expressive power but also increases the computational cost.
[0094] Time embedding vector dimension (tim_vex_dim): The time embedding vector dimension determines the expressive ability of time information, helps the model capture temporal relationships, and is especially suitable for temporal tasks.
[0095] Maximum number of iterations: The maximum number of iterations limits the number of times of the optimization process. A reasonable setting can avoid excessive training time and be combined with an early stopping strategy to improve efficiency.
[0096] 6) Model training: The data in the training set T is gradually input into the model. By calculating the loss value of each data sample and feeding it back to the model, the model adjusts its parameters based on the loss value. The final criterion is that the loss function tends to be stable, indicating that the model has learned the patterns from the data in the training set T and the parameters have been optimized to a relatively good state, thus completing the training task.
[0097] 7) Model application: Specifically, first, the trained model is applied to the test set S. Each quadruple in the test set S is scored through the scoring function in the model. The score of each quadruple reflects the predicted probability of that quadruple in the model or the degree of match with historical data. The model automatically evaluates the rationality of each quadruple by calculating the scores.
[0098] 8) Evaluation and verification: The trained dual-temporal fusion perception model is experimentally evaluated and verified. The performance of the model on the test set is comprehensively evaluated mainly by calculating and comparing the results of evaluation metrics such as MRR, Hit@1, Hit@3, and Hit@10. These evaluation metrics help us examine the prediction accuracy and recall ability of the model from different perspectives, and thus verify its performance.
[0099] The present invention is an intelligent completion method based on dual-temporal fusion perception. The method steps include S1 - S8:
[0100] S1. The present invention is based on two public datasets, ICEWS14 and ICEWS05 - 15, which are divided into a training set T and a test set S in a ratio of 8:2. ICEWS (International Crisis and Event Data System) collects and processes millions of data from multiple sources, aiming to help monitor and respond to global events, such as (Japan, South Korea, March 13, 2014). ICEWS14 and ICEWS05 - 15 cover the event data from January 1, 2014 to December 31, 2014 and from November 1, 2005 to December 31, 2015 respectively;
[0101] S2. Use the Language Technology Platform (LTP) to preprocess ICEWS14 and ICEWS05 - 15, successively performing entity extraction, relation extraction, time information annotation and standardization, entity unification, and anaphora resolution;
[0102] S3. Encode the structured data, and use one-hot encoding to encode the data in the structured dataset;
[0103] S4. Establish a dual-temporal fusion perception model, including a scoring function and a loss function;
[0104] S5. Implement the present invention by running PyTorch. Use the Adagrad optimizer to train the present invention. The batch size is fixed at 1000, and the optimal embedding dimension is k = 100;
[0105] S6. Respectively use the test sets in the ICEWS14 and ICEWS05-15 datasets to substitute into the loss function expression to start training the dual-temporal fusion perception model. The regularization weights λ a and λ b are in {0, …, 0.0025, 0.005, 0.0075, 0.01, …, 0.1}. Find the regularization weights λ a and λ b with the best results for ICEWS14 (0.0075, 0.01) and ICEWS05-15 (0.0025, 0.1);
[0106] S7. Calculate the scores of each quadruple in the ICEWS14 and ICEWS05-15 test sets through the scoring function of the trained dual-temporal fusion perception model. The quadruple with the highest score is selected as the completed data to complete the completion of the test set;
[0107] S8. Input the correlation coefficients of the experimental evaluation metrics MRR, Hit@1, Hit@3, and Hit@10 into the trained dual-temporal fusion perception model for calculation, and finally obtain excellent evaluation results as shown in Table 1.
[0108] Table 1 Summary of Experimental Evaluation Metrics
[0109]
[0110] Table 1 lists the results of the present invention and all baseline models on all datasets, showing the best performance. Generally speaking, the present invention outperforms all baseline models on all datasets. Compared with CEC-BD, the present invention increases the MRR by 27.4 points on ICEWS14 and 19.0 points on ICEWS05-15; compared with TPComplEx, the MRR value of the present invention increases by 0.9 points on ICEWS14 and 2.8 points on ICEWS05-15. ICEWS14, ICEWS18, and ICEWS05-15 all belong to ICEWS, so they share similar data types. Compared with the improvement on ICEWS14, the improvement of the present invention on ICEWS05-15 and ICEWS18 is more significant. Compared with ICEWS14, the time span covered by ICEWS05-15 is more than 10 times longer, which makes its data more dependent on time information. This supports the ability of the present invention in processing long-term time series.
[0111] The embodiments of this application are only preferred solutions and do not limit the scope of patent protection of the present invention. Those of ordinary skill in the art can make various modifications and changes based on the present invention without creative labor, and these modifications should be within the scope of protection of this patent.
Claims
1. An intelligent completion method based on dual-temporal fusion perception, characterized in that: The steps of the method are as follows: S1. Divide the dataset to be completed into a training set and a test set according to a ratio; S2. Conduct structured data preprocessing on the dataset, and successively perform entity extraction, relationship extraction, time information annotation and standardization, entity unification, and coreference resolution on the dataset; S3. Encode the preprocessed structured dataset using one-hot encoding; S4. Establish an intelligent completion model based on dual-temporal fusion perception, including a scoring function, regularization, and loss function; S5. Set hyperparameters for the intelligent completion model, including learning rate, batch size, entity embedding dimension, relationship embedding dimension, time embedding dimension, and maximum number of iterations; S6. Use the training set to train the above intelligent completion model, and substitute the data in the training set into the loss function of the intelligent completion model until the loss function of the model converges; S7. First, apply the trained model to the test set, score each quadruple in the test set through the scoring function in the model, and the score of each quadruple reflects the prediction possibility of the quadruple in the model or the matching degree with historical data. The model automatically evaluates the rationality of each quadruple by calculating the score; S8. Conduct experimental evaluation and verification on the trained intelligent completion model, and comprehensively evaluate the performance of the model on the test set by calculating and comparing the results of evaluation metrics such as MRR, Hit@1, Hit@3, and Hit@10; 2. The intelligent completion method based on dual-temporal fusion perception according to claim 1, wherein: The entity extraction in S2 is to identify meaningful entities from the text, including person names, place names, and dates. The relationship extraction is quadruple extraction, that is, a dataset is represented as a set of head entity, relationship, time, and tail entity. The time information annotation and standardization refer to annotating the time information and converting it into a unified format. The entity unification means unifying entities with different representations into a standard expression. The coreference resolution aims to solve the actual entity pointed to by pronouns or other referring words in the text; 3. The intelligent completion method based on dual-temporal fusion perception according to claim 1, characterized in that: The specific steps of S3 are as follows: For each entity e in the structured dataset i is represented by a one-hot encoded binary vector of f dimensions. Let the i-th element of be equal to 1, and other elements be set to 0; For each relationship r, it is represented by a one-hot encoded binary vector of l dimensions. Let and the j-th element be set to 1, and the remaining elements be set to 0; For each timestamp, first map it to a one-hot encoded binary vector of a fixed length according to the type of the timestamp, where the index position of the timestamp is 1, and other positions are set to 0, thus completing the encoding of time information. After the encoding processes of all entities, relationships, and timestamps are completed, a complete representation of the structured data is obtained.
4. The intelligent completion method based on dual-temporal fusion perception according to claim 1, wherein: The intelligent completion model based on dual-temporal fusion perception in S4 includes a diachronic temporal perception module and a synchronic temporal perception module; The diachronic temporal perception module is responsible for processing the perception of diachronic tense, and the specific process is as follows: Construct a diachronic timestamp, map each time point to an angle in turn, and use trigonometric functions to convert it into the scalar component of a quaternion; Use the Hamilton operator with a historical timestamp W τn to rotate the relation R r by quaternion into a historical relation Normalize to a unit quaternion. Rotate the head entity Q s by performing the Hamilton operator between the head entity Q and the normalized s : The synchronic temporal perception module is responsible for processing the perception of synchronic tense, and the specific process is as follows: Jointly recombine the synchronic timestamp and the relationship embedding quaternion to obtain two new quaternions; By using the Hamilton operator through to perform a rotation operation, where is the normalized representation of By performing a Hamilton operator between the head entity Q s and to rotate the head entity Q s : The scoring function is expressed as: The regularization and loss function use tensor nuclear norm regularization to enhance the temporal knowledge graph, and the following regularization is proposed: is the nuclear 4-norm. For the time constraint, a temporal regularizer is usually used to smooth the representations of adjacent timestamps. The following linear temporal regularizer is used: where: W b is the linear time constraint deviation, randomly initialized and learned during training; The loss function is expressed as: