A Spatio-Temporal Knowledge Graph Embedding Method Based on Variable Translation

Through the spatial and temporal knowledge graph embedding method based on variable translation, traditional translation principles are relaxed, auxiliary parameters and temporal attributes are introduced, link prediction and tripartite classification effects in the knowledge graph completion task are optimized, and the actual shortcomings of traditional models in modeling complex and diverse entities and dynamics are solved.

CN117216295BActive Publication Date: 2025-07-08GUANGXI NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311226561.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-22
Publication Date
2025-07-08
Estimated Expiration
2043-09-22

AI Technical Summary

Technical Problem

Traditional translation-based knowledge representation models cannot effectively model complex and diverse entities and relationships in the knowledge graph completion task, and cannot handle dynamic knowledge graph facts.

Method used

The spatial and temporal knowledge graph embedding method based on variable translation is adopted, and the model performance is optimized by relaxing traditional translation principles, auxiliary parameters and random numbers are introduced to adjust the embedding range, and the temporal properties of entities and relationships are integrated, and the scoring function and loss function are designed to optimize the model performance.

Benefits of technology

The effects of link prediction and triple classification are optimized, the model's ability to model complex and diverse entities and relationships is improved, and the completeness and accuracy of the knowledge graph are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117216295B_ABST
    Figure CN117216295B_ABST
Patent Text Reader

Abstract

The present invention discloses a spatio-temporal knowledge graph embedding method based on variable translation, comprising the following steps: 1 constructing a TKGE-VT model; 2 designing a loss function; 3 constructing negative triples; 4 training the model. This method relaxes the strict constraints of traditional translation principles, realizes flexible conversion between entity and relationship embeddings; also incorporates the temporal attributes of entities and relationships in triples, and applies the proposed variable translation principle to spatio-temporal knowledge graphs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of knowledge graph embedding and knowledge graph completion, and specifically relates to a spatio-temporal knowledge graph embedding method based on variable translation. Background Technique

[0002] A knowledge graph is a large-scale semantic network graph, where nodes represent entities and edges represent the relationships between entities. Each edge in the knowledge graph corresponds to a fact in the real world and can be represented by a triple (h, r, t), where h and t represent the head entity and the tail entity respectively, and r represents the relationship between them. Nowadays, many knowledge graphs have been constructed, such as Freebase, WordNet, YAGO, and NELL, etc. These large-scale knowledge graphs have been widely applied in various different fields, such as knowledge reasoning, intelligent question answering, information extraction, and recommendation systems, etc.

[0003] However, due to the fact that knowledge graphs are constructed manually or semi-manually, the phenomenon of knowledge graph incompleteness is very common. For example, there are approximately 3 million human entities in Freebase, but 71% of the entities have no birthplace information, and 91% of the entities have no education information. Therefore, predicting the missing entities or relationships in triples has always been a crucial issue, which is called knowledge graph completion. Knowledge graph completion aims to solve the problem of data sparsity in knowledge graphs and improve the integrity of knowledge graphs.

[0004] Due to the efficient computing power and low computational complexity of knowledge representation learning, it has been widely applied to knowledge graph completion tasks. Knowledge representation learning is the representation learning for entities and relationships in a knowledge base. By projecting entities or relationships into a low-dimensional continuous vector space, it can realize the representation of the semantic information of entities and relationships, and can efficiently calculate the complex semantic associations between entities, relationships, and themselves.

[0005] The translation-based knowledge representation learning methods perform well in knowledge graph completion tasks, so a series of translation-based methods have been proposed, such as TransE, TransH, TransR, and TransD, etc. However, these traditional translation-based methods all adopt the same translation principle, that is, h + r ≈ t. This translation principle is too strict and cannot well model complex and diverse entities and relationships. In addition, most of these traditional translation principles are used in static knowledge graphs. However, many facts in knowledge graphs are not static, and they usually only hold at specific time periods or timestamps. For example, the triple (Bill_Clinton, presidentof, US) is only true from 1993 to 2001, and the triple (Steve_Jobs, diedin, California) is only true on October 5, 2011. Summary of the Invention

[0006] The object of the present invention is to propose a spatio-temporal knowledge graph embedding method based on variable translation for the problem that the translation principle adopted by traditional translation-based knowledge representation models is too strict to model complex and diverse entities and relationships. This method relaxes the strict constraints of traditional translation principles and realizes flexible conversion between entity and relationship embeddings; it also incorporates the time attributes of entities and relationships in triples and applies the proposed variable translation principle to spatio-temporal knowledge graphs.

[0007] The technical solution for achieving the object of the present invention is as follows:

[0008] A spatio-temporal knowledge graph embedding method based on variable translation, comprising the following steps:

[0009] 1. Construct the TKGE-VT model:

[0010] 1.1 Construct the variable translation principle: The key to the variable translation principle is to provide greater freedom for the embeddings of entities and relationships to relax the strict constraints of the traditional translation principle h + r ≈ t, so that the TKGE-VT model can better model entities and relationships: for a given head entity h and relationship r, introduce an auxiliary parameter β t Set the embedding range of the tail entity t as a plane, and t + β t is not a fixed vector, but a set of vectors with different sizes in the same direction. Similarly, for a given head entity h and tail entity t, introduce an auxiliary parameter β r Set the embedding range of the relationship r as a plane, and r + β r is not a fixed vector, but a set of vectors with the same direction and different sizes. For a given tail entity t and relationship r, introduce an auxiliary parameter β h Set the embedding range of the head entity h as a plane, and h + β h is not a fixed vector, but a set of vectors with the same direction and different sizes. The constructed variable translation principle is: where β h β r and β t are auxiliary parameters of the head entity vector, relationship vector and tail entity vector, respectively setting the embedding ranges of the head entity, relationship and tail entity as a plane, λ and α are hyperparameters used to adjust the sizes of h + β h r + β r and t + β t respectively;

[0011] 1.2 Introduce random numbers into the variable translation principle: Considering the random error problem, that is, if h + r is not equal to t + βt +α, t - r is not equal to t - h is also not equal to r + β r +λ, introduce random numbers u1, u2, and u3 to adjust the random error;

[0012] 1.3 Incorporate the temporal attributes of entities and relationships: Many facts in the knowledge graph are dynamic, and these dynamic facts usually only apply to a specific time period or timestamp, that is, the triple facts in the knowledge graph have temporal attributes. Therefore, in the TKGE-VT model, incorporate the temporal attributes of triples, separately learn the temporal-aware embeddings of entities and relationships in the triples, and represent the temporal-aware embedding of the head entity h in the triple as h T , represent the temporal-aware embedding of the tail entity t as t T , represent the temporal-aware embedding of the relationship r as r T , therefore, the head entity h with temporal-aware embedding is represented as h + h T , the tail entity t with temporal-aware embedding is represented as t + t T , the relationship r with temporal-aware embedding is represented as r + r T ;

[0013] 1.4 Design the scoring function of the TKGE-VT model: Design a scoring function to score positive triples and negative triples. Through the scoring function, the score of positive triples will be less than the score of negative triples to test the performance of the TKGE-VT model. Design the scoring function of the spatio-temporal knowledge graph embedding model based on variable translation - TKGE-VT as: where both L1 and L2 represent regularization;

[0014] 2 Design the loss function: When training the TKGE-VT model, a loss function needs to be designed to describe the gap between the predicted value and the true value of the triple. Use L to represent the loss function. Through the iteration of the loss function, optimize the prediction and classification effects of the model, and at the same time update the parameters of the model. Use the following margin-based ranking loss function:

[0015]

[0016] where, f r (h, t) is the score of the positive triple, f r (h', t') is the score of the negative triple, γ is a hyperparameter representing the margin between the positive triple and the negative triple. When training the TKGE-VT model, maximize the difference between the positive triple score and the negative triple score by continuously adjusting the value of the margin γ, thereby optimizing the prediction and classification effects of the model;

[0017] 3 Construct negative triples:

[0018] 3.1 Replace the head entity or tail entity using the probability method: When training the TKGE-VT model, discriminative training needs to be carried out based on positive triples and negative triples, and the scores of positive triples and negative triples are sorted. To improve the quality of the generated negative triples, the probability method is used to construct negative samples, and the head entity h or the tail entity t is replaced with different probabilities, that is, different replacement strategies are set according to the relationship type. The relationship types are mainly one-to-many and many-to-one, and the replacement strategies are as follows: For the one-to-many relationship, that is, one head entity corresponds to multiple tail entities, the head entity is replaced with a higher probability; for the many-to-one relationship, that is, multiple head entities correspond to one tail entity, the tail entity is replaced with a higher probability. An entity has multiple attributes. When dealing with the one-to-many relationship, replacing the head entity can enable each attribute of the head entity to be fully trained. When dealing with the many-to-one relationship, replacing the tail entity can also enable each attribute of the tail entity to be fully trained;

[0019] 3.2 Select entities with similar semantics for replacement: In the vector space, entities of the same type are distributed in adjacent regions. For example, for the relationship "precident", the head entity is a person's name and the tail entity is a country name. Person's names will be concentrated in one region in the vector space, and country names will be concentrated in another region. However, in the vector space, there may be a situation where person's names and country names are distributed relatively close. If person's names and country names are distributed in adjacent regions, it is difficult to distinguish the entities gathered in one region. Therefore, when replacing the head entity and tail entity of a positive triple to generate a negative triple, entities with similar semantics are used for replacement to improve the model's ability to distinguish entities;

[0020] 4 Model training:

[0021] 4.1 Link prediction for spatio-temporal knowledge graph embedding based on variable translation: For link prediction, given a triple (?, r, t) or (h, r,?), the link prediction task is to predict the missing h or t based on the known facts in the triple. For the head entity, first randomly replace the head entity h of each triple with other entities, and then use the scoring function designed in step 1.4 to calculate the score of each triple. Sort the triples in ascending order according to the scores, record the ranking of the original correct triple, and then perform the same processing on the tail entity t as on the head entity h. Finally, select the entity in the triple with the smallest score as the head entity or tail entity of the missing triple;

[0022] 4.2 Triple Classification of Spatio-Temporal Knowledge Graph Embedding Based on Variable Translation: For triple classification, given a triple (h, r, t), the triple classification task determines whether the triple is correct, that is, whether it belongs to the factual triples existing in the knowledge graph, through the knowledge representation learned by the model. This is a binary classification task and also one of the classic tasks of knowledge graph completion: First, a relationship-specific threshold σ needs to be set. r For a given triple (h, r, t), if the score calculated by the scoring function is less than the set threshold σ r , then the triple will be determined as a positive triple; otherwise, it is a negative triple. Compared with existing knowledge graph embedding methods, the beneficial effects of this technical solution are:

[0023] 1. A translation principle of variable translation is proposed, relaxing the strict constraints of traditional translation principles, enabling the model to better model complex and diverse entities.

[0024] 2. During the process of embedding triple facts, the time attributes of entities and relationships are incorporated respectively, and the proposed variable translation principle is applied to the spatio-temporal knowledge graph.

[0025] 3. A new scoring function is designed, enabling the model to better represent complex and diverse entities and relationships in the knowledge graph.

[0026] Combining the above three points, this technical solution finally optimizes the effects of link prediction and triple classification, outperforming traditional baseline models. Description of the Drawings

[0027] Figure 1 It is a schematic diagram of the variable translation principle in the embodiment;

[0028] Figure 2 It is a schematic diagram of the variable translation principle with random numbers in the embodiment;

[0029] Figure 3 It is the overall architecture diagram of the TKGE-VT model in the embodiment;

[0030] Figure 4 It is an example diagram of easily confused triples in the embodiment. Detailed Implementation Modes

[0031] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, but it is not a limitation of the present invention.

[0032] Embodiment:

[0033] A spatio-temporal knowledge graph embedding method based on variable translation, comprising the following steps:

[0034] 1 Construct the TKGE-VT model:

[0035] 1.1 Construct the variable translation principle: As Figure 1 shown, in this example, the variable translation principle is designed as: For a given head entity h and relation r, introduce an auxiliary parameter β t Set the embedding range of the tail entity t as a plane, and the tail entity embedding vector t + β t is on the same straight line and its size is flexible. α is used as a hyperparameter to adjust t + β t 's size (see Figure 1 (a)). Similarly, for a given tail entity t and relation r, introduce an auxiliary parameter β h Set the embedding range of the head entity h as a plane, and the head entity embedding vector h + β h is on the same straight line and its size is flexible, and is used as a hyperparameter to adjust h + β h 's size (see Figure 1 (b)). If the given head entity h and tail entity t are provided, introduce an auxiliary parameter β r Set the embedding range of the relation r as a plane, and the relation embedding vector r + β r is on the same straight line and its size is flexible. λ is used as a hyperparameter to adjust r + β r 's size (see Figure 1 (c)). The variable translation principle relaxes the strict constraints of the traditional translation principle, realizes the flexible conversion between entities and relations, improves the model's ability to handle complex relations (i.e., one-to-many, many-to-one, and many-to-many relations), and enables the model to better model the triple facts in the knowledge graph;

[0036] 1.2 Introduce random numbers in the variable translation principle: Considering the random error problem, that is, if h + r is not equal to t + β t + α, t - r is not equal to t - h is also not equal to r + β r + λ, introduce random numbers u1, u2, and u3 to adjust the random error: In this example, the random number u1 is used to adjust the error of t + β t + α, the random number u2 is used to adjust 's error, and the random number u3 is used to adjust the error of r + β r + λ. Figure 2 This is the variable translation principle with random numbers. Therefore, the overall variable translation principle is designed as:

[0037] 1.3 Incorporating the Temporal Attributes of Entities and Relationships: Many facts in the knowledge graph are dynamic, and these dynamic facts usually only apply to a specific time period or timestamp, that is, the triple facts in the knowledge graph have temporal attributes. Therefore, in the TKGE-VT model, the temporal attributes of triples are incorporated, the temporal-aware embeddings of entities and relationships in the triples are learned respectively, and the temporal-aware embedding of the head entity h in the triple is represented as h T , the temporal-aware embedding of the tail entity t is represented as t T , and the temporal-aware embedding of the relationship r is represented as r T . Therefore, the head entity h with temporal-aware embedding is represented as h + h T , the tail entity t with temporal-aware embedding is represented as t + t T , and the relationship r with temporal-aware embedding is represented as r + r T ; The overall architecture of the TKGE-VT model is as shown in Figure 3 . Therefore, for the embeddings of the given head entity h and relationship r, the tail entity can be represented as t + t T +β t +α+u1 (see Figure 3 (a)); For the embeddings of the given tail entity t and relationship r, the head entity can be represented as (see Figure 3 (b)); For the embeddings of the given head entity h and tail entity t, the relationship can be represented as r + r T +β r +λ+u3 (see Figure 3 (c));

[0038] 1.4 Designing the Scoring Function of the TKGE-VT Model: To measure the credibility of triples and complete the representation of semantic information in the knowledge graph, a scoring function is designed to score positive and negative triples. Through the scoring function, the score of the positive triple will be less than that of the negative triple to test the performance of the TKGE-VT model. In this example, f r (h,t) is used to represent the scoring function. When the positive triple is input, a numerical value can be calculated through the scoring function, represented by f r (h,r,t). At the same time, f r (h',r,t') is used to represent the result of inputting the negative triple. The score f r (h,r,t) of the positive triple is lower than the score f r (h',r,t') of the negative triple. Combining the variable translation principle proposed in step 1.1 and the temporal attributes of entities and relationships in step 1.3, the scoring function of the spatio-temporal knowledge graph embedding model based on variable translation - TKGE-VT is designed as: where Both L1 and L2 represent regularization;

[0039] 2 Design the loss function: When training the TKGE-VT model, a loss function needs to be designed to describe the gap between the predicted value and the true value of the triple. Let L represent the loss function. Through the iteration of the loss function, the prediction and classification effects of the model are optimized, and at the same time, the parameters of the model are updated. The following margin-based ranking loss function is used as the training objective:

[0040]

[0041] where S is the set of positive triples, S' is the set of negative triples, max(x, y) represents returning the maximum value of x and y, and f r (h, t) is the score of the positive triple, and f r (h', t') is the score of the negative triple. γ is a hyperparameter representing the margin between the positive triple and the negative triple. When training the TKGE-VT model, by continuously adjusting the value of the margin γ, the difference between the positive triple score and the negative triple score is maximized, thereby optimizing the prediction and classification effects of the model;

[0042] The goal of the loss function is to separate the positive triples and negative triples as much as possible. During the model training process, stochastic gradient descent (SGD) is used to minimize the objective loss function. The set of positive triples (from the triples in the knowledge graph) is traversed multiple times. When a positive triple is visited, a negative triple is constructed according to step 3. After a small batch, the gradient is calculated and the model parameters are updated;

[0043] 3 Construct negative triples: To improve the training speed and the model prediction accuracy, negative triples need to be constructed based on the positive triples; the traditional method obtains negative triples by randomly shuffling the positive triples. For example, in TransE, for a positive triple (h, r, t), a corresponding negative triple is randomly selected from the entity set E as a pair of entities (h', t'). This random sampling method is only effective when dealing with one-to-one relationships. When dealing with one-to-many, many-to-one, and many-to-many relationships, the original correct triples may be marked as incorrect triples. Therefore, in this example, to improve the quality of the generated negative triples, the head entity or the tail entity is replaced with different probabilities, and entities with similar semantics are selected for replacement to improve the model's ability to distinguish entities;

[0044] 3.1 Replace the head entity or tail entity using the probability method: When training the TKGE-VT model, discriminative training needs to be carried out based on positive triples and negative triples, and the scores of positive triples and negative triples are sorted. To improve the quality of the generated negative triples, the probability method is used to construct negative samples, and the head entity h or the tail entity t is replaced with different probabilities, that is, different replacement strategies are set according to the relationship type. The relationship types are mainly one-to-many and many-to-one, and the replacement strategies are as follows: for the one-to-many relationship, that is, one head entity corresponds to multiple tail entities, the head entity is replaced with a high probability; for the many-to-one relationship, that is, multiple head entities correspond to one tail entity, the tail entity is replaced with a higher probability. An entity has multiple attributes. When dealing with the one-to-many relationship, replacing the head entity can enable each attribute of the head entity to be fully trained. When dealing with the many-to-one relationship, replacing the tail entity can enable each attribute of the tail entity to be fully trained;

[0045] In all triples of a relationship, in this example, the following data are first calculated: one is the average number of tail entities tph corresponding to each head entity, and the other is the average number of head entities hpt corresponding to each tail entity. Then, a Bernoulli distribution with the parameter p = tph / (tph + hpt) is defined for sampling. For a given positive triple, when constructing a negative triple, the head entity is replaced with the probability p, and the tail entity is replaced with the probability 1 - p. In this way, the total probability can be 1, and the sampling conforms to the Bernoulli distribution. Using the Bernoulli distribution can not only greatly reduce the probability of generating incorrect negative triples but also reduce the computational complexity of the model;

[0046] In this example, it is stipulated that if tph < 1.5 and hpt < 1.5, the relationship r is considered one-to-one; if tph > 1.5 and hpt > 1.5, the relationship r is considered many-to-many; if tph ≥ 1.5 and hpt < 1.5, the relationship r is considered one-to-many; if tph < 1.5 and hpt ≥ 1.5, the relationship r is considered many-to-one;

[0047] 3.2 Select semantically similar entities for replacement: In the vector space, entities of the same type are distributed in adjacent regions. For example, for the relationship precident, the head entity is a person's name, and the tail entity is a country name. Person's names will be concentrated in one region in the vector space, and country names will be concentrated in another region. However, in the vector space, there may be a situation where person's names and country names are distributed relatively close. If person's names and country names are distributed in adjacent regions, as Figure 4 shown, it is difficult to distinguish the entities gathered in one region. Therefore, when replacing the head entity and tail entity of a positive triple to generate a negative triple, semantically similar entities are used for replacement to improve the model's discrimination ability for entities;

[0048] When judging the similarity between entities in this example, it is necessary to first understand the semantic information of the entities. Generally speaking, the higher the entity similarity, the closer the semantics of the entities; in a knowledge graph, the attributes of entities can be used as the main basis for similarity judgment; in a distributed representation model, the semantic similarity between entities or relationships is usually represented by the similarity between vectors. Therefore, in this example, the similarity of entities is judged according to the semantic similarity of entities, and the similarity of entities is defined as:

[0049]

[0050] Therefore, given a positive triple (h, r, t), when generating a negative triple (h', r, t) by replacing the head entity, select h' such that dis(h, h') is minimized; when generating a negative triple by replacing the tail entity, select t' such that dis(t, t') is minimized;

[0051] 4 Model Training: For model training, in this example, the performance of the TKGE-VT model is evaluated on the knowledge graph completion task, and the model is mainly evaluated and trained on two subtasks of knowledge graph completion, namely link prediction and triple classification:

[0052] 4.1 Link Prediction of Spatiotemporal Knowledge Graph Embedding Based on Variable Translation: For link prediction, given a triple (?, r, t) or (h, r,?), the link prediction task is to predict the missing h or t based on the known facts in the triple. In this link prediction task, for each position of the missing entity, the TKGE-VT model is required to rank a set of candidate entities from the knowledge graph, rather than just giving a best result: First, randomly replace the head entity h of each triple with other entities, and then use the scoring function designed in step 1.4 to calculate the score of each triple. Sort the triples in ascending order of scores, record the ranking of the original correct triple, and then perform the same processing on the tail entity t as on the head entity h. Finally, select the entity in the triple with the smallest score as the head entity or tail entity of the missing triple. The specific operations are as follows:

[0053] 4.1.1 Data preprocessing: Analyze and preprocess the original data. In this task, two datasets are used, namely the WN18 dataset and the FB15K dataset. Among them, the WN18 dataset is a subset of WordNet, which is a database characterized by lexical relationships between words. FB15K is a subset of Freebase, which is a large-scale knowledge graph containing general knowledge facts. The WN18 dataset contains 18 relations and 40,943 entities, and the FB15K dataset contains 1,345 relations and 14,951 entities. Among them, 26.2% are one-to-one relations, 22.7% are one-to-many relations, 28.3% are many-to-one relations, and 22.8% are many-to-many relations. Therefore, FB15K is considered a large dense dataset. After analyzing the original datasets, the WN18 dataset and the FB15K dataset are respectively divided into training sets, validation sets, and test sets. The detailed division of the link prediction datasets is shown in Table 1:

[0054] Table 1 Division of link prediction datasets

[0055]

[0056] 4.1.2 Parameter initialization settings: During the experiment, to avoid the influence of random errors on the experimental results, each experiment is conducted 10 times, and then the average value is taken as the final result. The optimal model parameters are determined through the validation set. The traditional negative sampling method of randomly replacing the head entity / tail entity is represented by "unif", and the method adopting the Bernoulli sampling strategy is represented by "bern", that is, different probabilities are used to replace the head entity / tail entity according to different relation types. Under the "unif" setting, the optimal parameter configuration for the WN18 dataset is: random number μ = 0.1, learning rate α = 0.0001, margin γ = 5, embedding dimension k = 100, batch size B = 4800; the optimal parameter configuration for the FB15K dataset is: random number μ = 0.1, learning rate α = 0.0001, margin γ = 4, embedding dimension k = 100, batch size B = 4800; under the "bern" setting, the optimal parameter configuration for the WN18 dataset is: random number μ = 0.1, learning rate α = 0.0001, margin γ = 4, embedding dimension k = 200, batch size B = 4800; the optimal parameter configuration for the FB15K dataset is: random number μ = 0.1, learning rate α = 0.001, margin γ = 4, embedding dimension k = 100, batch size B = 4800. For these two datasets, all training triples are trained for 500 iterations;

[0057] 4.1.3 Model iterative training: First, the complete knowledge graph triples (h, r, t) in the training data are divided into batches and used as the input of the model. For the input triples, the head entity and tail entity embedding matrices and relationship embedding matrices (embeddings) are trained respectively. Then, the head and tail entity and relationship embeddings are respectively input into the prediction model TKGE-VT (i.e., the constructed head entity prediction (?, r, t) and tail entity prediction (h, r, ?)). Finally, the losses of these two parts are combined as the total loss for training;

[0058] 4.1.4 Testing and Evaluation: For each test triple (h, r, t), replace the tail entity t with each entity e in the knowledge graph and evaluate it according to the scoring function f r (h, t) calculates the corrupted triple (h, r, e), then sorts it in descending order, and finally gets the ranking of the original correct triple. Similarly, another ranking of (h, r, t) is obtained by corrupting the head entity h. The average MeanRank and Hit@10 of these predicted rankings are reported, that is, the proportion of correct entities ranked in the top 10. This evaluation setting is named "Raw";

[0059] In fact, there may also be a damaged triple in the knowledge graph, which should also be considered correct. However, the above evaluation may ignore the situation where these damaged but correct triplets are ranked high. Therefore, before sorting, it is necessary to remove the damaged triplets contained in the training set, validation set, and test set. This evaluation setting is called "Filter". In this example, the evaluation results of two settings will be reported. In these two settings, the lower the MeanRank, the better, and the higher the Hit@10, the better.

[0060] 4.2 Triple classification based on spatiotemporal knowledge graph embedding with variable translation: For triple classification, given a triple (h, r, t), the triple classification task uses the knowledge representation learned by the model to determine whether the triple is correct, that is, whether it belongs to the fact triple in the knowledge graph. This is a binary classification task and also one of the classic tasks of knowledge graph completion: First, a relation-specific threshold σ needs to be set. r , which is obtained when the accuracy of triple classification on the validation set is maximized. For a given triple (h, r, t), if the score calculated by the scoring function is less than the set threshold σ r , then the triple will be judged as a positive triple, otherwise, it is a negative triple; the specific operations are as follows:

[0061] 4.2.1 Data Preprocessing: In the data preprocessing stage, the original dataset is first analyzed and preprocessed. In triple classification, 3 datasets are used, namely a subset of WordNet, i.e., the WN11 dataset, and two subsets of Freebase, i.e., the FB13 dataset and the FB15K dataset. Among them, the test sets of the WN11 dataset and the FB13 dataset contain positive and negative triples. However, for the FB15K dataset, its test set only contains correct triples, and negative triples need to be constructed. In this example, the probability method is used to construct negative triples for the FB15K dataset. Then, the three datasets of WN11, FB13, and FB15K are respectively divided into training sets, validation sets, and test sets. The detailed information on the division of the triple classification datasets is shown in Table 2:

[0062] Table 2 Division of Triple Classification Datasets

[0063]

[0064] 4.2.2 Setting a Relationship-Specific Threshold: During the triple classification experiment, a relationship-specific threshold δ needs to be set. According to this threshold, the model can distinguish the positive and negative of the input triples. For the triple (h, r, t), if the score obtained through the scoring function f r is lower than the threshold δ r , the triple will be classified as a positive triple, otherwise as a negative triple. In this example, the threshold δ is optimized by maximizing the classification accuracy on the validation set r ; r ;

[0065] 4.2.3 Parameter Initialization Settings: The model hyperparameters are set according to the accuracy on the validation set in Step 4.2.1, that is, when the accuracy on the validation set reaches the optimum, the hyperparameters of the model are obtained: the random number μ, the learning rate α, the margin γ, the embedding dimension k, and the batch size B. The data in the training set and test set divided in Step 4.2.1 are processed in batches with B triple samples as a batch. The optimal parameter configuration on the WN11 dataset is: learning rate α = 0.001, margin γ = 10, embedding dimension k = 100, random number μ = 0.1, batch size B = 4800; the optimal parameter configuration on the FB13 dataset is: learning rate α = 0.001, margin γ = 5, embedding dimension k = 200, random number μ = 0.1, batch size B = 4800; the optimal parameter configuration on the FB15K dataset is: learning rate α = 0.001, margin γ = 5, embedding dimension k = 100, random number μ = 0.1, batch size B = 120;

[0066] 4.2.4 Model Iterative Training: Iteratively train TKGE-VT and update the coefficients after each iteration. In the triple classification experiment, 500 iterations were performed on all the training datasets to obtain the model TKGE-VT that meets the expectations. At the same time, the samples of the validation set divided in the data preprocessing stage were input into TKGE-VT in batches to calculate the accuracy of triple classification and the loss value of the loss function. After each epoch, check whether the model has reached the best validation loss. If so, save the parameters of the model at this time, and then use these optimal parameters to calculate the performance of the test set;

[0067] 4.2.5 Testing and Evaluation: Input the samples of the test set in the triple classification task training step 4.2.3 into the model TKGE-VT obtained in step 4.2.4 for testing, and record the test results at the same time. Then save and record the trained model parameters, and finally output the optimal model parameters;

[0068] In the triple classification task, the accuracy ACC is used as the evaluation metric. The higher the ACC, the better the model's performance in the triple classification task. The definition of ACC is as follows:

[0069]

[0070] where T p is the number of correctly predicted positive triples, T N is the number of correctly predicted negative triples, N pos and N neg represent the number of positive and negative triples in the training set, respectively.

Claims

1. A spatio-temporal knowledge graph embedding method based on variable translation, characterized in that , including the following steps:

1. Construct the TKGE-VT model: 1.1 Constructive Variable Translation Principle: The key to the variable translation principle lies in providing greater freedom for the embedding of entities and relationships to relax the strict constraints of traditional translation principles so that the TKGE-VT model can better model entities and relationships: For a given head entity and relationship , introduce auxiliary parameters to set the embedding range of the tail entity as a plane, and not a fixed vector, but a group of vectors with different magnitudes in the same direction. Similarly, for a given head entity and tail entity , introduce auxiliary parameters to set the embedding range of the relationship as a plane, and not a fixed vector, but a group of vectors with the same direction and different magnitudes. For a given tail entity and relationship , introduce auxiliary parameters to set the embedding range of the head entity as a plane, and not a fixed vector, but a group of vectors with the same direction and different magnitudes. The constructed variable translation principle is: , where , and are the auxiliary parameters of the head entity vector, relationship vector, and tail entity vector respectively. The embedding ranges of the head entity, relationship, and tail entity are set as a plane respectively, , and are used as hyperparameters to adjust the magnitudes of , and respectively; 1. Introduce random numbers in the variable translation principle: Considering the random error problem, that is, if is not equal to , is not equal to , is also not equal to , introduce random numbers , and to adjust the random error; 1.3 Incorporating Temporal Attributes of Entities and Relations: Many facts in knowledge graphs are dynamic, and these dynamic facts usually only apply to a specific time period or timestamp, that is, the triple facts in the knowledge graph have temporal attributes. Therefore, in the TKGE-VT model, the temporal attributes of triples are incorporated, the temporal-aware embeddings of entities and relations in the triples are learned respectively, and the learned temporal-aware embedding of the head entity is represented as , the temporal-aware embedding of the tail entity is represented as , and the temporal-aware embedding of the relation is represented as . Therefore, the head entity with temporal-aware embedding is represented as , the tail entity with temporal-aware embedding is represented as , and the relation with temporal-aware embedding is represented as ; 1.4 Design the scoring function of the TKGE-VT model: Design a scoring function to score positive triples and negative triples. Through the scoring function, the score of positive triples will be less than that of negative triples to test the performance of the TKGE-VT model. The scoring function of the spatio-temporal knowledge graph embedding model based on variable translation - TKGE-VT is designed as: , where , and both represent regularization; 2 Design the loss function: When training the TKGE-VT model, it is necessary to design a loss function to describe the gap between the predicted value and the true value of the triple. Use to represent the loss function. Through the iteration of the loss function, the prediction and classification effects of the model are optimized, and at the same time, the parameters of the model are updated. The following margin-based ranking loss function is used: wherein, is the positive triple score, is the negative triple score, is a hyperparameter representing the margin between the positive and negative triples. During training, by continuously adjusting the margin value, the difference between the positive triple score and the negative triple score is maximized, thereby optimizing the prediction and classification effects of the model; 3. Construct negative triples: 3.1 Replace the head entity or tail entity using the probability method: When training the TKGE-VT model, discriminative training needs to be performed based on positive and negative triples, and the scores of positive and negative triples are sorted. To improve the quality of the generated negative triples, the probability method is used to construct negative samples, and the head entity is replaced with different probabilities or the tail entity , that is, different replacement strategies are set according to the type of relationship. The relationship types are mainly one-to-many and many-to-one. The replacement strategy is as follows: for a one-to-many relationship, that is, one head entity corresponds to multiple tail entities, the head entity is replaced with a higher probability; for a many-to-one relationship, that is, multiple head entities correspond to one tail entity, the tail entity is replaced with a higher probability. An entity has multiple attributes. When dealing with a one-to-many relationship, replacing the head entity can fully train each attribute of the head entity. When dealing with a many-to-one relationship, replacing the tail entity can fully train each attribute of the tail entity; 3.2 Select semantically similar entities for replacement: In the vector space, entities of the same type are distributed in adjacent regions. Therefore, when generating negative triples by replacing the head entity and the tail entity of a positive triple, use semantically similar entities for replacement to improve the model's discrimination of entities; 4. Model training: 4.1 Link Prediction for Spatiotemporal Knowledge Graph Embedding Based on Variable Translation: For link prediction, given a triple or , the link prediction task is to predict the missing or according to the known facts in the triple: First, randomly replace the head entity of each triple with other entities, then use the scoring function designed in step 1.4 to calculate the score of each triple, sort the triples in ascending order according to the scores, record the ranking of the original correct triple, and then perform the same processing on the tail entity as on the head entity . Finally, select the entity in the triple with the smallest score as the head entity or tail entity of the missing triple; 4.2 Triple Classification Based on Spatiotemporal Knowledge Graph Embedding with Variable Translation: For triple classification, given a triple , the triple classification task determines whether the triple is correct, i.e., whether it belongs to the factual triples existing in the knowledge graph, through the knowledge representation learned by the model. This is a binary classification task and also one of the classic tasks in knowledge graph completion: First, a relation-specific threshold needs to be set. For a given triple , if the score calculated by the scoring function is less than the set threshold , then the triple will be judged as a positive triple; otherwise, it is a negative triple.

Citation Information

Patent Citations

  • Knowledge graph representation learning method based on dynamic translation

    CN113312492A

  • End-to-port language understanding without complete transcript

    CN116686045A