Knowledge Representation Learning Method Integrating Ordered Relationship Paths and Entity Description Information
By integrating ordered relationship paths and entity description information in the knowledge graph, using Transformer and relational attention mechanism for encoding and combining TransR model, the problem of inefficient computing in the existing technology is solved, and more efficient knowledge graph inference and complex relationship modeling is achieved.
Patent Information
- Application Number
- CN202210752210.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-29
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-06-29
AI Technical Summary
The existing knowledge graph representation learning methods fail to make full use of ordered relationship paths and entity description information, resulting in inefficient computing on large-scale knowledge graphs and difficulty in effectively conducting inference analysis.
A knowledge representation learning method that integrates ordered relational paths and entity description information is adopted. By projecting entities into different spaces and using pooling strategies to fusion path information, combining Transformer and relational attention mechanisms for encoding, and finally training and learning is carried out on the basis of TransR model, comprehensively utilizing triples, paths and description information.
It improves the ability to learn knowledge graph representation, enhances the accuracy of reasoning and modeling of complex relationships, and improves the performance of link prediction tasks.
Smart Images

Figure CN115858799B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of knowledge representation learning, and particularly to a knowledge representation learning method that fuses ordered relation paths and entity description information. Background Art
[0002] The statements in this section merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.
[0003] Knowledge graphs provide a better means for organizing, managing, and understanding massive amounts of Internet data information. As an important branch in the field of artificial intelligence, knowledge graphs are widely used in search engines, intelligent healthcare, question answering systems, etc., and have received extensive attention from both the academic and industrial communities.
[0004] A knowledge graph is composed of a large number of triples, which can be abbreviated as (h, r, t), indicating that two entities h and t are connected by a relation r. Based on the symbol-based triples of the knowledge graph, although concise, as the scale of the knowledge graph continues to increase, problems such as data sparsity become more prominent, resulting in low computational efficiency and making it difficult to achieve efficient reasoning on large-scale knowledge graphs. Based on this, as a supplement to symbol representation, knowledge graph representation learning is introduced, whose purpose is to project the entities and relations in the knowledge graph into a continuous low-dimensional vector space to improve computational efficiency and promote reasoning analysis on large-scale knowledge graphs.
[0005] Knowledge representation learning realizes the accurate description of semantic information of entities and relationships by vectorizing entities and relationships. In recent years, a variety of different types of knowledge representation learning models have been proposed. First, there is the translation model TransE proposed by Bordes et al. It regards the process of the head entity in a triple connecting to the tail entity through a relationship as a translation process, then uses a scoring function to measure the rationality of each triple, and finally obtains the most accurate vector representation result by continuously optimizing the loss function. Although TransE is simple and efficient, it is prone to semantic conflicts between different entities when dealing with complex relationships. To overcome this defect, Lin et al. proposed TransR, which maps entities and relationships to different spaces respectively, and then projects the entity vector representation from the entity space to the relationship space to realize the translation process, solving the problem that when entities and relationships belong to different objects, they cannot be represented in the same space. Although these methods have achieved good results, they only consider the triple structure information and do not make good use of multi-source information such as entity description information and relationship path information to further improve the semantic representation ability of the model. In terms of considering the semantic relationship of paths between entities, OPTransE projects the head and tail entities to different spaces to ensure the order in the path, and adopts a pooling strategy to extract complex features in different paths. DKRL integrates the entity description information in the knowledge graph into the knowledge graph representation learning, encodes and represents the entity description information using CNN and CBOW respectively, and at the same time uses fact triples and entity description information for learning. MCapsEED uses Transformer and a relationship attention mechanism to obtain the representation of entity descriptions, combines it with the structure information of triples using a dynamic gating mechanism, and finally inputs it into a capsule network model for processing.
[0006] There are often a large number of ordered relationship path information between entity pairs, and each entity usually has corresponding entity description information. These ordered relationship paths, entity descriptions and other information contain rich semantics, which can provide more accurate and reliable auxiliary information for reasoning, thus significantly improving the ability of knowledge graph representation learning and the accuracy of reasoning. However, the existing knowledge graph representation methods that introduce additional information only consider improving the ability of knowledge graph representation learning by fusing a single additional information, and do not make full use of the multi-source information in the knowledge graph. For example, the DKRL model and the MCapsEDD model that use entity description information, and the OPTransE model that uses ordered relationship path information. And these models still have the following deficiencies: The fusion training method of DKRL can well combine entity descriptions and the structural information of triples, but the continuous bag-of-words encoding and convolutional neural network encoding it uses have insufficient representation of entity description information; McapsEDD can efficiently and accurately encode entity descriptions using Transformer and relational attention mechanisms, but the capsule network model it uses has poor processing effect on large datasets; The TransE model, which is the basis of DKRL and OPTransE, has poor processing ability for complex information. Summary of the Invention
[0007] To solve the above problems, the present disclosure proposes a knowledge representation learning method that fuses ordered relationship paths and entity description information. The head and tail entities of each relationship in the knowledge graph are projected into different spaces to ensure the order of the relationship in the path, and a pooling strategy is used to obtain the representation of the ordered relationship path information; Transformer is used in combination with the relational attention mechanism to encode the entity description information to obtain the corresponding semantic representation; Combining with the TransR model that can handle more complex relationships in the knowledge graph, the triple information, ordered relationship path information and entity description information are fused and trained to improve the performance of knowledge graph reasoning.
[0008] According to some embodiments, the present disclosure adopts the following technical solutions:
[0009] A knowledge representation learning method that fuses ordered relationship paths and entity description information, comprising the following steps:
[0010] Extract and preprocess each path relationship information and entity description information in the knowledge graph triples in the database;
[0011] Project the head and tail entities of each extracted path relationship into different spaces, and use a pooling strategy to fuse different path information to obtain the representation of the ordered relationship path information;
[0012] Embed the extracted entity description information using a Transformer combined with a relational attention mechanism to obtain a vector representation of the entity description;
[0013] Fuse the representation of the ordered relational path information and the vector representation of the entity description, and learn the vector representation of the fused entity and relationship in the same continuous low-dimensional vector space.
[0014] Furthermore, for each triple in the knowledge graph, use vectors to represent the entity pair and the relationship; in each triple vector, it includes the head entity, the relationship, and the tail entity.
[0015] According to some other embodiments, the present disclosure also adopts the following technical solutions:
[0016] A knowledge representation learning system that fuses ordered relational path and entity description information, comprising:
[0017] A data extraction module, configured to extract and preprocess each path relationship information and entity description information in the knowledge graph triples in the database;
[0018] A data processing module, configured to project the head and tail entities of each extracted path relationship into different spaces, and use a pooling strategy to fuse different path information to obtain a representation of the ordered relational path information;
[0019] And embed the extracted entity description information using a Transformer combined with a relational attention mechanism to obtain a vector representation of the entity description;
[0020] A data fusion module, configured to fuse the representation of the ordered relational path information and the vector representation of the entity description, and learn the vector representation of the fused entity and relationship in the same continuous low-dimensional vector space.
[0021] According to some other embodiments, the present disclosure also adopts the following technical solutions:
[0022] A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the above method are implemented.
[0023] According to some other embodiments, the present disclosure also adopts the following technical solutions:
[0024] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are executed.
[0025] Compared with the prior art, the beneficial effects of the present disclosure are:
[0026] The present disclosure proposes a knowledge graph representation learning model that integrates multi-source information, including triple information, ordered relation path information, and entity description information, comprehensively improving the ability of knowledge graph representation learning and then performing reasoning.
[0027] Considering the semantic information in entity descriptions, use Transformer and relational attention mechanisms to effectively and accurately encode entity description information; effectively utilize the large amount of ordered relation path information existing in the knowledge graph, and can accurately infer the direct relationship between entity pairs. Here, not only the ordered relation information on the relation path is considered, but also the entity information on the relation path is considered. Combining with the TransR model that can handle more complex relations in the knowledge graph, train an integrated model to improve the performance of knowledge graph reasoning.
[0028] Experiments are carried out on the FB15K, WN18, FB15K-237, and WN18RR datasets. Compared with other benchmark models in the link prediction task, good results are obtained. Moreover, OPDRL shows better complex relation modeling ability in the Hits@10 metric for N-1 and N-N complex relations. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The specification drawings constituting a part of the present disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure.
[0030] Figure 1 It is the overall training model architecture diagram of the present disclosure;
[0031] Figure 2 It is the structural diagram of the ordered relation path representation of the present disclosure;
[0032] Figure 3 It is the entity description embedding learning architecture diagram of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The present disclosure will be further described below in conjunction with the drawings and embodiments.
[0034] It should be noted that the following detailed description is illustrative and is intended to provide a further description of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.
[0035] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0036] Embodiment 1
[0037] The present disclosure provides a knowledge representation learning method that fuses ordered relationship paths and entity description information, and adopts three algorithm models: OPDRL, DKRL(TA) combined with TransR, and OPTransR. 1) OPDRL: This model mainly includes two parts. One part is in the representation of ordered relationship path information. First, in order to distinguish the relationship order on the path, entities and relationships are projected into different spaces. Then, in order to infer missing relationships with these latent representations, a two-layer pooling strategy is used to fuse information from different paths. The other part is in the representation of entity description information. First, the preprocessed entity description information is embedded using Transformer combined with a relationship attention mechanism to obtain a vector representation of the entity description. Then, it is learned in the same vector space as the vector representation of the entity. Finally, the results of these two tasks are further comprehensively co-trained on the basis of TransR to obtain the vector representations of entities and relationships after model fusion. 2) DKRL(TA)+TransR: This model mainly considers the representation of entity description information. On the basis of the DKRL model, Transformer (T) combined with a relationship attention mechanism (A) is used to process the entity description information instead of CNN to obtain a vector representation of the entity description. Then, it is co-trained with the TransR model that can solve the complex relationships between entity pairs. 3) OPTransR: This model mainly considers the representation of ordered relationship path information. On the basis of the OPTransE model, the TransR model that can solve the complex relationships between entity pairs is used to replace the TransE model for training and learning.
[0038] Based on the above models, the following method steps are mainly implemented:
[0039] Step 1: Extract and preprocess each path relationship information and entity description information in the knowledge graph triples in the database;
[0040] Step 2: Project the head and tail entities of each extracted path relationship into different spaces, and use a pooling strategy to fuse information from different paths to obtain a representation of the ordered relationship path information;
[0041] Step 3: Embed the extracted entity description information by using Transformer combined with the relational attention mechanism to obtain the vector representation of the entity description;
[0042] Step 4: Fuse the representation of the ordered relational path information and the vector representation of the entity description, and learn the vector representation of the fused entity and relationship in the same continuous low-dimensional vector space.
[0043] Furthermore, each triple in the knowledge graph is represented by a vector for entity pairs and relationships; in each triple vector, it includes the head entity, the relationship, and the tail entity. Specifically, in the knowledge graph (KG), for each triple (h, r, t) in the KG, vectors are used to represent entity pairs and relationships. represents the head entity, represents the tail entity, represents the relationship r.
[0044] In Step 1, first extract each path relationship information and entity description information in the knowledge graph triples in the database, and perform preprocessing;
[0045] In Step 2, project the head and tail entities of each extracted path relationship into different spaces, and use the pooling strategy to fuse different path information to obtain the representation of the ordered relational path information;
[0046] The specific implementation process is as Figure 2 shown in the structural diagram of the ordered relational path representation. Assume that the path connecting two entities contains the indicative information of the direct relationship between these two entities. To measure such indicative effects and ensure the order of relationships in the path, an energy function E(h, p s=n , t) is defined, where p s=n represents one of the n-step paths from h to t, that is The energy function is as follows:
[0047]
[0048] where,
[0049] h p = f(p, h), t p = g(p, t) (2)
[0050]
[0051] h p and t p respectively represent the representations of the head entity h and the tail entity t in the ordered relational path p, represents the sequence matrix regarding the i-th relationship in the given path p.
[0052] If the relation path from h to t is reasonable, it will obtain a lower energy value.
[0053] The triple (h, r, t) in the knowledge graph can be regarded as a single-step path between h and t. Therefore, the value of E(h, r, t) can be obtained by substituting the direct relation r as p s=1 into Equation (1).
[0054] As can be seen from Equation (1) above, the sequence matrix i before each relation r is different. If the order of several relations in a path changes, the value of the energy function will also change accordingly. Therefore, in our model, paths with the same relation set but different relation orders will infer very different direct relations, and the specific representation of the ordered relation path will be demonstrated in the following content.
[0055] To maintain the ordered information of the relations in the path, two matrices and and are introduced for each relation, representing the projection matrices of the head entity and the tail entity of relation r respectively. Using these two matrices, the head and tail entities are projected into different spaces with respect to the same relation. Specifically, by introducing two matrices for each relation, the head and tail entities of the relation are projected into different spaces. Let and represent the projection matrices of the head entity and the tail entity of relation r respectively. Using these two matrices, the head and tail entities are projected into different spaces with respect to the same relation. Suppose there is a path r1, r2,..., r n , from h to t, ideally, we define the following equation:
[0056]
[0057] where t (i) represents the i-th passed node on the path.
[0058] For entity pairs with relation paths, we obtain their representations after eliminating the transfer nodes from Equation (4). Therefore, the specific forms of the variables in Equation (1) are as follows:
[0059]
[0060]
[0061] where
[0062]
[0063]
[0064] Denote the path as p s=n 's projection matrix, whose purpose is to project the tail entity in the path onto p s=n 's space. Moreover, I in Equation (8) represents the identity matrix, Denote as r k 's head entity space to r k-1 's tail entity space's space transfer matrix, that is
[0065] In Step 2, it is mentioned to use a pooling strategy to fuse different path information to obtain a representation of ordered relational path information. Specifically, a two-layer pooling strategy is designed to fuse different path information. First, the first-layer strategy is to extract feature information from i-step paths using the min-pooling method and define an energy function as follows:
[0066]
[0067] where, represents the set of all i-step paths related to the relationship r from the head entity h to the tail entity t.
[0068] To obtain introduce the conditional probability Pr(r|p s=i ) to represent the reliability of the path p s=i associated with the given relationship r, as follows:
[0069]
[0070] where, Pr(r, p s=i ) represents the joint probability of r and p s=i , Pr(p s=i ) represents the marginal probability of p s=i . In addition, N(r, p s=i ) represents the number of cases where r and p s=i link the same entity pair in the knowledge graph, N(p s=i ) represents the number of paths p s=i in the knowledge graph, and N(p) represents the total number of paths in the knowledge graph. Since N(p) can be removed from both the numerator and denominator of the fractional expression, we finally convert the probability to frequency for calculation.
[0071] By selecting all p s=i for which Pr(r|p s=i ) > 0 from h to t to filter the paths. Therefore, is all the filtered ps=i a set. Sometimes we can infer facts not from the direct relation r, but from paths, which means the value may be less than the value of E(h, r, t).
[0072] The second-layer pooling strategy is to use the min-pooling method to fuse information from paths of different lengths, and define the following energy function:
[0073] E final (h, r, t) = Min[E(h, r, t), E(h, p s=1 , t), E(h, p s=2 , t),..., E(h, p s=n , t)] (11)
[0074] where E(h, r, t) represents the energy value of the direct relation r. It is calculated by substituting r as p s=1 into formula (1).
[0075] E(h, p s=n ,t ) is initialized to infinity, so if there is no i-step path between h and t, it will not affect the result of the final energy function.
[0076] As mentioned above, we adopted the min-pooling strategy twice in the model. For the purpose of min-pooling is to select the path that best matches r among all i-step paths. And for the final energy function E final (h, r, t), we use min-pooling to extract non-linear features from paths of different lengths. In addition, the min-pooling method also solves the problem that there may be no relational path between h and t.
[0077] Furthermore, as mentioned in step 3, the extracted entity description information is embedded using Transformer combined with the relation attention mechanism to obtain the vector representation of the entity description; specifically, as described below, the input layer is used to extract keywords from the entity description information, and then the Transformer encoder and the relation attention mechanism are used to extract the key semantic features contained in the entity description.
[0078] The DKRL model is a classic model that performs knowledge graph representation learning by integrating entity descriptions. This model first extracts keywords from entity descriptions and then uses CBOW and CNN for encoding to obtain corresponding representations. Since these representations do not contain all the semantic information of entity descriptions, there will be a certain loss of semantic information. Therefore, in this part, we use the Transformer encoder and relational attention mechanism used in MCapsEED to effectively and accurately extract the important semantic features contained in entity descriptions. Since the process of obtaining description representations for the head entity and the tail entity is the same, take the head entity as an example.
[0079] The representation model of entity description information consists of three layers, as Figure 3 shown, namely the input layer, the Transformer Encoder layer, and the relational attention layer; in the input layer, the word embedding and position embedding of the entity description are concatenated to form a sentence embedding, which serves as the input information of the model; a Transformer Encoder with a multi-head self-attention mechanism is used to obtain more semantic features in the entity description and learn the dependencies between words in the description; a relational attention mechanism is used to calculate the weights of each word in the entity description to obtain the global vector representation of the entity description.
[0080] Specifically, in the input layer, the word embedding and position embedding of the entity description are concatenated to form a sentence embedding, which serves as the input information of the model. Due to the simplicity and effectiveness of Skip-gram, it is used to train word embeddings from a large number of entity descriptions, setting the size of the skip window to 5 and the dimension of the word embedding to d = 100. After training, the word embedding matrix WordVec ∈ R k×d , where k is the vocabulary size of the entity description.
[0081] By looking up in the obtained word embedding matrix, the embedding e n of each word w i in the entity description desc = {w1, w2,..., w i} can be obtained. Furthermore, the word embedding matrix E = (e1, e,..., e n} of the entity description is obtained, where e i ∈ R d×1 is the embedding of the i-th word of the entity description, and n is the length of the entity description.
[0082] To capture the sequential order of entity descriptions, a position embedding matrix PosVec ∈ R k×d is obtained through random initialization, and this matrix is updated during training. The position embedding is obtained by looking up the position embedding matrix. After converting each position index to a position embedding, the position embedding matrix P = (p1, p2,..., pn ), where p i ∈R d×1 is the position embedding of the i-th word in the description, and n is the length of the entity description.
[0083] The conjunction embedding and the position embedding are concatenated to obtain the output vector S = (s1, s2,..., s n ), where s i ∈R 2d ×1 is the concatenation of e i and p i .
[0084] In the Transformer encoder layer, to obtain more semantic features in the entity description and learn the dependencies between words in the description, a Transformer Encoder with a multi-head self-attention mechanism is adopted. The encoder consists of N = 6 identical layers. Each layer has two sub-layers, a multi-head self-attention layer and a simple position-wise fully-connected feed-forward network. Residual connections are used after each sub-layer, followed by layer normalization. Finally, the output of the Transformer encoder layer is the representation of each word and its contextual semantic information H = (h1, h2,..., h n ), which serves as the input to the relation attention layer.
[0085] Furthermore, to further improve the processing ability in the entity description, power normalization PN is used instead of layer normalization LN for normalization in the Transformer encoding layer. PN replaces the variance with the quadratic mean and uses an approximate backpropagation to capture the running statistics, achieving the optimal effect in multiple tasks of NPL compared to the Transformer using LN.
[0086] In the relation attention layer, to obtain the global vector representation of the entity description, a relation attention mechanism is adopted to calculate the weight of each word in the entity description, and the global representation of the entity description is taken as the weighted sum of the representations of each word in the entity description. To identify which words in the entity description are closely related to the entity and the relation, a simple fully-connected neural network is used to calculate the weight of each word in the entity description. The input of this network is the head entity representation h s and the relation representation h r , as well as the contextual feature representation h i of each word. The weights are calculated in (12)(13)(14).
[0087] f i = W a [h i ; h s + Vhr (12)
[0088]
[0089]
[0090] where W a is the weight matrix, V ∈ R d×1 is the parameter vector, and h d is the representation of the head entity description. In the same way, the representation t of the tail entity description can be obtained d .
[0091] By combining the triples in the knowledge graph with the entity description information, it is possible to better learn the optimal vector representations of entities and relationships. The structure-based representation can better capture the factual triple information in the knowledge graph, while the entity-description-based representation can better capture the text information. Usually, similar entities should have similar description information and similar keywords. These relationships are difficult to obtain directly from the structural information, but they may be discovered through the internal connections of the keywords. Learning the structure-based representation and the entity-description-information-based representation simultaneously in the same continuous low-dimensional vector space will likely result in better representation capabilities.
[0092] Based on the above analysis, the energy function based on the entity-description information representation is defined as:
[0093] E d = E dd + E ds + E sd (15)
[0094] where E dd = ||h d + r - t d ||, E ds = ||h d + r - t s ||, E sd = ||h s + r - t d ||. h s and t s represent the structure-based representation; h d and t d represent the entity-description-information-based representation. E dd represents the definition of the energy function where both the head and tail entities are based on the entity-description information representation; E ds and E sdIt represents the definition of an energy function that is based on structural representation and another based on entity description information. By defining the energy function in the above manner, information such as structural entity description can be applied to training and learning simultaneously, thereby better obtaining the vector representations of entities and relationships.
[0095] In step 4, it is mentioned that the representation of ordered relationship path information and the vector representation of entity description are fused, and in the same continuous low-dimensional vector space, the vector representation of entity and relationship fusion is learned.
[0096] In specific implementation, in order to better learn the optimal representation of entities and relationships, the OPDRL model of the present disclosure combines the triple structure information, relationship paths, and entity descriptions of the knowledge graph to comprehensively train the model. In the same continuous low-dimensional vector space, the vector representations of entities and relationships are learned. The comprehensive energy function is defined as follows:
[0097] E = E s + E p + E d (16)
[0098] Where E s is the energy function based on structural representation, E p is the energy function based on relationship path representation defined in Equation (11), and E d is the energy function based on entity description information representation defined in Equation (15).
[0099] Based on Equation (16), the energy function based on structural representation and entity description representation can be obtained, E s Adopting the definition of TransR, the specific calculation expression is as follows.
[0100]
[0101] Based on the above calculation and analysis, a loss function is further constructed. The optimization method based on the margin is defined and used as the training objective, and the model is optimized by minimizing the loss function, as shown in Equation (18):
[0102]
[0103] Where L(h, r, t) represents the loss function of the triple (h, r, t), and L(h, p s=i , t) represents the loss value with respect to the relationship path p s=i . The probability pr(p s=i |h, t) represents the reliability of the relationship path p s=i given the entity pair (h, t), and Pr(r|p s=i ) represents the path p related to the given relationship rs=i Reliability is a normalization factor used to balance the triple loss and the path loss
[0104]
[0105]
[0106] -E(h′, p s=i , t′), 0) (21)
[0107] γ i is the boundary parameter for separating positive and negative samples. We use different boundary parameters for paths with different numbers of steps. E(h, r, t) is the energy function based on the structural and entity description information, and E(h, p s=i , t) is the energy function based on the relation path. T is the set of positive examples composed of the correct triples (h, r, t), and T′ is the set of negative examples composed of the incorrect triples (h′, r, t′). The definition of T′ is given as follows:
[0108] T′ = ((h', r, t) ∪ (h, r, t')} (22)
[0109] Here, T is obtained by randomly replacing the head entity or the tail entity in the set of positive examples to get the corresponding negative instances. During the model training process, the stochastic gradient descent method is used for optimization to minimize the value of the loss function
[0110] The above method of the present disclosure is verified by experiments. The datasets used in the experiments of the present disclosure are the standard datasets of FB15K, WN18, FB15K - 237, and WN18RR. Each entity in the datasets has corresponding short description information. Among them, FB15K is extracted from the large - scale knowledge base FreeBase, FB15K - 237 is a subset of FB15K with the reverse relations in FB15K removed; WN18 is extracted from the WordNet knowledge base, and WN18RR is a subset of WN18 with the reverse relations in WN18 removed. The present disclosure divides the datasets into training datasets, validation datasets, and test datasets. The relevant situations of the datasets used are shown in Table 1
[0111] Table 1 Statistical situation of the datasets used
[0112]
[0113] The representative baseline models used in the comparative analysis of the present disclosure experiment include: TransE, TransH, TransR, DKRL(CNN)+TransE, PTransE, TEKE_H, RPE, DistMult, ComplEx, ConvE, RotatE, OPTransE, Interstellar, MCapsEED. Among them, PTransE, RPE, and OPTransE utilize the path information between entity pairs, and DKRL(CNN)+TransE, TKHE_H, and MCapsEED use entity description information.
[0114] Based on the model framework proposed in the present disclosure, the achievable prediction models include: 1) DKRL(TA)+TransR, which combines the entity description vector representations encoded by the Transformer and the relational attention module and trains and learns together with TransR; 2) OPTransR, which is based on OPTransE and trains and learns the ordered relational path together with TransR; 3) OPDRL, which combines the information based on the relational path, the information based on the entity description, and the structural information based on TransR and trains and learns together.
[0115] To evaluate the performance of the models, a link prediction experiment was conducted. Link prediction refers to predicting the missing triples in the knowledge graph, that is, the missing entities or relations in the triples. In the experiment, for the missing entities or relations, a set of candidate entities is selected from the knowledge graph entity set for ranking, rather than only giving a best result. All entities are used to replace the head entity or the tail entity to create a set of negative example triples. A scoring function is used to calculate their similarity scores for these negative example triples, and they are ranked based on this. The higher the similarity, the higher the ranking, and thus the true ranking of the correct entity can be obtained.
[0116] We use the following three metrics as evaluation criteria: (1) The average rank MR (MeanRank) of the correct entity, the smaller the better; (2) The average reciprocal rank MRR (Mean Reciprocal Rank) of the correct entity, the larger the better; (3) The percentage Hits@10 of the correct entity entering the top 10, the larger the better. In fact, there may also be incorrect triples in the knowledge graph that are considered correct. Therefore, the present disclosure filters out the incorrect triples from the dataset. The filtered setting is called Filter, and the original is called Raw. Additionally, on the datasets FB15K-237 and WN18RR, only the Filter setting is used, that is, incorrect triples are not considered.
[0117] The optimal parameter settings of the disclosed model are as follows: On FB15K, the learning rate λ = 0.0005, the margin γ = 4.0, γ1 = 4.5, γ2 = 5.0, and the dimensions d of the entity, relation, and entity description vector representations are 100. On WN18, the learning rate λ = 0.0001, the margin γ = 5.0, γ1 = 5.0, γ2 = 5.5, and d = 100. On FB15K-237, the learning rate λ = 0.0001, the margin γ = 4.0, γ1 = 4.0, γ2 = 4.5, and d = 100. On WN18RR, the learning rate λ = 0.0001, the margin γ = 5.0, γ1 = 5.5, γ2 = 5.5, and d = 100. In addition, we use the L1 norm for scoring. Since there will be many relationship paths between entity pairs, if all lengths of relationship paths are traversed, the computational cost will be very high. Therefore, in the model effect verification and analysis of the present disclosure, the two-step relationship paths between entity pairs in the knowledge graph are mainly focused on.
[0118] The experimental results of link prediction on the FB15K and WN18 datasets for multiple models implemented in this disclosure and the baseline model are shown in Table 2. From the results in Table 2, it can be observed and analyzed that: 1) On the two datasets, the DKRL(TA)+TransR model performs better than the DKRL(CNN)+TransE model. On the FB15K dataset, Hits@10 (filter) has increased by 9%. This indicates that using Transformer and the relational attention mechanism can better obtain the semantic representation of entity description information. Moreover, this also shows that jointly training with the TransR-based model can better utilize the structural information of the knowledge graph, process relatively complex relational information, and achieve better results. 2) The performance of the OPTransR model is superior to that of the OPTransE model, which also indicates that when modeling and training ordered relational paths, the effect of combining with TransR is better than that of TransE. This is because TransR can handle relatively complex relational information in the knowledge graph. 3) On the FB15K and WN18 datasets, DKRL(TA)+TransR and OPTransR are superior to TransR in all evaluation metrics. This shows that on the TransR-based structured model, whether fusing ordered relational path information or entity description information, it can improve the representation ability of entities and relationships in the knowledge graph to a certain extent, and further promote entity prediction. 4) On the two datasets, the OPDRL model proposed in this disclosure is superior to all baseline models in the evaluation metrics MR and Hits@10. Compared with the best-performing baseline model on the FB15K dataset, MR (filter) has decreased by 13%, and Hits@10 (filter) has increased by 3%. This comparison result shows that the OPDRL model has a more accurate knowledge representation ability than other baseline models. Integrating the information based on ordered relational paths, entity descriptions, and TransR-based structural information can well represent the entities and relationships in the knowledge graph, promote the reasoning performance of entity prediction, and improve the prediction accuracy to a certain extent. 5) Comparing the experimental results of DKRL(TA)+TransR, OPTransR, and OPDRL, it can be found that the effect of OPDRL is better than that of the other two models. On the WN18 dataset, the evaluation metric Hits@10 (filter) has increased by 4% compared with DKRL(TA)+TransR and by 1% compared with OPTransR, and MR (filter) has decreased by 9% compared with DKRL(TA)+TransR and by 6% compared with OPTransR. This comparison result shows that simultaneously fusing the semantic information of ordered relational paths and entity descriptions is more effective than only using relational path or entity description information in improving the knowledge graph representation learning ability of the model.
[0119] Table 2. Link prediction experiment results on FB15K and WN18 datasets
[0120]
[0121] We further compared the performance of OPDKL and baseline models on different types of relationships on the FB15K dataset. As can be seen from Table 3, OPDKL achieved the highest scores in all subtasks, especially performing best on complex relationships of N-to-1 and N-to-N. When predicting the tail entity of the N-to-1 relationship, ODPKE increased Hits@10 to 97.6%, and the average prediction accuracy for the N-to-N relationship on the two datasets reached 92.8%. In the representation of ordered relation paths, we projected the head and tail entities of the triple into different relation spaces respectively to better distinguish relevant entities. In the representation of entity description information, the relation attention mechanism we used can effectively associate entities with relations. And our model is co-trained with TransR that can handle relatively complex relationships. Therefore, these results show that integrating ordered relation paths and entity descriptions simultaneously can be used as a supplement to structured models, effectively improving the ability to model complex relationships and promoting relation prediction.
[0122] Table 3. Evaluation results of mapped relation attributes on FB15K (%)
[0123]
[0124] As shown in Table 4, in order to understand the impact of these two parts, namely the relation attention mechanism and PowerNormalization, on the model in the representation of entity description information, we conducted ablation experiments. To further improve the ability to capture entity descriptions, we used Power Normalization instead of LayerNormalization for normalization in the Transformer Encode layer. Compared with OPDRL using LN, using PN increased the model by 0.4% in Hits@10. To make full use of entity description information related to relations, we added a relation attention mechanism to OPDRL, denoted as A, which also improved the knowledge representation effect by 0.3%. This fully shows that these two modules play a very good role in processing entity descriptions and improve the overall performance of the model.
[0125] Table 4. Model performance on the FB15k-237 dataset with different integration methods and optimization strategies.
[0126]
[0127] In the link prediction experiment, the present disclosure also evaluated the experimental effects of the OPDRL model and some advanced benchmark models on the FB15K-237 and WN18RR datasets. As can be seen from Table 4: 1) Compared with other advanced benchmark models, the OPDRL model proposed in the present disclosure reached a higher level. On the two datasets, compared with the best results of the benchmark models, the OPDRL model improved by 4% and 2% respectively on the evaluation metric Hits@1. This indicates that the OPDRL model effectively supplements the structured TransR model by integrating relationship paths and entity description information, and can better improve the ability of knowledge graph representation learning, thereby improving the prediction performance. 2) By comparing the experimental results of OPDRL and the current state-of-the-art relationship path model Interstellar, it can be found that OPDRL has varying degrees of improvement in various metrics on the two datasets compared to Interstellar. This shows that on a structured model, introducing ordered relationship paths and entity description information performs better than using only relationship paths to predict missing information in the knowledge graph. 3) Compared with the McapsEED model that uses capsule networks to process triples, the model of the present disclosure performs poorly in terms of Hits@10 and MRR metrics on the FB15K-237 dataset. This may be due to the small size of the dataset, which cannot fully utilize entity description information and ordered relationship path information, and capsule networks perform better on datasets with closer relationships. On the larger-scale dataset WN18RR, all metrics of the model of the present disclosure are better than those of the McapsEED model, with Hit@10 increasing by 2 percentage points and MRR increasing by 14%. This shows that using relationship paths and entity description information simultaneously as supplementary structural information can effectively improve the model's performance and is more suitable for larger-scale datasets.
[0128] Table 5. Link prediction results on datasets B15K-237 and WN18RR
[0129]
[0130] The present disclosure conducted experiments on the FB15K, WN18, FB15K-237, and WN18RR datasets and achieved good results when compared with other benchmark models in the link prediction task. Moreover, OPDRL demonstrated better complex relationship modeling capabilities in terms of the Hits@10 metric for N-1 and N-N complex relationships.
[0131] Example 2
[0132] The present disclosure provides a knowledge representation learning system that integrates ordered relationship paths and entity description information, including:
[0133] A data extraction module, configured to extract and preprocess each path relationship information and entity description information in the knowledge graph triples in the database;
[0134] A data processing module, configured to project the head and tail entities of each extracted path relationship into different spaces, and use a pooling strategy to fuse different path information to obtain a representation of ordered relationship path information;
[0135] And use a Transformer combined with a relational attention mechanism to embed the extracted entity description information to obtain a vector representation of the entity description;
[0136] A data fusion module, configured to fuse the representation of ordered relationship path information and the vector representation of entity description, and learn the vector representation of entity and relationship fusion in the same continuous low-dimensional vector space.
[0137] Embodiment 3
[0138] The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above method are implemented.
[0139] Embodiment 4
[0140] The purpose of this embodiment is to provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the above method are executed.
[0141] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0142] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1means for the functions specified in one or more boxes.
[0143] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one Figure 1 one or more processes and / or boxes Figure 1 or more boxes.
[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one or more processes and / or boxes Figure 1 or more boxes.
[0145] The above is only the preferred embodiment of the present disclosure and is not intended to limit the present disclosure. For those skilled in the art, the present disclosure can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
[0146] Although the specific implementation manners of the present disclosure have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that, based on the technical solutions of the present disclosure, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present disclosure.
Claims
1. A knowledge representation learning method that fuses ordered relation paths and entity description information, characterized in that It includes the following steps: Extract and preprocess each path relation information and entity description information in the knowledge graph triples in the database; Project the head and tail entities of each extracted path relation into different spaces, and use a pooling strategy to fuse different path information to obtain a representation of ordered relation path information; Embed the extracted entity description information using Transformer combined with relation attention mechanism to obtain a vector representation of the entity description; Fuse the representation of ordered relation path information and the vector representation of the entity description, and in the same continuous low-dimensional vector space, learn the vector representation of the fused entity and relation; Among them, the step of projecting the head and tail entities of each extracted path relation into different spaces and using a pooling strategy to fuse different path information to obtain a representation of ordered relation path information includes: Design a two-layer pooling strategy to fuse information from different paths. First, the first-layer strategy is to extract feature information from the i-step path using the min-pooling method and define an energy function as follows: Among them, represents the set of all i-step paths related to the relation r from the head entity h to the tail entity t; The second-layer pooling strategy is to use the min-pooling method to fuse information from paths of different lengths, and define the following energy function: E final (h, r, t) = Min[E(h, r, t), E(h, p s=1 , t), E(h, p s=2 , t), …, E(h, p s=n , t)] Among them, E(h,r,t) represents the energy value of the direct relation r; The step of embedding the extracted entity description information using Transformer combined with relation attention mechanism to obtain a vector representation of the entity description includes: Use the DKRL model to extract keywords from the entity description information, and then use the Transformer encoder and relation attention mechanism to extract the key semantic features contained in the entity description.
2. The knowledge representation learning method for integrating ordered relation paths and entity description information according to claim 1, wherein For each triple (h,r,t) in the knowledge graph, use vectors to represent the entity pair and the relation; in each triple vector, it includes the head entity h, the relation r, and the tail entity t.
3. A knowledge representation learning method that fuses ordered relation paths and entity description information, characterized in that To measure the direct relationship between two entities and ensure the order of relationships in the path, an energy function E(h, p s=n , t) is defined, where p s=n represents one of the n-step paths from h to t, that is The energy function is as follows: Among them, h p = f(p, h), t p = g(p, t) h p and t p respectively represent the representations of the head entity h and the tail entity t in the ordered relation path p. represents the sequence matrix for the i-th relation in the given path p.
4. The knowledge representation learning method that fuses the ordered relationship path and entity description information as described in claim 1, characterized in that The representation model of the entity description information consists of three layers, namely the input layer, the Transformer Encoder layer, and the relation attention layer; in the input layer, connect the word embedding and position embedding of the entity description to form a sentence embedding as the model input information; use a Transformer Encoder with a multi-head self-attention mechanism to obtain more semantic features in the entity description and learn the dependency relationship between words in the description; use the relation attention mechanism to calculate the weight of each word in the entity description to obtain the global vector representation of the entity description.
5. A knowledge representation learning system that integrates ordered relation paths and entity description information, specifically executing a knowledge representation learning method that integrates ordered relation paths and entity description information as described in any one of claims 1-4, characterized in that, It includes: A data extraction module for extracting and preprocessing each path relation information and entity description information in the knowledge graph triples in the database; A data processing module for projecting the head and tail entities of each extracted path relation into different spaces and using a pooling strategy to fuse different path information to obtain a representation of ordered relation path information; and for embedding the extracted entity description information using Transformer combined with relation attention mechanism to obtain a vector representation of the entity description; A data fusion module for fusing the representation of ordered relation path information and the vector representation of the entity description, and in the same continuous low-dimensional vector space, learning the vector representation of the fused entity and relation.
6. A computer device, characterized in that, Comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor implements the steps of any one of the methods as described in claims 1-4 above when executing the program.
7. A computer-readable storage medium, characterized in that, A computer program is stored, and when the program is executed by a processor, it executes the steps of any one of the methods as described in claims 1-4 above.