A Knowledge Graph Representation Method Based on the Embedding of Hypercomplex Numbers in Arbitrary Dimensions
By introducing a knowledge graph representation method of super-complex embedding in any dimension, the problem of predefined dimension limitation in the existing technology is solved, and more efficient knowledge graph representation is achieved, which improves the flexibility and performance of the model.
Patent Information
- Application Number
- CN202210848061.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-19
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-07-19
AI Technical Summary
The existing knowledge graph representation methods have limited predefined dimensions in hyperplural space, which limits the flexibility and efficiency of the model and makes it difficult to effectively process incomplete knowledge graph data.
Using a knowledge graph representation method based on hyper-complex embedding in any dimension, the hyper-complex embedding linear layer HyperE is constructed, and the model parameters are optimized using the Kronecker product and self-adversarial negative sampling loss function to achieve weight sharing and greater structural flexibility.
It reduces memory usage, improves the performance and efficiency of the model, and achieves excellent embedding results, adapts to the training needs of different data sets.
Smart Images

Figure CN115168612B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a knowledge graph representation method, and in particular to a knowledge graph representation method based on arbitrary dimensional hypercomplex embedding. Background Art
[0002] A knowledge graph is a collection of triples, where each triple consists of a head entity h, a relation r, and a tail entity t. Existing knowledge graph datasets include Freebase, Yago, and WordNet. Knowledge graphs are widely used in applications such as question answering, information retrieval, recommendation systems, and natural language processing. Research on knowledge graphs is attracting increasing attention from the industry. Knowledge graphs are often incomplete, so a fundamental problem in knowledge graphs is predicting missing links. To facilitate the application of knowledge in downstream tasks, many researchers have attempted to model this data in a computable manner, and a variety of knowledge graph representation methods have emerged.
[0003] 1. Knowledge Graph Representation Method
[0004] The general knowledge graph representation method is to model and infer the connectivity patterns in the knowledge graph based on the observed knowledge facts. For example, some relationships are symmetric (such as marriage), some relationships are antisymmetric (such as parent-child relationships); some relationships are opposite (such as hyponyms); and some relationships are composed of other relationships (for example, my mother's husband is my father). Find ways to model and infer these patterns from the observed knowledge graph dataset to predict missing. Many existing methods have been trying to implicitly or explicitly model one or several of the above relationship patterns. For example, the classic distance-based embedding model TransE aims to simulate opposite and combined relationships; the DisMult model models the three-way interaction between the head entity, relationship, and tail entity, aiming to model symmetric relationships; the PairRE model can simultaneously encode complex relationships and multiple relationship patterns, using two vectors to represent the relationship, and projecting the corresponding head entity and tail entity into the Euclidean space with these vectors to minimize the distance between the projected vectors.
[0005] 2. Hypercomplex Method
[0006] Among existing knowledge graph embedding methods, previous attempts have included complex number embedding (ComplEx) and quaternion embedding (QuatE). The ComplEx model uses complex number embedding to handle various binary relationships, including symmetric and antisymmetric relationships, using the Hermitian product. The QuatE model uses quaternion embedding, a hypercomplex number embedding with three imaginary parts, to capture potential interdependencies using the Hamilton product. Quaternions can also express rotations in four-dimensional space, providing more degrees of freedom than rotations in complex space.
[0007] Recent research has shown that in the hypercomplex space, representation learning has achieved good results, such as the complex embedding ComplEx and the quaternion embedding QuatE mentioned above. However, the hypercomplex space only exists in very few predefined dimensions (4D, 8D, and 16D), which limits the flexibility of models that utilize hypercomplex multiplication. Hypercomplex embeddings of arbitrary dimensions adopt a method of parameterizing hypercomplex multiplication, allowing the model to learn the multiplication rules from the data, regardless of whether these rules are predefined. Operations in an arbitrary n-dimensional hypercomplex space use 1 / n of the learnable parameters and provide greater architectural flexibility compared to the method corresponding to a fully connected layer. This technique has been applied to practical applications such as natural language processing, machine translation, and text style transfer, demonstrating the architectural flexibility and effectiveness of the hypercomplex embedding method. Summary of the Invention
[0008] The present invention discloses a knowledge graph representation method based on hypercomplex embeddings of arbitrary dimensions. The main feature of this method is that the quaternion embedding part in the original knowledge graph embedding method is replaced with a linear layer of hypercomplex embeddings. It specifically includes the following steps: 1. Preprocess the knowledge graph data, and preprocess the traditional knowledge graph into structured data according to the model requirements; 2. Use the deep learning framework pytorch to construct a preliminary embedding and build a new linear layer, that is, a hypercomplex embedding linear layer, to learn the vector representations of entities and relationships on the graph; 3. Use the knowledge graph validation set for verification and adjust to the optimal network parameters; 4. Test the knowledge graph test set and count the results. The present invention improves an existing quaternion knowledge graph embedding method QuatE, introduces a hypercomplex strategy, reduces memory occupancy, reduces parameters, and at the same time maintains excellent embedding results.
[0009] To implement this embedding method, it is necessary to deploy and configure a python3.7 and pytorch1.4.0 running environment under the Linux environment.
[0010] A knowledge graph representation method based on hypercomplex embeddings of arbitrary dimensions. The specific implementation process of this solution is as follows: Step 1 is specifically as follows: First, preprocess the knowledge graphs in different fields into five files. The processed files include the knowledge graph triple training set, the knowledge graph triple validation set, the knowledge graph triple test set, the entity ID set, and the relationship ID set.
[0011] Step 2 is specifically as follows: First, embed the entities and relationships in the knowledge graph obtained in Step 1 into initial vectors to prepare for the subsequent training. Construct a hypercomplex embedding (Hypercomplex Embedding) linear layer, that is, a HyperE layer, and obtain the initial embedding result I of an n-ary number from the input, where n is the set arity, and the dimension of I is divisible by n.
[0012] I = [I1, I2, I3, …, I n (1)
[0013] I1 represents the real part of the n - ary number embedding, and I i , i ∈ [2, 3, …, n] represents the imaginary part of the n - ary number embedding. These n parts are concatenated along a given axis to form a vector I, which is used as the input of the HyperE layer. The HyperE layer adopts the same form as the standard translation model: y = HyperE(x) = Ux + b. The key idea is to construct U as a parameter matrix through the sum of Kronecker products, where x is the input embedding vector to be trained, b is the bias, and y is the embedding vector of the entity or relationship;
[0014] Calculate the scores of positive and negative samples, and calculate the loss of each batch of data through the scores for iterative optimization.
[0015] The pre - processing operation described in step 1 is specifically as follows: Randomly divide the entire knowledge graph triple dataset into a training set, a validation set, and a test set according to 8:1:1, and output the entity - corresponding IDs and relationship - corresponding IDs of the entire knowledge graph.
[0016] The operation of constructing the hyper - complex linear layer HyperE described in step 2 is specifically as follows:
[0017] Obtain the entity embeddings and relationship embeddings in the knowledge graph. For any triple with a head entity h, a relationship r, and a tail entity t, next, the HyperE layer converts the entity and relationship embeddings into high - order embeddings, y = HyperE(x) = Ux + b. Different learning matrices U are constructed according to different arities in the way of Kronecker product. The Kronecker product generalizes the outer product of vectors to matrices. Let X ∈ R m*n , Y ∈ R p*q , and the Kronecker product is:
[0018]
[0019] where x ij =(X) i,j . Let n be the dimension of the hyper - complex embedding HyperE, and k be a user - defined hyperparameter representing the dimension of entity and relationship embeddings. The above - mentioned U matrix is obtained from n Kronecker products:
[0020]
[0021] where C i ∈R n*n represents the contribution matrix, Denote the component weight matrix, and at the same time, \(k\) also represents the input and output sizes of the linear transformation, and the contribution matrix \(C\) i is selected as a full-rank matrix, whose rows and columns are linearly independent, and all elements belong to \(\{-1, 0, 1\}\). Let
[0022]
[0023] 1 and -1 alternate on the diagonal, and each contribution matrix \(C\) i is initialized as the matrix which is the product with the power of the cyclic permutation matrix \(P\) n . The role of the cyclic permutation matrix \(P\) n is to shift the columns to the right,
[0024]
[0025] where when \(j - 1 = 1\) and \(i = n, j = 1\), \((P\) n ) i,j = 1, and all other terms are 0; in addition, this is not the only initialization method, and random initialization can also be used. The contribution matrix \(C\) i is set to be learnable. When the arity of the hypercomplex linear layer is set to binary and quaternary, that is, \(n = 2\) or 4, the contribution matrix is not constructed as above, but an initialization matrix is constructed in the same way as complex and quaternion algebras.
[0026] When \(n = 2\), set
[0027]
[0028] When \(n = 4\), set
[0029]
[0030] Among them, in step 2, the triple score is calculated through the scoring function, and the specific operation is as follows:
[0031] Let \(y_1\) be obtained by adding the head entity \(h\) and the relation \(r\):
[0032] \(y_1 = h + r\) (6)
[0033] And \(y_2\) is the tail entity \(t\):
[0034] \(y_2 = t\) (7)
[0035] Pass \(y_1, y_2\) through the HyperE layer to get:
[0036] \(y'_1 = GHyperE(y_1), y'_2 = HyperE(y_2)\) (8)
[0037] Define the score obtained through the distance function as:
[0038] d r (h, t) = ||y′1 - y′2|| = ||HyperE(h + r) - HyperE(t)|| (9).
[0039] In step 2, learn the vector representations of entities and relationships through the loss function. The specific operation is as follows:
[0040] Negative samples are very effective for learning knowledge graph embeddings and word embeddings. Use a loss function similar to the negative sampling loss to optimize the distance-based model, that is, the self-adversarial negative sampling method. Sample negative triples according to the current embedding model. Let:
[0041]
[0042] where α is the sampling parameter, (h i , r i , t i ) is a positive triple, and (h′ i , r, t′ i ) is a negative triple.
[0043] The final sampling loss of self-adversarial training is:
[0044]
[0045] where γ is a fixed margin and σ is the sigmoid function.
[0046] Compared with the prior art, the advantages of the present invention are as follows: A knowledge graph embedding model based on hypercomplex numbers proposed by the present invention introduces the mechanism of arbitrary arity embedding in the model, realizes the hypercomplex embedding layer HyperE layer, which inherently includes a weight sharing mechanism, and specifies that the multiplication rule of the algebra itself is inferred from the data during training. By adjusting the arity n of the hypercomplex numbers, a better model can be obtained according to different training data. At the same time, increasing the arity of the hypercomplex numbers can obtain a model with higher memory efficiency and performance advantages. When the arity of the hypercomplex embedding is set appropriately, better results can be achieved compared with traditional methods, and it has greater architectural flexibility than the fixed arity hypercomplex embedding. Brief Description of the Drawings
[0047] Figure 1 is the overall flowchart of the present invention;
[0048] Figure 2 is the schematic diagram of the HyperE layer;
[0049] Figure 3 is the flowchart comparison diagram between the traditional method and the inventive method. Detailed implementation mode
[0050] The technical solution of the present invention will be described in detail below in conjunction with the drawings and embodiments.
[0051] The knowledge graph representation technology based on hypercomplex embedding proposed in the present invention includes a training part and a testing part. As Figure 1 shown, the network needs to be trained first. After the training is completed, the embedded results are saved, and then the embedded results are used to test the test set.
[0052] Embodiment 1: A method for representing a knowledge graph based on hypercomplex embedding in any dimension. The specific implementation process of this solution is as follows: Step 1 is specifically as follows: First, preprocess the knowledge graphs in different fields into five files. The processed files include the knowledge graph triple training set, the knowledge graph triple validation set, the knowledge graph triple test set, the entity ID set, and the relationship ID set.
[0053] Step 2 is specifically as follows: First, embed the entities and relationships in the knowledge graph obtained in Step 1 into initial vectors to prepare for the next training. Construct a hypercomplex embedding (Hypercomplex Embedding) linear layer, that is, the HyperE layer. Obtain the initial embedding result I of the n - number from the input. n is the set number of elements, and the dimension of I is divisible by n
[0054] I = [I1, I2, I3, …, I n (1)
[0055] I1 represents the real - part of the n - number embedding, and I i , i ∈ [2, 3, …, n] represents the imaginary - part of the n - number embedding. These n parts are concatenated along the given axis to form the vector I, which is used as the input of the HyperE layer. The HyperE layer adopts the same form as the standard translation model: y = HyperE(x) = Ux + b. The key idea is to construct U as a parameter matrix through the sum of Kronecker products, where x is the input embedding vector to be trained, b is the bias, and y is the embedding vector of the entity or relationship;
[0056] Calculate the scores of positive and negative samples, and calculate the loss of each batch of data through the scores for iterative optimization.
[0057] The preprocessing operation described in Step 1 is specifically: Randomly divide the entire knowledge graph triple data set into a training set, a validation set, and a test set according to 8:1:1, and output the entity - corresponding IDs and relationship - corresponding IDs of the entire knowledge graph.
[0058] The operation of constructing the hypercomplex linear layer HyperE described in Step 2 is specifically:
[0059] Obtain the entity embeddings and relationship embeddings in the knowledge graph. For any triple with head entity h, relationship r, and tail entity t, next, the HyperE layer converts the entity and relationship embeddings into high-order embeddings, y = HyperE(x) = Ux + b. Different learning matrices U are constructed according to different arities through the Kronecker product. The Kronecker product generalizes the outer product of vectors to matrices. Let X ∈ R m*n , Y ∈ R p*q , and the Kronecker product is:
[0060]
[0061] where x ij = (X) i,j . Let n be the dimension of the hypercomplex embedding HyperE, and k be a user-defined hyperparameter representing the dimension of the entity and relationship embeddings. The U matrix described above is obtained from n Kronecker products:
[0062]
[0063] where C i ∈ R n*n represents the contribution matrix, represents the component weight matrix. At the same time, k also represents the input and output size of the linear transformation. The contribution matrix C i is selected as a full-rank matrix, whose rows and columns are linearly independent, and all elements belong to {-1, 0, 1}. Let
[0064]
[0065] have 1 and -1 alternating on the diagonal. Each contribution matrix C i is initialized as the product between the matrix and the powers of the cyclic permutation matrix P n . The role of the cyclic permutation matrix P n is to shift the columns of to the right.
[0066]
[0067] where when j - 1 = 1 and i = n, j = 1, (P n ) i,i = 1, and all other terms are 0; in addition, this is not the only initialization method, and random initialization can also be used. The contribution matrix C i is set to be learnable. When the arity of the hypercomplex linear layer is set to binary and quaternary, that is, n = 2 or 4, instead of constructing the contribution matrix as above, an initialization matrix is constructed in the same way as complex and quaternion algebras.
[0068] When n = 2, set
[0069]
[0070] When n = 4, set
[0071]
[0072] Among them, the scoring function described in step 2, the specific operation is:
[0073] Let y1 be obtained by adding the head entity h and the relation r:
[0074] y1 = h + r (6)
[0075] And y2 is the tail entity t:
[0076] y2 = t (7)
[0077] Pass y1 and y2 through the HyperE layer to obtain:
[0078] y′1 = GyperE(y1), y′2 = HyperE(y2) (8)
[0079] Define the score obtained through the distance function as:
[0080] d r (h, t) = ||y′1 - y′2|| = ||HyperE(h + r) - HyperE(t)|| (9).
[0081] The loss function described in step 2, the specific operation is:
[0082] Negative samples are very effective for learning knowledge graph embeddings and word embeddings. Use a loss function similar to the negative sampling loss to optimize the distance-based model, that is, the self-adversarial negative sampling method. Sample negative triples according to the current embedding model. Let:
[0083]
[0084] Among them, α is the sampling parameter, (h i , r i , t i ) is a positive triple, (h′ i , r, t′ i ) is a negative triple.
[0085] The final sampling loss of the self-adversarial training is:
[0086]
[0087] where γ is a fixed boundary and σ is the sigmoid function.
[0088] Example 2:
[0089] Taking the FB15k-237 knowledge graph dataset as an example, the steps of the present invention will be described in detail by performing knowledge graph embedding representation on it.
[0090] Experimental conditions: A computer is selected for network training. The computer is configured with an Intel(R) processor (3.2 GHz) and 124 GB of random access memory (RAM), an Ubuntu 14.04 64-bit operating system, and an NVIDIA GTX 3090 (24 GB) graphics card; the software environment is the deep learning framework pytorch1.4.0.
[0091] Experimental object: The training dataset comes from FB15k-237. The validation set and test dataset used during network training also come from this dataset, and the three are strictly distinguished. FB15k-237 is a subset of the knowledge graph Freebase. 15k means there are 15k subject words in the knowledge base, and 237 means there are a total of 237 types of relationships.
[0092] Experimental steps:
[0093] Step 1: Preprocessing operation. The entire FB15k-237 dataset is divided into a training set, a validation set, and a test set according to a certain ratio, and the entity corresponding IDs and relationship corresponding IDs of the entire knowledge graph are obtained. The results after preprocessing are shown in Table 1:
[0094] Table 1 Results after preprocessing of the FB15k-237 dataset
[0095] Statistical attribute \ set Training set Validation set Test set Number of entities 13781 7652 8171 Total number of triples 272115 17535 20466 Number of relation types 237 223 224
[0096] Step 2: Construct preliminary embedding. Use the knowledge graph training set to generate 272,115 groups of positive triples. For each positive triple, change its correct head entity or tail entity to obtain 256 negative triples, that is, incorrect head entities or tail entities. Similarly, the validation set and test set are respectively divided into 17,535 groups of triples and 20,466 groups of triples. During training, for positive triples, the input sizes of the head entity, relationship, and tail entity are [1024, 1, 1000], where 1024 is the batch size and 1000 is the dimension size. For negative triples, the input size of the head entity or tail entity is [1024, 256, 1000], and 256 is the number of negative triples corresponding to each positive triple.
[0097] After that, use pytorch1.4.0 to construct a new hypercomplex embedding linear layer HyperE. As Figure 2As shown, the hypercomplex technology is used here to decompose the parameter matrix U into Hyper_rule and W. It should be noted that is a three-dimensional tensor, After that, the Kronecker product of W and Hyper_rule is performed to obtain the Kronecker product The summation operation is performed on the first dimension to obtain After that, the sum (h+r) of the head entity and the relationship embedding result obtained in the first step and the tail entity embedding result t are input into the linear layer. The calculation form of the linear layer HyperE is: y = HyperE(x) = Ux + b.
[0098] Taking the positive triple as an example, after passing through the embedding layer, the encoding result of a certain triple is obtained:
[0099] Head entity: / m / 027rn:
[0100] h = [2.85*10 -3 , 2.35*10 -2 Z, 2.57*10 -2 , ……, 1.28*10 -3
[0101] Relationship: / location / country / form_of_government:
[0102] r = [-3.57*10 -2 , -1.47*10 -2 , 2.47*10 -2 , ……, -3.43*10 -2
[0103] Tail entity: / m / 06cx9:
[0104] t = [-1.42*10 -2 , 6.13*10 -3 , 9.39*10 -3 , ……, -6.48*10 -3
[0105] Among them, the head and tail entities represent the address in the knowledge graph.
[0106] Let y1 be the sum of the head entity h and the relationship r:
[0107] y1 = h + r
[0108] And y2 be the tail entity t:
[0109] y2 = t
[0110] y1 and y2 are processed through the HyperE layer to obtain:
[0111] y′1 = GyperE(y1), y′2 = HyperE(y2)
[0112] y′1 = [-1.99*10 -2 , -3.66*10 -2 , 2.45*10 -2 , ……, -1.21*10 -2
[0113] y′2 = [-2.17*10 -2 , -5.98*10 -2 , 1.97*10 -2 , ……, 1.13*10 -2
[0114] Define the score obtained through the distance function as:
[0115] d r (h, t) = ||y′1 - y′2|| = ||HyperE(h + r) - HyperE(t)||
[0116] Construct a deep learning network. Stack different layers of the network in the manner shown in Table 2. The number of layers required for knowledge graph embedding is relatively small, which is for the consideration of streamlining the network. During training, the Adam optimization method is adopted. The initial learning rate is set to 0.0001, the batchSize is set to 1024, the weight coefficient adopts the L2 regularization method, the coefficient is set to 0.0001, the beta1 parameter of adam is set to 0.99, the beta2 parameter is set to 0.999, epsilon is set to 1e-8, and the learning rate decay rate is set to 0.001. A total of 150,000 epochs of training are required.
[0117] Table 2 Knowledge Graph Representation Network Structure Based on Hypercomplex Embedding of Arbitrary Dimensions
[0118]
[0119] Step 3: Use the FB15k-237 validation set for verification and adjust to the best network parameters.
[0120] Step 4: Use the FB15k-237 test set for testing. Keep the embedding results of the network obtained after training. Calculate evaluation metrics such as MR, MRR, and HIT10 for the ranking of the embedding results of the head entity, relation, and tail entity after passing through the model.
[0121] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A knowledge graph representation method based on arbitrary-dimensional hypercomplex embedding, characterized in that, The method includes the following steps: Step 1, preprocessing of knowledge graph data, preprocessing the traditional knowledge graph into structured data according to the model requirements; Step 2, constructing a preliminary embedding using the deep learning framework pytorch, and building a new linear layer, namely a hypercomplex embedding linear layer, to learn the vector representations of entities and relationships on the graph; Step 3, validating with the knowledge graph validation set and adjusting to the optimal network parameters; Step 4, testing the knowledge graph test set, counting the test results, and evaluating the model using evaluation metrics such as MR (Mean Rank), MRR (Mean Reciprocal Ranking), and HIT10 (average proportion of triples with a rank less than 10 in link prediction); Among them, Step 1 is specifically as follows: First, preprocess the knowledge graphs in different fields into five files. The processed files include the knowledge graph triple training set, the knowledge graph triple validation set, the knowledge graph triple test set, the entity ID set, and the relationship ID set; Step 2 is specifically as follows: First, embed the entities and relationships in the knowledge graph obtained in Step 1 into initial vectors to prepare for the next training. Build a hypercomplex embedding linear layer, namely the HyperE layer, and obtain the initial embedding result I of the n - ary number from the input. n is the set number of arities, and the dimension of I is divisible by n I = [I1, I2, I3, …, I n (1) I1 represents the real part of the n - ary number embedding, and I i , i ∈ [2, 3, …, n] represents the imaginary part of the n - ary number embedding. These n parts are concatenated along a given axis to form a vector I, which serves as the input to the HyperE layer. The HyperE layer has the same form as the standard translation model: y = HyperE(x) = Ux + b. The key idea is to construct U as a parameter matrix through the sum of Kronecker products, where x is the input embedding vector to be trained, b is the bias, and y is the embedding vector of the entity or relation; Calculate the scores of positive and negative samples, and calculate the loss of each batch of data through the scores for iterative optimization.
2. The knowledge graph embedding method based on hypercomplex embedding according to claim 1, wherein The preprocessing operation described in Step 1 is specifically: Randomly divide the entire knowledge graph triple data set into a training set, a validation set, and a test set according to 8:1:1, and output the corresponding entity IDs and relationship IDs of the entire knowledge graph.
3. The method for representing a knowledge graph based on arbitrary - dimensional hyper - complex embedding according to claim 2, wherein, The operation of building the hypercomplex embedding linear layer HyperE in Step 2 is specifically: Obtain entity embeddings and relationship embeddings in the knowledge graph. For any triple with head entity h, relationship r, and tail entity t, next, the HyperE layer converts the entity and relationship embeddings into high-order embeddings, y = HyperE(x) = Ux + b. Different learning matrices U are constructed according to different arities by means of the Kronecker product. The Kronecker product generalizes the outer product of vectors to matrices. Let X ∈ R m*n , Y ∈ R p*q , and the Kronecker product is: where x ij =(X) i,j , let n be the dimension of the hypercomplex embedding HyperE, and k be a user-defined hyperparameter representing the dimensions of entity and relation embeddings. The learning matrix U described above is obtained from n Kronecker products: where C i ∈R n*n represents the contribution matrix, represents the component weight matrix, and at the same time k also represents the input and output size of the linear transformation. The contribution matrix C i is selected as a full-rank matrix, whose rows and columns are linearly independent, and all elements belong to {-1, 0, 1}. Let The diagonal alternates between 1 and -1, and each contribution matrix C i is initialized as the matrix which is the product with powers of the cyclic permutation matrix P n The cyclic permutation matrix P n has the effect of shifting the columns to the right. Where when j - 1 = 1 and i = n, j = 1, (P n ) i,j = 1, and all other terms are 0; When n = 2, set When n = 4, set 4. The knowledge graph representation method based on arbitrary-dimensional hypercomplex embedding according to claim 3, wherein The operation of calculating the triple score through the scoring function in Step 2 is specifically: Let y1 be obtained by adding the head entity h and the relationship r: y1 = h + r (6) And y2 be the tail entity t: y2 = t (7) Pass y1 and y2 through the HyperE layer to get: y′1 = HyperE(y1), y′2 = HyperE(y2) (8) Define the score obtained through the distance function as: d r (h, t) = ||y′1 - y′2|| = ||HyperE(h + r) - HyperE(t)|| (9).
5. A method for representing a knowledge graph based on embedding of hypercomplex numbers in any dimension according to claim 1, characterized in that, The operation of learning the vector representations of entities and relationships through the loss function in Step 2 is specifically: Negative samples are very effective for learning knowledge graph embeddings and word embeddings. Use a loss function similar to the negative sampling loss to optimize the distance - based model, that is, the self - adversarial negative sampling method. Sample negative triples according to the current embedding model. Let: where α is the sampling parameter, (h i , r i , t i ) is a positive triple, and (h′ i , r, t′ i ) is a negative triple; The final sampling loss of self - adversarial training is: Where γ is a fixed margin and σ is the sigmoid function.
Citation Information
Patent Citations
Knowledge graph embedding model training method and system and electronic equipment
CN112182245A
Time sequence knowledge graph completion method and device based on QR decomposition and electronic equipment
CN114691890A