Knowledge graph representation learning method based on shared encoder
By introducing a shared Transformer encoder and position embedding matrix in knowledge graph representation learning, the limitations of existing methods in dealing with complex relationships and deep interactions are solved, and efficient knowledge graph representation learning is achieved, which improves the accuracy of ternary combination rational prediction and the scalability of the model.
Patent Information
- Application Number
- CN202510233393.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
AI Technical Summary
Existing knowledge graph representation learning methods have limitations in dealing with complex relationships and capturing deep interactions between entities, and Transformer-based models have problems with computing efficiency and scalability.
A knowledge graph representation learning method based on shared encoder is proposed, the entity-relationship pair is encoded through a shared Transformer encoder, and the sequence position information is retained through the position embedding matrix, the feature interaction is performed using the Hadamard product, and the rationality score is calculated through linear transformation.
Excellent performance is achieved in a lower embedding dimension, reducing computation and storage overhead, improving model scalability, and suitable for representation learning and reasoning tasks of large-scale knowledge graphs.
Smart Images

Figure CN120069039A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of knowledge graph representation learning, and particularly relates to a knowledge graph representation learning method based on a shared encoder. Background Art
[0002] A knowledge graph is an effective way of knowledge representation and is widely applied in multiple fields. The information in the knowledge graph is huge and usually implicit or deep, so it is challenging to directly utilize this information to obtain valuable knowledge. To address this issue, research on knowledge representation learning methods such as knowledge graph embedding has received extensive attention. Knowledge graph embedding aims to map symbolic entities and relationships into a low-dimensional dense vector space, thereby facilitating subsequent calculations and applications in tasks such as knowledge graph completion and triple classification.
[0003] Traditional knowledge graph representation learning methods such as TransE, DistMult, etc., map entities and relationships into a low-dimensional vector space and use specific scoring functions to evaluate the rationality of triples. However, these methods have certain limitations in dealing with complex relationships and capturing deep interactions between entities.
[0004] Transformer-based models have made significant progress in processing high-dimensional embedding representations and have achieved better performance compared to traditional models. However, they still face problems of scalability and computational efficiency. Most existing Transformer models use multiple encoder layers to enhance the model's expressive power, and each encoder layer uses the multi-head self-attention mechanism to capture the complex dependencies between entities and relationships. However, although this design has greatly improved accuracy, it increases redundant free parameters, resulting in a significant increase in computational time and memory consumption. Especially when dealing with large-scale data, the storage and training efficiency of the model are significantly reduced. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a knowledge graph representation learning method based on a shared encoder, including the following steps:
[0006] S1: Embed entities and relationships in the knowledge graph through an embedding method, and map the entity embeddings and relationship embeddings into a low-dimensional vector space;
[0007] S2: Construct a position embedding matrix P, denoted as where L is the sequence length and D is the embedding dimension.
[0008] S3: Encode the head entity-relationship pair using a shared Transformer encoder, and encode the tail entity-relationship pair using a shared Transformer encoder;
[0009] S4: Perform pooling operations on the outputs of different encoded entity-relationship pairs to obtain different global representations respectively;
[0010] S5: Perform feature interaction on the global representations of different entity-relationship pairs to obtain a comprehensive feature representation;
[0011] S6: Calculate the rationality score of the comprehensive feature representation through a linear transformation;
[0012] S7: Use the binary cross-entropy loss function for training to optimize the model parameters.
[0013] Furthermore, the specific steps of step 1 are as follows:
[0014] Assume that the knowledge graph contains N entities and M relationships, and use randomly initialized embedding matrices and where d is the embedding dimension. For the head entity h, relationship r, and tail entity t of the triple, extract their corresponding embedding vectors e h , e r and e t , that is, e h = E entity [h], e r = E relation [r], e t = E entiyy [t] to obtain the low-dimensional vector representations of entities and relationships.
[0015] Furthermore, the construction of the position embedding matrix P is as follows:
[0016] Each row of the position embedding matrix P corresponds to a position in the sequence, and its initialization method adopts the Xavier uniform distribution:
[0017]
[0018] where, P ij is the element in the i-th row and j-th column of the position embedding matrix P, and XavierUniform represents the weight initialization.
[0019] Furthermore, the specific steps of step 3 are as follows:
[0020] First, construct two input sequences: for (h, r), stack the head entity h and the relationship r to form the sequence [h, r]; for (r, t), stack the relationship r and the tail entity t to form the sequence [r, t];
[0021] Then add the corresponding position information to the above input sequences, that is, P hr and P rt, corresponding to the position embeddings of (h, r) and (r, t) respectively;
[0022] Process the two input sequences with position information through a shared Transformer encoder to obtain an encoded output.
[0023] Furthermore, the feature interaction is performed using the Hadamard product.
[0024] Furthermore, the rationality score is calculated by the following formula:
[0025]
[0026] where is the rationality score, W and b are the weights and biases of the linear transformation, σ is the Sigmoid activation function, and v hrt is the comprehensive feature representation;
[0027] Furthermore, the shared Transformer encoder adopts a single-layer Transformer encoding layer design, specifically including a multi-head self-attention mechanism and a feed-forward neural network, where the number of attention heads of the multi-head self-attention mechanism is 32, 64, or 128; the feed-forward neural network performs a non-linear transformation on the output of the multi-head self-attention mechanism.
[0028] Furthermore, the pooling operation includes a sum pooling or an average pooling operation.
[0029] Furthermore, the calculation formula of the binary cross-entropy loss function is as follows:
[0030]
[0031] where y is the true label; during the optimization process, the loss function is minimized to update the parameters of the model, including weights and biases.
[0032] The present invention has the following beneficial effects:
[0033] The present invention proposes a knowledge graph representation learning method based on a shared encoder. Compared with the prior art, this method significantly reduces redundant parameters by introducing a shared Transformer encoder and increasing the number of self-attention heads. By constructing position embeddings to preserve the position information of the sequence, the modeling ability of the model in dealing with the local and global interaction characteristics of the sequence is improved. In addition, the design of separately encoding different entity-relationship pairs further enhances the performance of the model in dealing with different interaction patterns. This method achieves excellent performance under a lower embedding dimension while maintaining low computational and storage overheads, greatly improving the scalability of the model. This method can efficiently learn the representations of entities and relationships in the knowledge graph, improve the accuracy of triple rationality prediction, is applicable to the representation learning and reasoning tasks of large-scale knowledge graphs, and has good generalization ability and practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic flowchart of a knowledge graph representation learning method based on a shared encoder according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] To make the objectives, technical solutions, and beneficial effects of the present invention clearer and more understandable, the following will further describe in detail the specific embodiments of the present invention in conjunction with specific embodiments.
[0036] Referring to Figure 1 , the present invention provides a knowledge graph representation learning method based on a shared encoder, including the following steps:
[0037] S1: Embed the entities and relationships in the knowledge graph through an embedding method, and map the entity embeddings and relationship embeddings into a low-dimensional vector space.
[0038] In this embodiment, a knowledge graph of superalloy materials is taken as the research object, and this knowledge graph contains detailed information of 205 superalloy grades. The knowledge graph contains 69 ontology network triples and 7338 entity network triples, covering the following key features, including physical and chemical parameters, shear modulus, microstructure, preparation process, misfit degree, performance parameters, element composition, and creep temperature, etc.
[0039] First, embed the entities (such as alloy grades, elements, process parameters, etc.) and relationships (such as "contains element", "heat treatment process", etc.) in the knowledge graph of superalloy materials. Assuming that the knowledge graph contains N entities and M relationships, an embedding method is used to map each entity and relationship into a low-dimensional vector. Using a randomly initialized embedding matrix and where d is the embedding dimension (taking the value of 128). For the head entity h, relation r, and tail entity t of these triples, their corresponding embedding vectors e h , e r and e t are extracted from the embedding matrix, that is, e h = E entity [h], e r = E relation [r], e t = E entitu [t] to obtain the low-dimensional vector representations of entities and relations.
[0040] S2: Construct a position embedding matrix P, denoted as where L is the sequence length and D is the embedding dimension.
[0041] In the triple representation of the superalloy material knowledge graph, the order of entities and relations is important. To maintain the order relationship of information such as element composition and process parameters in the superalloy material knowledge graph, a position embedding matrix P is constructed. Its dimension is L×D, where L represents the sequence length. For example, in a triple sequence, the sequence length is 3, that is, (h, r, t), and D is the embedding dimension. Each row of this position embedding matrix P corresponds to a position in the sequence, and its initialization method uses the Xavier uniform distribution to ensure the rationality and stability of the embedding representation:
[0042]
[0043] where, P ij is the element in the i-th row and j-th column of the position embedding matrix P, and XavierUniform represents weight initialization. The position embedding matrix is used to add position encoding to each element in the triple sequence of the superalloy knowledge graph, and the order information in the sequence is retained through position embedding. The position embedding matrix P is a trainable parameter matrix obtained through learning to enhance the model's understanding of the position order and capture sequence dependencies.
[0044] S3: Use a shared encoder to encode different entity-relation pairs respectively.
[0045] Use a shared Transformer encoder to encode different entity-relation pairs respectively. For example, for the triple (GH4169, contains element, Ni), use the shared Transformer encoder to encode different entity-relation pairs, and encode (h, r) and (r, t) respectively. First, construct two input sequences: for (h, r), stack the head entity h and the relation r to form the sequence [h, r]. For (r, t), stack the relation r and the tail entity t to form the sequence [r, t].
[0046] Add the corresponding position information to these input sequences, namely P hr and P rt , corresponding to the position embeddings of (h, r) and (r, t) respectively. These sequences with position information are input into a shared Transformer encoder for processing. The shared encoder adopts a single-layer Transformer encoding layer design, and its core structure consists of a multi-head self-attention mechanism and a feed-forward neural network. Different from the traditional multi-layer encoder where each layer only contains 8 attention heads, the present invention innovatively adopts a single-layer structure and increases the number of attention heads to 64. This "single-layer multi-head" design avoids the parameter redundancy of multi-layer stacking by adding a large number of attention heads in a single layer, while maintaining a strong feature extraction ability. The 64 attention heads are responsible for capturing the feature relationships in different dimensions through parallel computing, realizing fine-grained feature processing. This design not only reduces the storage overhead but also improves the processing efficiency by reducing the inter-layer propagation, making the model have better scalability when processing large-scale knowledge graphs. The feed-forward neural network is located after the self-attention layer and performs a non-linear transformation on the output of the attention mechanism to further enhance the expressive ability of the model. The shared encoder uses the same model structure for different interaction calculations, reducing the number of model parameters while improving the learning efficiency and performance of the model.
[0047] Specifically, for the [h, r] sequence, the input representation is:
[0048] x hr = Stack(e h , e r )
[0049] For the [r, t] sequence, the input representation is:
[0050] x rt = Stack(e r , e t )
[0051] Among them, the Stack stacking operation is a method of merging two embedding vectors with the same feature dimension along the batch dimension. Specifically, given the head entity h, the relation r, and the tail entity t, assuming their embedding vectors have shapes (1024, 128) respectively, where 1024 represents the batch size and 128 represents the embedding dimension. The stacking operation merges these vectors along the batch dimension to form a new tensor with a shape of (1024, 2, 128).
[0052] Then, add the corresponding position information to the input sequence:
[0053] x hr ' = x hr + Phr , x rt ′ = x rt + P rt
[0054] These position - embedded input triple sequences will be input into the shared Transformer encoder to obtain its encoded output:
[0055] z hr = Transformer(x ′ hr ), z rt = Transformer(x ′ ri )
[0056] S4: Perform pooling operations on the encoded outputs of different entity - relation pairs to obtain different global representations respectively.
[0057] After encoding, the Transformer encoder outputs a vector sequence containing all element information of the sequence (GH4169, contains element, Ni). To obtain a fixed - dimensional representation of this sequence, a pooling operation is adopted. Using sum pooling or mean pooling operations, the output representation of the Transformer encoder is compressed into a fixed - dimensional representation vector for subsequent scoring calculations. Specifically, for the sequences [h, r] and [r, t] of the triple (GH4169, contains element, Ni), we obtain two feature vectors as their global representations through sum pooling operations respectively:
[0058]
[0059] where z jr,i and z rt,i are the output vectors at the i - th position in the sequences [h, r] and [r, t] respectively.
[0060] S5: Perform feature interaction on the feature representations of different entity - relation pairs to obtain a comprehensive feature representation.
[0061] Perform interaction on the two feature vectors v hr and v rt obtained by pooling to obtain a comprehensive feature representation. Use Hadamard product (element - wise multiplication) for feature interaction:
[0062] v hrt = v hr × v rt
[0063] This method combines the interaction information between the head entity and the relation with the interaction information between the relation and the tail entity to obtain the comprehensive feature representation v hrt , and this feature representation contains all the interaction information in the triple (GH4169, contains element, Ni).
[0064] S6: Calculate the rationality score of the triple through linear transformation.
[0065] Input the comprehensive feature representation v hrt into a linear layer, and calculate the rationality score of the triple (GH4169, contains element, Ni) by performing a linear transformation on it. The rationality score is calculated by the following formula:
[0066]
[0067] where W and b are the weights and biases of the linear transformation, and σ is the Sigmoid activation function, which is used to map the score to the interval [0, 1] to represent the rationality probability of the triple; the pooled representation is mapped to the score space through linear transformation, and the score value is used to judge the rationality of the triple.
[0068] S7: Use the binary cross-entropy loss function for training to optimize the model parameters.
[0069] Use known correct triples (such as verified alloy compositions, process parameters, etc.) as positive samples, and use randomly generated incorrect triples (such as unreasonable heat treatment temperatures, incorrect element ratios, etc.) as negative samples, and train through the binary cross-entropy loss function. The calculation formula of the loss function is:
[0070]
[0071] where y is the true label, indicating whether the triple is correct. By minimizing the loss function, the parameters of the model can be updated, thereby improving the accuracy of triple rationality prediction.
[0072] During the optimization process, we use gradient descent or other optimization algorithms (such as Adam) to minimize the loss function and update the weights W and biases b of the model, ultimately enabling the model to accurately predict the rationality of the triple.
[0073] It is understood that the present invention is described by way of some embodiments, and those skilled in the art will know that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. A knowledge graph representation learning method based on a shared encoder, characterized in that: The following steps are involved: S1: Embed entities and relations in the knowledge graph through embedding methods, and map entity embedding and relationship embedding into a low-dimensional vector space; S2: Construct the position embedding matrix P, denoted as Where L is the sequence length and D is the embedding dimension. S3: Use a shared Transformer encoder to encode the head entity-relation pair, and use a shared Transformer encoder to encode the tail entity-relation pair; S4: Pooling operations are performed on the outputs of different entity-relation pairs after encoding to obtain different global representations; S5: Perform feature interactions on the global representations of different entity-relation pairs to obtain a comprehensive feature representation; S6: Calculate the rationality score of the comprehensive feature representation through linear transformation; S7: Use the binary cross entropy loss function for training and optimize the model parameters.
2. According to the knowledge graph representation learning method based on shared encoder according to claim 1, it is characterized in that: The step 1 is specifically as follows: Assume that the knowledge graph contains N entities and M relations, using a randomly initialized embedding matrix and Where d is the embedding dimension. For the head entity h, relation r, and tail entity t of the triple, extract its corresponding embedding vector e from the embedding matrix. h 、e r and e t , that is, e h =E entity [h],e r =E relation [r],e t =E entity [t] Obtain low-dimensional vector representations of entities and relations.
3. A knowledge graph representation learning method based on a shared encoder according to claim 2, characterized in that: The position embedding matrix P is constructed as follows: Each row of the position embedding matrix P corresponds to a position in the sequence, and its initialization method adopts Xavier uniform distribution: Among them, P ij It is the element in the i-th row and j-th column of the position embedding matrix P, and XavierUniform represents weight initialization.
4. A knowledge graph representation learning method based on a shared encoder according to claim 3, characterized in that: The step 3 is as follows: First, construct two input sequences: for (h, r), stack the head entity h and the relation r to form the sequence [h, r]; for (r, t), stack the relation r and the tail entity t to form the sequence [r, t]; Then add the corresponding position information to the above input sequence, that is, P hr and P rt , corresponding to the position embedding of (h,r) and (r,t) respectively; The two input sequences with position information are processed into a shared Transformer encoder to obtain the encoded output.
5. A knowledge graph representation learning method based on a shared encoder according to claim 4, characterized in that: The feature interaction is performed using Hadamard product.
6. A knowledge graph representation learning method based on a shared encoder according to claim 5, characterized in that: The plausibility score is calculated by the following formula: in, is the rationality score, W and b are the weight and bias of the linear transformation, σ is the Sigmoid activation function, and v hrt It is a comprehensive feature representation.
7. A knowledge graph representation learning method based on a shared encoder according to claim 6, characterized in that: The shared Transformer encoder adopts a single-layer Transformer encoding layer design, specifically including a multi-head self-attention mechanism and a feedforward neural network, wherein the number of attention heads of the multi-head self-attention mechanism is 32, 64 or 128; the feedforward neural network performs a nonlinear transformation on the output of the multi-head self-attention mechanism.
8. A knowledge graph representation learning method based on a shared encoder according to claim 7, characterized in that: The pooling operation includes a sum pooling operation or an average pooling operation.
9. A knowledge graph representation learning method based on a shared encoder according to claim 7, characterized in that: The calculation formula of the binary cross entropy loss function is as follows: Among them, y is the true label; in the optimization process, the loss function is minimized and the parameters of the model, including weights and biases, are updated.