Knowledge graph completion model based on relation awareness common attention
By introducing a relationship-perceived co-attention mechanism and a co-attention relationship convolution decoder into the knowledge graph completion model, the shortcomings of existing methods in capturing complex nonlinear interactions and global semantic associations are solved, and a more efficient and accurate knowledge graph completion effect is achieved.
Patent Information
- Application Number
- CN202510629472.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing knowledge graph completion method has shortcomings in capturing complex nonlinear interactions between entities and relationships, utilizing graph structure information and multimodal additional information, reducing the risk of overfitting, and capturing global semantic associations across triples.
A knowledge graph completion model based on relationship perception co-attention is proposed. Through the relationship perception attention network encoder and the co-attention relationship convolution decoder, combined with the graph convolution network and convolution operation, it captures the complex nonlinear interaction between entities and relationships, and captures the global feature interaction through the entity relationship co-attention mechanism.
Effectively break through the linear interaction assumption of traditional embedding methods, enhance the relationship directional modeling ability of graph structure methods, reduce the risk of overfitting, realize the accurate alignment and fusion of multimodal information, capture the semantic associations across triples, and improve the performance and practicality of knowledge graph completion.
Smart Images

Figure CN120144835A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge graph completion, and specifically to a knowledge graph completion model based on relational perception co-attention. Background Art
[0002] Existing knowledge graph completion methods mainly include four categories: traditional embedding methods, such as TransE, RotatE, etc., which model entity relationships through low-dimensional space translation or complex space rotation, or linear combination models based on semantic matching such as RESCAL, ComplEx, etc.; graph structure methods, such as R-GCN, GraIL, etc., which use graph neural networks to aggregate neighborhood information or extract subgraph structure features; additional information methods, such as LiteralE, IKRL, etc., which fuse multi-modal data such as text and images to enhance entity representations; entity-relationship interaction methods, such as ConvE, InteractE, etc., which capture local interaction features between entities and relationships within triples through convolution or reshaping strategies.
[0003] The existing technologies have the following main problems and defects: traditional embedding methods (such as TransE, RotatE, RESCAL, etc.) rely on linear transformation assumptions or shallow linear combinations, making it difficult to capture complex non-linear interactions between entities and relationships, and generally ignoring the graph structure information and multi-modal additional information of the knowledge graph, resulting in limited representation capabilities; graph structure methods (such as R-GCN, GraIL, etc.) can aggregate neighborhood information or extract subgraph features through graph neural networks, but they utilize relationship directionality information insufficiently, and graph neural networks are prone to overfitting due to parameter redundancy, affecting the generalization performance of the model; additional information methods (such as LiteralE, IKRL, etc.) rely on external data (such as text, images) for pre-training, increasing the dependence of the model on additional data sources. At the same time, there are semantic alignment deviation problems in the multi-modal information fusion mechanism, making it difficult to effectively integrate heterogeneous data; entity-relationship interaction methods (such as ConvE, InteractE, etc.) mainly focus on local interactions within triples, unable to capture global semantic associations across triples, and the interaction granularity and the number of interactions are limited, and the graph structure information and multi-modal additional information are not effectively combined, resulting in insufficient comprehensiveness and accuracy of interaction modeling. These problems limit the performance and practicality of existing methods in complex knowledge graph completion tasks. Summary of the Invention
[0004] The purpose of this part is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this part, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this part, the abstract, and the title. However, such simplifications or omissions shall not be used to limit the scope of the present invention.
[0005] Therefore, the objective of the present invention is to provide a knowledge graph completion model based on relation-aware co-attention, which breaks through the dependence of traditional embedding methods on the linear interaction hypothesis, captures the complex non-linear interaction between entities and relations, enhances the modeling ability of the graph structure method for relation directionality, and reduces the risk of overfitting.
[0006] To solve the above technical problems, according to one aspect of the present invention, the following technical solutions are provided: A knowledge graph completion model based on relation-aware co-attention, comprising: A relation-aware attention network encoder, which is used to utilize the directed graph structure information and multi-relation type information of the multi-relation knowledge graph, aggregate the entity information and attribute information of the nodes connected by different relations in the knowledge graph to obtain entity embeddings, and increase the difference between the head entity and the tail entity in the entity embedding representation through a residual connection layer; A co-attention relation convolution decoder, which receives the output vector from the relation-aware attention network encoder and another set of relation vectors as inputs, captures the local feature interaction between entities and relations through relation convolution, and captures the global feature interaction by introducing an entity-relation co-attention mechanism, so as to maximize the number of feature interactions at different granularities between entities and relations.
[0007] As a preferred solution of a knowledge graph completion model based on relation-aware co-attention according to the present invention, wherein the relation-aware attention network encoder includes a relation specialization component and an attention feature integration component; The relation specialization component defines the propagation model as:
[0008] Wherein, represents the output feature of node at the th layer, represents the non-linear activation function, represents the weight matrix corresponding to relation in the th layer, represents the output feature of neighbor node at the th layer, represents the weight matrix of the node's own features, represents node at the th layer of the original feature, represents relation under node 's neighbor index set, is a problem-specific normalization constant that can be learned or selected in advance, is the relation The corresponding adjacency matrix is the node representation matrix of the layer, and is the weight matrix of the relationship The attention feature integration component performs self-attention mechanism calculation on the nodes to calculate the attention coefficients ; where represents the importance degree of node to node , represents the self-attention mechanism function, represents the shared weight matrix, represents the feature vector of node , represents the feature vector of node ; Among them, to make the coefficients easy to compare among different neighborhood nodes, the softmax function is used to normalize it, and the formula is as follows: ; where exp represents the exponential operation, represents a certain neighborhood of node , represents the attention score of node on the th neighborhood.
[0009] As a preferred solution of the knowledge graph completion model based on relationship-aware co-attention described in the present invention, among them, the network layer of the relationship-aware attention network encoder is expressed as:
[0010]
[0011] where represents the weight vector, represents the feature vector on the th neighborhood, T represents the transpose, and || is the concatenation operation.
[0012] As a preferred solution of the knowledge graph completion model based on relationship-aware co-attention described in the present invention, among them, in the co-attention relationship convolution decoder, the received entity vector is reshaped into a 2D convolution matrix, the relationship vector is further segmented into block structures of the same size, each block is reshaped into the shape of a 2D convolution kernel, and the convolution kernel weight coefficient is calculated by the co-attention mechanism, and ; where is a vector reshaping function, represents a relationship vector, and each convolutional kernel .
[0013] As a preferred solution of a knowledge graph completion model based on relationship-aware co-attention according to the present invention, for each relationship convolutional kernel , after performing a convolution operation, a convolution feature map is generated, where and represent the length and width of the entity vector matrix respectively, and represent the length and width of the convolutional kernel respectively, and the calculation formula is as follows: ; where represents a convolution operator, represents a two-dimensional matrix after reshaping the head entity embedding vector, is an activation function.
[0014] As a preferred solution of a knowledge graph completion model based on relationship-aware co-attention according to the present invention, after calculating the convolution feature maps corresponding to each relationship convolutional kernel, the convolution feature maps are stacked and tiled to form a long vector : ; where represents a vector flattening function, represents a vector concatenation function, and then this vector passes through a fully connected layer to obtain a vector .
[0015] As a preferred solution of a knowledge graph completion model based on relationship-aware co-attention according to the present invention, in the entity-relationship co-attention mechanism of the co-attention relationship convolutional decoder, given the two-dimensional vector representation of the relationship and the two-dimensional vector representation of the entity, the affinity matrix between the entity and the relationship is calculated as: ; where represents a non-linear activation function, represents a weight.
[0016] As a preferred solution of a knowledge graph completion model based on relationship-aware co-attention according to the present invention, after calculating the affinity matrix, the attention maps of the predicted relationship and entity are learned in the following way:
[0017] Among them, represents the relational attention map, represents the entity attention map, , is the weight parameter, and are the attention probabilities of each relationship and entity feature respectively; Based on the above attention weights, the relational attention vector is calculated as the weighted sum of relational features:
[0018] Among them, represents the weighted relational feature, represents the attention weight of the relational vector, represents the relational vector. After the obtained relational attention vector is normalized by softmax, each value in the vector corresponds to the weight of each relational convolution kernel, that is: ; Among them, represents the weight of each relational convolution kernel, each element value in the relational attention vector.
[0019] As a preferred solution of the knowledge graph completion model based on relational perception co-attention described in the present invention, among them, the knowledge graph completion model is trained using a cross-entropy loss function, and the loss function is: ; Among them represents the number of candidate entities, represents the triple score, is the binary label value. If is a positive triple, its value is 1, otherwise the value is 0.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows: The proposed Relational-Aware Co-Attention Convolutional Network model (RACOCN) effectively overcomes the defects of the existing technologies through a relational-aware attention mechanism and a co-attention relational convolution decoder, combining graph convolutional networks and convolutional operations, in the following ways: breaking through the dependence of traditional embedding methods on the linear interaction hypothesis and capturing the complex non-linear interactions between entities and relationships; enhancing the modeling ability of graph structure methods for relationship directionality and reducing the risk of overfitting; avoiding the strong dependence of additional information methods on external data and achieving precise alignment and fusion of multi-modal information; and transcending the local interaction limitations of existing interaction methods to capture semantic associations across triples through a global context awareness mechanism. It achieves significant advantages in terms of feature interaction ability, comprehensiveness of information fusion, computational efficiency, and application universality, providing efficient and precise technical support for knowledge graph completion and related fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below with reference to the drawings and specific embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts. Among them: Figure 1 It is a schematic structural diagram of a knowledge graph completion model based on relational-aware co-attention of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] To make the above objects, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the drawings.
[0023] The present invention provides a knowledge graph completion model based on relational-aware co-attention, which breaks through the dependence of traditional embedding methods on the linear interaction hypothesis, captures the complex non-linear interactions between entities and relationships, enhances the modeling ability of graph structure methods for relationship directionality, and reduces the risk of overfitting.
[0024] Such as Figure 1As shown in the figure, the invention model consists of a relation-aware attention network encoder and a co-attention relation convolution decoder, which combines the advantages of graph convolutional neural networks and convolutional operations. Among them, the encoder makes full use of the directed graph structure information and multi-relation type information of the multi-relation knowledge graph, represents entity embeddings by aggregating the entity information and attribute information of the nodes connected by different relations in the knowledge graph, and uses a residual connection layer to increase the difference between the head entity and the tail entity in the entity embedding representation. The co-attention relation convolution decoder captures the feature information interaction at different granularities between entities and relations, captures local feature interaction through the way of relation convolution, and captures global feature interaction by introducing an entity-relation co-attention mechanism, so as to maximize the number of feature interactions at different granularities between the two.
[0025] The relation-aware attention network encoder consists of a relation specialization component and an attention feature integration component. The following propagation model is defined in the relation specialization component:
[0026] Where represents the output feature of node at the -th layer, represents a non-linear activation function, represents the weight matrix corresponding to relation in the -th layer, represents the output feature of neighbor node at the -th layer, represents the weight matrix of the node's own feature, represents node at the -th layer's original feature, represents the neighbor index set of node under relation , is a problem-specific normalization constant that can be learned or selected in advance, is the adjacency matrix corresponding to relation , is the node representation matrix at the -th layer, is the weight matrix of relation .
[0027] In the attention feature integration component, a self-attention mechanism is executed on the nodes to calculate the attention coefficient:
[0028] Where represents node For the node , the importance is represented by the self-attention mechanism function , the shared weight matrix is represented by the eigenvector of the node is represented by the eigenvector of the node . And in order to make the coefficients easy to compare among different neighboring nodes, the softmax function is used to normalize them among all the node
[0029] choices: where represents a certain neighborhood of the node is represented by at the th neighborhood, and the attention score
[0030] In this component, the attention mechanism is a single-layer feed-forward neural network, parameterized by the weight vector , and applying the LeakyReLU non-linear activation function. After fully expanding it, the coefficients calculated by the attention mechanism can be expressed as:
[0031] where represents the weight vector, represents the eigenvector at the th neighborhood, T represents transpose, and || is the concatenation operation.
[0032] By specializing the connection between the component and the attention feature integration component through the above relationship, a network layer of the relation-aware attention network is combined, and this network layer can be expressed as:
[0033] where the attention set coefficient is:
[0034] For the decoder, the co-attention relation convolution receives the output vector from the encoder and another set of relation vectors as inputs. The entity vector received by the decoder is reshaped into a 2D convolution matrix (where ), and the relation vector is further segmented into block structures of the same size , c is the number of convolution kernels, and each block is reshaped into the shape of a 2D convolution kernel.
[0035]
[0036] where represents the weighted convolution kernel, is the weight coefficient of each convolution kernel, which is calculated by the co-attention mechanism, represents different relational convolution kernels, which can be specifically expressed as:
[0037] where is the vector reshaping function, represents the relational vector, and each convolution kernel , for each relational convolution kernel , after performing the convolution operation, a convolution feature map will be generated , where, and represent the length and width of the entity vector matrix respectively, and represent the length and width of the convolution kernel respectively, and their calculation methods are shown in the following formula:
[0038] where represents the convolution operator, represents the two-dimensional matrix after reshaping the head entity embedding vector, is the activation function, such as the ReLU function.
[0039] After calculating the convolution feature maps corresponding to each relational convolution kernel, we stack and tile the convolution feature maps to form a long vector :
[0040] where represents the vector flattening function, represents the vector concatenation function. Then this vector passes through the fully connected layer to obtain the vector .
[0041] For the entity-relation co-attention mechanism in the decoder, given the two-dimensional vector representation of the relation and the two-dimensional vector representation of the entity, the affinity matrix between the entity and the relation is calculated as follows:
[0042] where represents the non-linear activation function, Denote the weights. After calculating the affinity matrix, regard the affinity matrix as a feature, and learn to predict the attention maps of relationships and entities in the following way:
[0043] where denotes the relationship attention map, denotes the entity attention map, , is the weight parameter, and are the attention probabilities of each relationship and entity feature respectively, and the affinity matrix converts the relationship attention space to the entity attention space ( vice versa). Based on the above attention weights, calculate the relationship attention vector as the weighted sum of relationship features:
[0044] where denotes the weighted relationship feature, denotes the relationship vector attention weight, denotes the relationship vector. After the obtained relationship attention vector is normalized by softmax, each value in the vector corresponds to the weight of each relationship convolution kernel, that is:
[0045] where, denotes the weight of each relationship convolution kernel, each element value in the relationship attention vector.
[0046] For model training, the model adopts the following cross-entropy loss function: .
[0047] where denotes the number of candidate entities, denotes the triple score, is the binary label value. If is a positive triple, its value is 1, otherwise the value is 0.
[0048] Although the present invention has been described above with reference to the embodiments, various modifications can be made thereto and components thereof can be replaced with equivalents without departing from the scope of the present invention. In particular, as long as there is no structural conflict, the features in the embodiments disclosed in the present invention can be combined with each other in any way, and the exhaustive description of these combinations is not given in this specification only for the sake of saving space and resources. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A knowledge graph completion model based on relation-aware co-attention, characterized in that: include: The relation-aware attention network encoder is used to utilize the directed graph structure information and multi-relation type information of the multi-relation knowledge graph to aggregate the entity information and attribute information of the nodes connected by different relations in the knowledge graph to obtain entity embedding, and increase the difference between the head entity and the tail entity in the entity embedding representation through the residual connection layer; The co-attention relation convolution decoder receives the output vector from the relation-aware attention network encoder and another set of relation vectors as input, captures the local feature interactions between entities and relations through relation convolution, and captures the global feature interactions by introducing the entity-relation co-attention mechanism, thereby maximizing the number of feature interactions of different granularities between entities and relations.
2. According to claim 1, a knowledge graph completion model based on relationship-aware co-attention is characterized in that: The relation-aware attention network encoder includes a relation-specialization component and an attention feature integration component; The relationship specialization component defines the propagation model as: ; in, Representation Node In the The output features of the layer, represents a nonlinear activation function, Indicates Layer-by-layer relations The corresponding weight matrix is, Represents neighbor nodes In the The output features of the layer, The weight matrix representing the node’s own features, Representation Node In the The original characteristics of the layer, Representing relationships Next Node The neighbor index set of is a problem-specific normalization constant that can be learned or chosen in advance, It's a relationship The corresponding adjacency matrix is, It is The nodes of a layer represent matrices, It's a relationship The weight matrix of The attention feature integration component performs the self-attention mechanism on the node to calculate the attention coefficient ; in Representation Node For Node The importance of represents the self-attention mechanism function, represents the shared weight matrix, Representation Node The characteristic vector of Representation Node The eigenvector of In order to make the coefficients easier to compare between different neighborhood nodes, the softmax function is used to normalize them. The formula is as follows: ; Among them, exp represents exponential operation, Representation Node A neighborhood of Representation Node In the The attention score on the neighborhood.
3. A knowledge graph completion model based on relationship-aware co-attention according to claim 2, characterized in that: The network layer of the relation-aware attention network encoder is represented as: ; ; in, represents the weight vector, Indicated in The eigenvectors in the neighborhood, T represents transposition, and || is a concatenation operation.
4. A knowledge graph completion model based on relationship-aware co-attention according to claim 1, characterized in that: In the co-attention relation convolution decoder, the received entity vector is reshaped into a 2D convolution matrix, and the relation vector is further divided into blocks of the same size. Each block is reshaped into the shape of a 2D convolution kernel, and the convolution kernel weight coefficient is calculated by the co-attention mechanism, and ; in, is a vector reshaping function, Represents the relationship vector, each convolution kernel .
5. A knowledge graph completion model based on relationship-aware co-attention according to claim 1, characterized in that: For each relation convolution kernel , after performing the convolution operation, a convolution feature map is generated ,in, and Respectively represent the length and width of the entity vector matrix, and Respectively represent the length and width of the convolution kernel, and the calculation formula is as follows: ; in, represents the convolution operator, Represents the reshaped two-dimensional matrix of the head entity embedding vector, is the activation function.
6. A knowledge graph completion model based on relationship-aware co-attention according to claim 5, characterized in that: After calculating the convolution feature maps corresponding to each relational convolution kernel, the convolution feature maps Stack and tile to form a long vector : ; in represents the vector flattening function, Represents the vector connection function, and then the vector is passed through the fully connected layer to obtain the vector .
7. A knowledge graph completion model based on relationship-aware co-attention according to claim 1, characterized in that: The entity relation co-attention mechanism in the co-attention relation convolutional decoder, given a two-dimensional vector representation of the relation and a two-dimensional vector representation of the entity , calculate the affinity matrix between entities and relations for: ; in, represents a nonlinear activation function, Represents weight.
8. A knowledge graph completion model based on relationship-aware co-attention according to claim 7, characterized in that: After computing the affinity matrix, the attention map for predicting relations and entities is learned as follows: ; in, represents the relational attention graph, represents the entity attention map, , is the weight parameter, and are the attention probabilities of each relation and entity feature, respectively; Based on the above attention weights, the relation attention vector is calculated as the weighted sum of relation features: ; in, represents the weighted relationship feature, represents the relation vector attention weight, Represents the relationship vector. After the obtained relationship attention vector is normalized by softmax, each value in the vector corresponds to the weight of each relationship convolution kernel, that is: ; in, represents the weight of each relational convolution kernel, The value of each element in the relation attention vector.
9. A knowledge graph completion model based on relationship-aware co-attention according to claim 1, characterized in that: The graph recognition and spectrum completion model is trained using the cross entropy loss function, and the loss function is: ; in represents the number of candidate entities, represents the triple score, is a binary label value, if If it is a positive triple, its value is 1, otherwise it is 0.