Tense induction knowledge reasoning method based on attention mechanism and GNN
By introducing attention mechanism and GNN methods in tense knowledge inference, the problem that the existing technology fails to effectively perform tense knowledge graph inductive reasoning is solved, and better predictive capabilities for new entities and relationships are achieved.
Patent Information
- Application Number
- CN202510036072.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
Existing tense knowledge inference methods fail to effectively consider the inductive reasoning of tense knowledge graphs and cannot predict new entities and relationships that may appear in the future.
The tense induction knowledge inference method based on attention mechanism and GNN is adopted to capture the different degrees of influence of relationships and relationship paths within the path through the bilayer attention mechanism, combine MLP to achieve the fusion of local features and global features, model the semantic features of new entities, and design the scoring function and loss function.
It improves the ability to reason tense inductive knowledge, and can more effectively utilize the structured information and time evolution characteristics of tense knowledge graphs, thereby improving the ability to reason about unseen data.
Smart Images

Figure CN119940546A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of knowledge reasoning, and particularly to a temporal induction knowledge reasoning method based on attention mechanism and GNN. Background Art
[0002] Knowledge reasoning aims to explore the relationships and patterns hidden in data through analysis and modeling of knowledge representation, and to make inferences or predictions based on existing knowledge. Knowledge reasoning is widely used in tasks such as knowledge graph construction, question-answering systems, and recommendation systems, and can improve the machine's ability to understand and solve complex problems. Temporal knowledge reasoning refers to reasoning and analysis based on the temporal knowledge graph to obtain the temporal patterns and trends of temporal knowledge, and to predict future facts based on the historical state of things.
[0003] Most existing temporal knowledge reasoning methods consider extrapolation reasoning, that is, predicting events at future time points based on historical information. The predicted events are entities or relationships that have already appeared in the historical time, and new entities that will appear at future time points are not considered, that is, the inductive reasoning of temporal knowledge graphs is not considered. Temporal knowledge graph inductive reasoning requires not only the ability to predict events at future time points, but also the ability to reason and predict new entities and relationships that may appear in the future. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a temporal inductive knowledge reasoning method based on attention mechanism and GNN. Aiming at the global features of the node, a double-layer attention mechanism is designed to capture the different degrees of influence of the relationship within the path and the relationship path to model the global path features. Aiming at the local features of the node, GNN is introduced to encode the neighbor nodes and neighbor relationships, and the fusion of local features and global features is realized based on MLP.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] A temporal inductive knowledge reasoning method based on attention mechanism and GNN, characterized in that the method comprises the following steps:
[0007] S1: Modeling knowledge graph quadruple global features and node local features;
[0008] S2: Design position encoding based on relative order and time difference;
[0009] S3: Determine the importance of relations based on the intra-path relation attention mechanism;
[0010] S4: Determine path importance based on path attention mechanism;
[0011] S5: Encode local features based on GNN;
[0012] S6: Fusion of global features and local features based on MLP;
[0013] S7: Modeling new entity semantic features;
[0014] S8: Modeling scoring and loss functions.
[0015] Further, in step S1, the modeling of the knowledge graph quadruple global features and node local features specifically includes: modeling the local features and global features of each node in the knowledge graph; s The local features of a are defined as the neighbor entities e that are directly connected to it v The features of and the features of its neighbor relationship r are modeled as:
[0016] F(e s )={(h v ,r s,v )|(h s ,r s,v ,h v ,t)∈G}
[0017] Where G = {(h s ,r s,v ,h v ,t)} represents the set of four tuples of knowledge graph, h s For entity e s The characteristic representation of h v For the tail entity e v The characteristic representation of s,v For entity e u and e v The relationship characteristics between , t is the corresponding time characteristics;
[0018] The global features of each entity are modeled as a relation path in the knowledge graph. Each path consists of a series of relation paths connecting two entities. That is, for a given entity pair e s and e o , connect e s and e o The path set is defined as {(e s ,r s,1 ,e 1 ,t 1 ),...,(e K ,r K,o ,e o ,t K )},make Represents entity e s To e o The lth relationship path of is modeled as:
[0019]
[0020] Where K is the path length, r l Represents the relationship in the lth path.
[0021] Further, in step S2, the position coding is designed based on the relative order of the relationship and the time difference, specifically including: for the lth relationship path between the entity pairs, determining the relationship position based on the relative order of the relationship and the time difference, specifically, let Represents an entity pair (e s ,e o ) is modeled as:
[0022]
[0023] in, Represents the relationship r in the lth path k and r j The difference, Represents the time difference between the kth and jth relations in the lth path, that is: Δr k,j =r j -r k , Δt k,j =t j -t k ;make express The output after position encoding is defined as:
[0024]
[0025] in, represents the vector concatenation operation; the encoded relational features are used as the input of the relational attention layer, represents the input matrix, defined as:
[0026] Further, in step S3, determining the importance of the relationship based on the intra-path relationship attention mechanism specifically includes: determining the importance of the relationship based on the intra-path attention mechanism for the matrix Perform a linear transformation:
[0027]
[0028] in, They are The query, key and value matrices obtained after linear transformation, are all linear transformation matrices; let Represents an entity pair (e s ,e o) in the lth path, is defined as:
[0029]
[0030] Where d represents the feature dimension; based on the attention weight matrix pair matrix Perform weighted summation to generate the aggregate representation of the lth path It is expressed as:
[0031]
[0032] Further, in step S4, the path importance is determined based on the inter-path attention mechanism, specifically including: let P s,o Represents a node pair (e s ,e o ) is defined as:
[0033]
[0034] Based on the inter-path attention mechanism s,o Perform linear transformation and model as follows:
[0035]
[0036] Among them, Q s,o , K s,o 、V s,o P s,o The query, key and value matrices obtained after linear transformation, are all linear transformation matrices; let β represent the inter-path attention matrix, defined as:
[0037]
[0038] Based on the calculated attention weight β, the global path feature of the node pair is generated, and g s,o Represents a node pair (e s ,e o ), defined as:
[0039] g s,o =βV p .
[0040] Further, in step S5, the encoding of local features based on GNN specifically includes: Represents node e in the nth layer GNN u The feature representation of , 1≤n≤N, N is the number of layers of GNN, and the node feature update formula is expressed as follows:
[0041]
[0042] in, and Represents the node e of the n+1th layer respectively v and the weight matrix of relation r, b n+1 Represents the bias vector of the n+1th layer.
[0043] Further, in step S6, the fusion of global features and local features based on MLP specifically includes: and Transformer output g s,o Splice and get the feature vector z s,o ,Right now:
[0044]
[0045] The concatenated feature vector z s,o As the input of the multi-layer MLP, let y represent the final feature fusion output, and the feature update formula is as follows:
[0046]
[0047] Among them, σ is the activation function, and They represent the weight matrix and bias matrix of the nth layer in MLP respectively.
[0048] Further, in step S7, the semantic features of the modeling new entity specifically include: modeling the new entity e u The initialization embedding representation h u , defined as follows:
[0049]
[0050] Among them, r u For the new entity e u Neighbor relationship, N u Represents the set of neighbor relationships of the new entity.
[0051] Further, in step S8, the modeling scoring function and loss function specifically include: let f represent the prediction score of the quadruple, which is defined as follows:
[0052] f=σ(W s y s,o +b s ),
[0053] Among them, σ is the activation function, W s and b s are the weight matrix and the bias matrix respectively; the constructed model is trained using the cross entropy loss function, which is defined as:
[0054]
[0055] Where M represents the number of samples, c i is the label of the i-th sample, c i =1 indicates a positive sample, c i =0 indicates a negative sample; the loss function is derived through the back-propagation algorithm, and the model parameters are updated using the gradient descent method to minimize the loss and maximize the score of the true quadruple. After the model converges, the missing elements in the quadruple can be predicted.
[0056] The beneficial effect of the present invention is that the method described in the present invention can effectively improve the reasoning ability of temporal inductive knowledge. By combining the relative position of the relationship, the time sequence and the attention mechanism between paths, the method of the present invention makes full use of the structured information and time evolution characteristics in the temporal knowledge graph, thereby improving the reasoning ability of unseen data.
[0057] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0059] Figure 1 It is a schematic diagram of the inductive temporal knowledge reasoning framework;
[0060] Figure 2 The figure is a flowchart of a temporal inductive knowledge reasoning method based on attention mechanism and GNN. DETAILED DESCRIPTION
[0061] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0062] See also Figure 1-2 , Figure 1This is a schematic diagram of the inductive temporal knowledge reasoning framework, such as Figure 1 As shown, the method includes a global feature encoder and a local feature encoder, which encode the relationship path and neighbor information respectively, and realizes feature fusion based on MLP, and trains the model by designing a scoring function and a loss function.
[0063] Figure 2 This is a flowchart of a temporal inductive knowledge reasoning method based on attention mechanism and GNN. Figure 2 As shown, the method specifically comprises the following steps:
[0064] Step 1: Modeling knowledge graph quadruple global features and node local features
[0065] Model the local and global features of each node in the knowledge graph; s The local features of a are defined as the neighbor entities e that are directly connected to it v The features of and the features of its neighbor relationship r are modeled as:
[0066] F(e s )={(h v ,r s,v )|(h s ,r s,v ,h v ,t)∈G}
[0067] Where G = {(h s ,r s,v ,h v ,t)} represents the set of four tuples of knowledge graph, h s For entity e s The characteristic representation of h v For the tail entity e v The feature representation of s,v For entity e u and e v The relationship characteristics between , t is the corresponding time characteristics;
[0068] The global features of each entity are modeled as a relation path in the knowledge graph. Each path consists of a series of relation paths connecting two entities. That is, for a given entity pair e s and e o , connect e s and e o The path set is defined as {(e s ,r s,1 ,e 1 ,t 1 ),...,(e K ,r K,o ,e o ,tK )},make Represents entity e s To e o The lth relationship path of is modeled as:
[0069]
[0070] Where K is the path length, r l Represents the relationship in the lth path.
[0071] Step 2: Design position encoding based on relative order and time difference
[0072] For the lth relationship path between entity pairs, the relationship position is determined based on the relative order and time difference of the relationship. Specifically, let Represents an entity pair (e s ,e o ) is modeled as:
[0073]
[0074] in, Represents the relationship r in the lth path k and r j The difference, Represents the time difference between the kth and jth relations in the lth path, that is: Δr k,j =r j -r k , Δt k,j =t j -t k ;make express The output after position encoding is defined as:
[0075]
[0076] in, represents the vector concatenation operation; the encoded relational features are used as the input of the relational attention layer, represents the input matrix, defined as:
[0077] Step 3: Determine the importance of relations based on the intra-path relation attention mechanism
[0078] Based on the in-path attention mechanism Perform a linear transformation:
[0079]
[0080] in, They are The query, key and value matrices obtained after linear transformation, are all linear transformation matrices; let Represents an entity pair (e s ,e o ) in the lth path, is defined as:
[0081]
[0082] Where d represents the feature dimension; based on the attention weight matrix pair matrix Perform weighted summation to generate the aggregate representation of the lth path It is expressed as:
[0083]
[0084] Step 4: Determine path importance based on inter-path attention mechanism
[0085] Let P s,o Represents a node pair (e s ,e o ) is defined as:
[0086]
[0087] Based on the inter-path attention mechanism s,o Perform linear transformation and model as follows:
[0088]
[0089] Among them, Q s,o , K s,o 、V s,o P s,o The query, key and value matrices obtained after linear transformation, are all linear transformation matrices; let β represent the inter-path attention matrix, defined as:
[0090]
[0091] Based on the calculated attention weight β, the global path feature of the node pair is generated, and g s,o Represents a node pair (e s ,e o ), defined as:
[0092] g s,o =βV p .
[0093] Step 5: Encode local features based on GNN
[0094] make Represents node e in the nth layer GNN u The feature representation of , 1≤n≤N, N is the number of layers of GNN, and the node feature update formula is expressed as follows:
[0095]
[0096] in, and Represents the node e of the n+1th layer respectively v and the weight matrix of relation r, b n+1 Represents the bias vector of the n+1th layer.
[0097] Step 6: Fusion of global and local features based on MLP
[0098] Output of GNN and Transformer output g s,o Splice and get the feature vector z s,o ,Right now:
[0099]
[0100] The concatenated feature vector z s,o As the input of the multi-layer MLP, let y represent the final feature fusion output, and the feature update formula is as follows:
[0101]
[0102] Among them, σ is the activation function, and They represent the weight matrix and bias matrix of the nth layer in MLP respectively.
[0103] Step 7: Modeling semantic features of new entities
[0104] Modeling new entities u The initialization embedding representation h u , defined as follows:
[0105]
[0106] Among them, r u For the new entity e u Neighbor relationship, N u Represents the set of neighbor relationships of the new entity.
[0107] Step 8: Modeling the scoring function and loss function
[0108] Let f denote the prediction score of a quadruple, defined as follows:
[0109] f=σ(W s y s,o +b s ),
[0110] Among them, σ is the activation function, W s and b s are the weight matrix and the bias matrix respectively; the constructed model is trained using the cross entropy loss function, which is defined as:
[0111]
[0112] Where M represents the number of samples, c i is the label of the i-th sample, c i =1 indicates a positive sample, c i =0 indicates a negative sample; the loss function is derived through the back-propagation algorithm, and the model parameters are updated using the gradient descent method to minimize the loss and maximize the score of the true quadruple. After the model converges, the missing elements in the quadruple can be predicted.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A temporal inductive knowledge reasoning method based on attention mechanism and GNN, characterized in that: The method specifically comprises the following steps: S1: Modeling knowledge graph quadruple global features and node local features; S2: Design position encoding based on relative order and time difference; S3: Determine the importance of relations based on the intra-path relation attention mechanism; S4: Determine path importance based on path attention mechanism; S5: Encode local features based on GNN; S6: Fusion of global features and local features based on MLP; S7: Modeling new entity semantic features; S8: Modeling scoring and loss functions.
2. The temporal inductive knowledge reasoning method based on attention mechanism and GNN according to claim 1 is characterized in that: In step S1, the modeling of the knowledge graph quadruple global features and node local features specifically includes: Model the local and global features of each node in the knowledge graph; s The local features of a are defined as the neighbor entities e that are directly connected to it v The features of and the features of its neighbor relationship r are modeled as: F(e s )={(h v ,r s,v )|(h s ,r s,v ,h v ,t)∈G} Where G = {(h s ,r s,v ,h v ,t)} represents the set of four tuples of knowledge graph, h s For entity e s The characteristic representation of h v For the tail entity e v The characteristic representation of s,v For entity e u and e v The relationship characteristics between , t is the corresponding time characteristics; The global features of each entity are modeled as a relation path in the knowledge graph. Each path consists of a series of relation paths connecting two entities. s and e o , connect e s and e o The path set is defined as {(e s ,r s,1 ,e1,t1),...,(e K ,r K,o ,e o ,t K )},make Represents entity e s To e o The lth relationship path of is modeled as: Where K is the path length, r l Represents the relationship in the lth path.
3. The temporal inductive knowledge reasoning method based on attention mechanism and GNN according to claim 2 is characterized in that: In step S2, the position coding is designed based on the relative order of the relationship and the time difference, specifically including: For the lth relationship path between entity pairs, the relationship position is determined based on the relative order and time difference of the relationship. Specifically, let Represents an entity pair (e s ,e o ) is modeled as: in, Represents the relationship r in the lth path k and r j The difference, Represents the time difference between the kth and jth relations in the lth path, that is: Δr k,j =r j -r k , Δt k,j =t j -t k ;make express The output after position encoding is defined as: in, represents the vector concatenation operation; the encoded relational features are used as the input of the relational attention layer, represents the input matrix, defined as:
4. The temporal inductive knowledge reasoning method based on attention mechanism and GNN according to claim 3 is characterized in that: In step S3, determining the importance of the relationship based on the intra-path relationship attention mechanism specifically includes: Based on the in-path attention mechanism Perform a linear transformation: in, They are The query, key and value matrices obtained after linear transformation, are all linear transformation matrices; let Represents an entity pair (e s ,e o ) in the lth path, is defined as: Where d represents the feature dimension; based on the attention weight matrix pair matrix Perform weighted summation to generate the aggregate representation of the lth path It is expressed as:
5. The temporal inductive knowledge reasoning method based on attention mechanism and GNN according to claim 4 is characterized in that: In step S4, the path importance is determined based on the inter-path attention mechanism, specifically including: Let P s,o Represents a node pair (e s ,e o ) is defined as: Based on the inter-path attention mechanism s,o Perform linear transformation and model as follows: Among them, Q s,o , K s,o 、V s,o P s,o The query, key and value matrices obtained after linear transformation, are all linear transformation matrices; let β represent the inter-path attention matrix, defined as: Based on the calculated attention weight β, the global path feature of the node pair is generated, and g s,o Represents a node pair (e s ,e o ), defined as: g s,o =βV p 。 6. The temporal inductive knowledge reasoning method based on attention mechanism and GNN according to claim 5 is characterized in that: In step S5, encoding the local features based on GNN specifically includes: make Represents node e in the nth layer GNN u The feature representation of , 1≤n≤N, N is the number of layers of GNN, and the node feature update formula is expressed as follows: in, and Represents the node e of the n+1th layer respectively v and the weight matrix of relation r, b n+1 Represents the bias vector of the n+1th layer.
7. The temporal inductive knowledge reasoning method based on attention mechanism and GNN according to claim 6 is characterized in that: In step S6, the fusion of global features and local features based on MLP specifically includes: Output of GNN and the output g of Transformer s,o Splice and get the feature vector z s,o ,Right now: The concatenated feature vector z s,o As the input of the multi-layer MLP, let y represent the final feature fusion output, and the feature update formula is as follows: Among them, σ is the activation function, and They represent the weight matrix and bias matrix of the nth layer in MLP respectively.
8. The temporal inductive knowledge reasoning method based on attention mechanism and GNN according to claim 7 is characterized in that: In step S7, the semantic features of the modeled new entity specifically include: Modeling new entities u The initialization embedding representation h u , defined as follows: Among them, r u For the new entity e u Neighbor relationship, N u Represents the set of neighbor relationships of the new entity.
9. The temporal inductive knowledge reasoning method based on attention mechanism and GNN according to claim 8 is characterized in that: In step S8, the modeling scoring function and loss function specifically include: Let f denote the prediction score of a quadruple, defined as follows: f=σ(W s y s,o +b s ), Among them, σ is the activation function, W s and b s are the weight matrix and the bias matrix respectively; the constructed model is trained using the cross entropy loss function, which is defined as: Where M represents the number of samples, c i is the label of the i-th sample, c i =1 indicates a positive sample, c i =0 indicates a negative sample; the loss function is derived through the back-propagation algorithm, and the model parameters are updated using the gradient descent method to minimize the loss and maximize the score of the true quadruple. After the model converges, the missing elements in the quadruple can be predicted.