Graph neural network knowledge graph inference learning method based on attention mechanism
By combining multi-head attention mechanism and graph neural network, the problems of incomplete information and complex structure in knowledge graph reasoning are solved, realizing an efficient and interpretable knowledge graph reasoning method, which improves semantic accuracy and computational efficiency.
Patent Information
- Application Number
- CN202511382941.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-10-31
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies in knowledge graph reasoning suffer from incomplete information, low efficiency and poor scalability of traditional methods, difficulty in capturing complex structural information based on embedding, and challenges in multi-source information fusion and key information focusing of graph neural networks.
By employing a multi-head attention mechanism combined with a graph neural network, entity sets are processed layer by layer through a multi-layer cascaded structure to calculate message and edge weights. Combined with Gumbel-Top-K sampling technology and GRU gating mechanism, automatic decay of the influence of high-level entities and dynamic weight calculation are achieved.
It improves the semantic accuracy, computational efficiency, and interpretability of knowledge graph reasoning, and enhances the reasoning accuracy and scalability of graph neural networks.
Smart Images

Figure CN120875002A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of knowledge graph reasoning and learning technology, and in particular to a graph neural network-based knowledge graph reasoning and learning method based on an attention mechanism. Background Technology
[0002] With the development of artificial intelligence, knowledge graphs are widely used in many fields, but they suffer from incomplete information, necessitating effective reasoning methods to mine implicit knowledge. Traditional reasoning methods each have their limitations: rule-based reasoning relies on manually formulated rules, resulting in low efficiency and poor scalability; while embedding-based reasoning is computationally efficient, it has limited ability to express semantic relationships and struggles to capture complex structural information. Graph neural networks have significant advantages in processing graph-structured data; however, in knowledge graph reasoning, they face challenges in multi-source information fusion and focusing on key information. Attention mechanisms can allocate weights based on information importance and have yielded significant results in other fields. Combining attention mechanisms with graph neural networks for knowledge graph reasoning is expected to overcome the shortcomings of existing methods and improve reasoning performance; this is the research background of this invention.
[0003] Therefore, this invention proposes an attention-based graph neural network knowledge graph reasoning learning method to solve the above problems. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention develops a graph neural network knowledge graph reasoning learning method based on an attention mechanism. This invention employs a multi-head attention mechanism to capture the association information of nodes and edges in the propagation path, which can improve the prediction accuracy of the model.
[0005] The technical solution of this invention to solve the technical problem is a graph neural network knowledge graph reasoning learning method based on an attention mechanism, comprising the following steps: S1. Load the knowledge graph dataset and generate the corresponding reverse triples based on the data in the knowledge graph dataset; S2. Construct a knowledge graph for training based on the loaded knowledge graph dataset and the generated inverse triples. Read the query information and initialize the entity set of the 0th layer propagation path of the query entity in the query information; S3. The entity set of the propagation path of the query entity at layer 0 is processed layer by layer through the multi-head attention mechanism to obtain the entity set of each layer. Then, the message and edge weight are calculated through the multi-head attention mechanism based on the entity, relationship and position information in the entity set of each layer. The message is then updated according to the edge weight to obtain the updated message of each triple. Finally, the updated messages are concatenated and linearly mapped to obtain the target entity. S4. Define a scoring function to calculate the score of the target entity. Calculate the probability of generating the target entity based on its score. Use Gumbel-Top-K sampling technique to sample the top K highest-scoring entities from the probability distribution of the target entities and merge them with the entity set of the previous layer to form the entity set of the current layer. S5. Repeat steps S3-S4 until the propagation path depth reaches... At that time, the first The complete entity set of each layer, calculate the first layer respectively. The target entity with the highest score in the layer entity set is selected as the final target entity to be queried.
[0006] S1 is as follows: Load the knowledge graph dataset. Each data point in the knowledge graph dataset is represented as a triple. A set of triplets is constructed based on each triplet. ,in, Indicates the head entity. Indicates the tail entity. express and Relationship, Describe the set of triples The set of head and tail entities. Describe the set of triples China-US relations A set; According to the set of triples Generate a set of reverse triples , , express and Relationship, Represents the set of reverse triples China-US relations The set, , Describe the set of triples The number of relations between China and the United States exist The triple represents the head entity. exist The triple represents the tail entity.
[0007] S2 is as follows: S2.1, The set of triples formed by the data in the knowledge graph dataset. and based on triplet sets The generated set of reverse triples Generate a set of triples for training. , ; Based on the set of triples used for training Building knowledge graphs The specific process is as follows: Set the triplet set In and Mapped to knowledge graph The nodes in the set of triples The relationships in the graph are mapped to a knowledge graph. From the edges in the graph, we obtain the knowledge graph. , , Representation of knowledge graph A set of relationships; S2.2, Read query information , , Indicates the entity being queried. Indicates the target entity to be queried. This indicates the relationship between the queried entity and the entity to be queried, for the queried entity The set of entities in the propagation path at layer 0. Perform initialization.
[0008] S3 is as follows: The multi-head attention mechanism employs a multi-layered cascading structure to process the entity set along the 0th layer propagation path of the query entity, layer by layer. During message passing at each layer, according to the first layer Layer Entity Collection From knowledge graphs The set of direct neighbor entities queried in the middle Calculate the new target entity set ,from Remove from The entity in, and guarantee The calculation formula is as follows: ; from The triplet information obtained from the data is represented as follows: , , The number of target entities is The number of triples is , No. Layer query entities Indicates the first The target entity of the layer, the first target entity of layer Indicates the first The query entity of the layer, Indicates the first Layer query entities and the target entity of layer The relationship between them; for Each triple in the set will Embedded feature vector representation , Embedded feature vector representation , Embedded feature vector representation and location information The message is obtained by performing calculations using a multi-head attention mechanism. Location information , The initial value is 1, where , , This represents the embedding dimension of the entity vector. The embedding dimension of the resulting message vector is represented by the following: , , and The inputs are fed into the multi-head attention mechanism module to calculate the edge weights. By edge weight Regarding the message The update is performed, and then all the updated messages are concatenated to obtain the target entity based on the concatenated messages.
[0009] The specific computational process of the multi-head attention mechanism is as follows: (1) , , respectively with By splicing the components, the features are obtained separately. , , , Then, a linear transformation is performed to obtain the query matrix of the current layer's multi-head attention. Key matrix Sum matrix The calculation formula is as follows: , , , , , , in, This indicates a splicing operation. , , These represent the trainable weight matrices used for the query matrix, key matrix, and value matrix, respectively. , , Let represent the trainable bias vectors used for the query matrix, key matrix, and value matrix, respectively. , Indicates the number of long positions; (2) , , Divided into Size, Attention weights are calculated for each head, specifically through matrix multiplication and... Function to calculate attention weights , The calculation formula is as follows: , in, Indicates transpose; Next and Perform a weighted summation to obtain the weighted output. , Then through a fully connected layer Restored to the original dimensions, the calculation formula in the fully connected layer is as follows: , in, Indicates the relationships between nodes. , This represents the operation of expanding a matrix into row vectors. This represents a trainable weight matrix. , This represents a trainable bias vector. ; (3) Through a linear layer pair Projection, and through The function normalizes its values to (0,1) and then uses a gating mechanism to fuse them. and The message indicated The calculation formula is as follows: , , in, Represents the weight vector. express function, Represents element-wise product; (4) , , and The inputs are processed together using multi-head attention. The calculation method for each input is the same as in steps (1)-(3), but the computational parameters are independent of each other. After multi-head attention calculation, the inputs are processed by ReLU activation function, linear mapping, and... The edge weights are obtained after the function processes the edge layer by layer. edge weight It is a one-dimensional vector with edge weights. The calculation formula is as follows: , in, This represents multi-head attention computation; (5) Based on edge weights Update message Generate updated information The calculation formula is as follows: ; For the l Layer Repeat steps (1)-(4) to obtain three triples. One reason Received message Then, the message is updated, and subsequently, the updated messages of all candidate entities in that layer are concatenated to obtain the messages of all target entities in that layer. The formula is as follows: , Messages pointing to the same target entity are aggregated and, after a linear mapping, the target entity is obtained. Feature representation The calculation formula is as follows: , in, , express l The number of target entities in the layer. This represents the trainable weight matrix, and ACG represents the summation operation. Indicates query information The target entity to be queried .
[0010] S4 is as follows: For target entity Define the scoring function The calculation formula is as follows: , in, Indicates trainable parameters, Representation device; Based on score Calculate and generate target entity probability distribution The calculation formula is as follows: , in, This represents the temperature parameter used to adjust the smoothness of the distribution. , express l All candidate target entities in the layer; Using Gumbel-Top-K technology to analyze the target entity probability distribution The K highest-scoring target entities before sampling are denoted as And the sampled entities are compared with the historical entity set. Merge, form l The complete set of entities in a layer is calculated using the following formula: , Update the feature representation of the target entity using the GRU gating mechanism. The updated data is then passed to the next layer of the neural network in the multi-head attention mechanism.
[0011] S5 is detailed below: Repeat steps S3-S4 until the propagation path depth reaches At that time, the first The complete entity collection of a layer Through the first The complete entity collection of a layer Feature representation of target entities Predict a score for each target entity, and select the target entity with the highest score as the final target entity to be queried. The formula for calculating the target entity score is as follows: , in, This represents a trainable weight matrix.
[0012] The effects described in the invention are merely those of the embodiments, and not all the effects of the invention. The above technical solutions have the following advantages or beneficial effects: This invention combines triples with a multi-head attention mechanism. Specifically, it concatenates and fuses the relation feature vector of the query triple with the corresponding layer feature vector to form the query matrix Q in the attention mechanism; it concatenates and fuses the head entity feature vector of the query with the layer feature vector to form the key matrix K in the attention mechanism; and it concatenates and fuses the feature vectors of possible tail entities in the path with the layer feature vector to form the value matrix V. Simultaneously, it introduces a layer decay factor, dynamically adjusting the importance weights of entities at different levels by taking the reciprocal of the layer number. This enables automatic decay of the influence of high-level entities and dynamic weight calculation for different level paths, enhancing the modeling ability of coefficient relationships. It provides an extensible and general framework for complex relational reasoning, thereby improving the semantic accuracy, computational efficiency, and interpretability in knowledge graph reasoning. This invention also employs a dual-path parallel multi-head attention mechanism architecture in the attention calculation module. The parameter matrices of the two multi-head attention sub-modules are independent of each other and they perform forward propagation calculations separately. The first attention sub-module is dedicated to calculating the message vector in the propagation path, and the second attention sub-module is dedicated to calculating the weight coefficients of the edges. The dual-path attention mechanism can decouple the functions of message passing and edge weight calculation through independent parameter space learning, providing an interpretable and scalable solution for complex relationship modeling.
[0013] In summary, by employing a multi-head attention mechanism to capture the correlation information between nodes and their edges in the propagation path, this invention can effectively improve the accuracy of graph neural network knowledge graph reasoning and learning. Attached Figure Description
[0014] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0015] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation
[0016] To clearly illustrate the technical features of this solution, the invention will be described in detail below through specific embodiments and in conjunction with the accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure of the invention, the components and arrangements of specific examples are described below.
[0017] Example 1 An attention-based graph neural network-based knowledge graph reasoning learning method, with the following specific steps: S1. Load the knowledge graph dataset and generate the corresponding reverse triples based on the data in the knowledge graph dataset; S2. Construct a knowledge graph for training based on the loaded knowledge graph dataset and the generated inverse triples. Read the query information and initialize the entity set of the 0th layer propagation path of the query entity in the query information; S3. The entity set of the propagation path of the query entity at layer 0 is processed layer by layer through the multi-head attention mechanism to obtain the entity set of each layer. Then, the message and edge weight are calculated through the multi-head attention mechanism based on the entity, relationship and position information in the entity set of each layer. The message is then updated according to the edge weight to obtain the updated message of each triple. Finally, the updated messages are concatenated and linearly mapped to obtain the target entity. S4. Define a scoring function to calculate the score of the target entity. Calculate the probability of generating the target entity based on its score. Use Gumbel-Top-K sampling technique to sample the top K highest-scoring entities from the probability distribution of the target entities and merge them with the entity set of the previous layer to form the entity set of the current layer. S5. Repeat steps S3-S4 until the propagation path depth reaches... At that time, the first The complete entity set of each layer, calculate the first layer respectively. The target entity with the highest score in the layer entity set is selected as the final target entity to be queried.
[0018] In a specific implementation, S1 is as follows: Load the knowledge graph dataset. Each data point in the knowledge graph dataset is represented as a triple. A set of triplets is constructed based on each triplet. ,in, Indicates the head entity. Indicates the tail entity. express and Relationship, Describe the set of triples The set of head and tail entities. Describe the set of triples China-US relations A set; According to the set of triples Generate a set of reverse triples , , express and Relationship, Represents the set of reverse triples China-US relations The set, , Describe the set of triples The number of relations between China and the United States exist The triple represents the head entity. exist The triple represents the tail entity.
[0019] In a specific implementation, S2 is as follows: S2.1, The set of triples formed by the data in the knowledge graph dataset. and based on triplet sets The generated set of reverse triples Generate a set of triples for training. , ; Based on the set of triples used for training Building knowledge graphs The specific process is as follows: Set the triplet set In and Mapped to knowledge graph The nodes in the set of triples The relationships in the graph are mapped to a knowledge graph. From the edges in the graph, we obtain the knowledge graph. , , Representation of knowledge graph A set of relationships; S2.2, Read query information , , Indicates the entity being queried. Indicates the target entity to be queried. This indicates the relationship between the queried entity and the entity to be queried, for the queried entity The set of entities in the propagation path at layer 0. Perform initialization.
[0020] In a specific implementation, S3 is as follows: The multi-head attention mechanism employs a multi-layered cascading structure to process the entity set along the 0th layer propagation path of the query entity, layer by layer. During message passing at each layer, according to the first layer Layer Entity Collection From knowledge graphs The set of direct neighbor entities queried in the middle Calculate the new target entity set ,from Remove from The entity in, and guarantee The calculation formula is as follows: ; from The triplet information obtained from the data is represented as follows: , , The number of target entities is The number of triples is , No. Layer query entities Indicates the first The target entity of the layer, the first target entity of layer Indicates the first The query entity of the layer, Indicates the first Layer query entities and the target entity of layer The relationship between them; for Each triple in the set will Embedded feature vector representation , Embedded feature vector representation , Embedded feature vector representation and location information The message is obtained by performing calculations using a multi-head attention mechanism. Location information , The initial value is 1, where , , This represents the embedding dimension of the entity vector. The embedding dimension of the resulting message vector is represented by the following: , , and The inputs are fed into the multi-head attention mechanism module to calculate the edge weights. By edge weight Regarding the message The update is performed, and then all the updated messages are concatenated to obtain the target entity based on the concatenated messages.
[0021] In a specific implementation, the computational process of the multi-head attention mechanism is as follows: (1) , , respectively with By splicing the components, the features are obtained separately. , , , Then, a linear transformation is performed to obtain the query matrix of the current layer's multi-head attention. Key matrix Sum matrix The calculation formula is as follows: , , , , , , in, This indicates a splicing operation. , , These represent the trainable weight matrices used for the query matrix, key matrix, and value matrix, respectively. , , Let represent the trainable bias vectors used for the query matrix, key matrix, and value matrix, respectively. , Indicates the number of long positions; (2) , , Divided into Size, Attention weights are calculated for each head, specifically through matrix multiplication and... Function to calculate attention weights , The calculation formula is as follows: , in, Indicates transpose; Next and Perform a weighted summation to obtain the weighted output. , Then through a fully connected layer Restored to the original dimensions, the calculation formula in the fully connected layer is as follows: , in, Indicates the relationships between nodes. , This represents the operation of expanding a matrix into row vectors. This represents a trainable weight matrix. , This represents a trainable bias vector. ; (3) Through a linear layer pair Projection, and through The function normalizes its values to (0,1) and then uses a gating mechanism to fuse them. and The message indicated The calculation formula is as follows: , , in, Represents the weight vector. express function, Represents element-wise product; (4) , , and The inputs are processed together using multi-head attention. The calculation method for each input is the same as in steps (1)-(3), but the computational parameters are independent of each other. After multi-head attention calculation, the inputs are processed by ReLU activation function, linear mapping, and... The edge weights are obtained after the function processes the edge layer by layer. edge weight It is a one-dimensional vector with edge weights. The calculation formula is as follows: , in, This represents multi-head attention computation; (5) Based on edge weights Update message Generate updated information The calculation formula is as follows: ; For the l Layer Repeat steps (1)-(4) to obtain three triples. One reason Received message Then, the message is updated, and subsequently, the updated messages of all candidate entities in that layer are concatenated to obtain the messages of all target entities in that layer. The formula is as follows: , Messages pointing to the same target entity are aggregated and, after a linear mapping, the target entity is obtained. Feature representation The calculation formula is as follows: , in, , express l The number of target entities in the layer. This represents the trainable weight matrix, and ACG represents the summation operation. Indicates query information The target entity to be queried .
[0022] In a specific implementation, S4 is as follows: For target entity Define the scoring function The calculation formula is as follows: , in, Indicates trainable parameters, Representation device; Based on score Calculate and generate target entity probability distribution The calculation formula is as follows: , in, This represents the temperature parameter used to adjust the smoothness of the distribution. , express l All candidate target entities in the layer; Using Gumbel-Top-K technology to analyze the target entity probability distribution The K highest-scoring target entities before sampling are denoted as And the sampled entities are compared with the historical entity set. Merge, form l The complete set of entities in a layer is calculated using the following formula: , Update the feature representation of the target entity using the GRU gating mechanism. The updated data is then passed to the next layer of the neural network in the multi-head attention mechanism.
[0023] In a specific implementation, S5 is as follows: Repeat steps S3-S4 until the propagation path depth reaches At that time, the first The complete entity collection of a layer Through the first The complete entity collection of a layer Feature representation of target entities Predict a score for each target entity, and select the target entity with the highest score as the final target entity to be queried. The formula for calculating the target entity score is as follows: , in, This represents a trainable weight matrix.
[0024] Example 2 To demonstrate that the method of this invention improves upon existing technologies, experiments were conducted on the WN18RR knowledge graph inference dataset, comparing the method of this invention with the current best knowledge graph inference models. The current best knowledge graph inference models include ConvE (Convolutional Knowledge Graph Embedding), CompGCN (Combined Graph Convolutional Network), NBFNet (Neural Bellman-Ford Network), RED-GNN (Relation Evolution Detection Graph Neural Network), and AdaProp (Adaptive Propagation Model). The evaluation metrics were MRR (MRR represents the average of the inverse ranking of correct predictions; a smaller value indicates a higher ranking) and Hit@k (Hit@k represents the proportion of correct predictions within the top-k range; a larger value is better, with k taking values of 1 and 10 respectively). Table 1 shows that the scheme of this invention has a more significant advantage in inference ability, and the method of this invention has higher prediction accuracy and better performance.
[0025] Table 1. Performance comparison of the method of this invention with the current best knowledge graph reasoning model on the WN18RR dataset in reasoning tasks. Example 3 To demonstrate the practicality of the method of this invention, it is applied to a smart home system. Since the linkage rules between devices in a smart home system (such as "if the temperature sensor detects a high temperature, the air conditioner will automatically turn on") typically rely on manually predefined rules or simple rule engines, it is difficult to handle complex scenarios (such as multi-device collaboration, user habit learning, etc.). Therefore, the attention-based graph neural network knowledge graph reasoning method proposed in this invention is applied to a smart home system. By constructing a smart home knowledge graph, the implicit relationships between devices can be automatically mined, enabling more intelligent decision-making.
[0026] The specific operation process is as follows: Building a smart home knowledge graph Original triple example: (temperature sensor, trigger, air conditioner), (human sensor, detection, unoccupied state), (light, off condition, unoccupied state); generate inverse triple: (air conditioner, triggered, temperature sensor), (unoccupied state, detected, human sensor). Map devices (such as "temperature sensor", "air conditioner") as nodes, and relationships (such as "trigger", "detection") as edges to form a graph structure.
[0027] Generate a query query= based on the above triples. ,in , This represents the target entity to be queried, with the goal of inferring the devices that the "temperature sensor" might trigger; initialize the level 0 entity set. Then eigenvector representation , eigenvector representation ,as well as Embedded feature vector representation Separately with location information By concatenating the vectors, we obtain the corresponding position vectors: , and ,in =1 indicates that the current layer is the 1st layer of the network; right , and Perform a linear transformation to obtain the query matrix of the current layer's multi-head attention. Key matrix Sum matrix Subsequently through and Calculate attention weights ; Next and Perform a weighted summation to obtain the weighted output. And through a fully connected layer Restored to the original dimension; Through a linear layer pair Projection, and through The function normalizes its values to (0,1) and then uses a gating mechanism to fuse them. and The message indicated ; The same will , , and Multi-head attention calculation of edge weights ; Based on edge weights Update message Generate updated information ; For the l Repeat the above steps for all triples in the layer to obtain the message corresponding to each triple. With edge weight After completing the message update, the updated messages of all candidate entities in this layer are concatenated to obtain the messages of all target entities in this layer. ; Then, messages pointing to the same target entity are aggregated, and after a linear mapping, the target entity is obtained. Feature representation ; Based on score Calculate and generate target entity probability distribution ; Using Gumbel-Top-K technology to analyze the target entity probability distribution The K highest-scoring target entities before sampling are denoted as And the sampled entities are compared with the historical entity set. Merge, form l The complete collection of entities in a layer; Update the feature representation of the target entity using the GRU gating mechanism. The updated data is then passed to the next layer of the neural network.
[0028] Repeat the above steps until the propagation path depth reaches When the layer is reached, the first layer is obtained. The complete entity collection of a layer Through the first The complete entity collection of a layer Feature representation of target entities Predict the score for each target entity; The final scores for the target entities were: "Air Conditioner" 0.9, "Fan" 0.08, and "Humidifier" 0.02. The target entity with the highest score, "Air Conditioner," was selected as the final prediction result. This allows for the selective activation of the air conditioner or fan based on the threshold of the indoor temperature sensor.
[0029] Although the specific embodiments of the invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the invention. Based on the technical solutions of the invention, various modifications or variations that can be made by those skilled in the art without creative effort are still within the scope of protection of the invention.
Claims
1. A graph neural network-based knowledge graph reasoning learning method based on an attention mechanism, characterized in that, Includes the following steps: S1. Load the knowledge graph dataset and generate the corresponding reverse triples based on the data in the knowledge graph dataset; S2. Construct a knowledge graph for training based on the loaded knowledge graph dataset and the generated inverse triples. Read the query information and initialize the entity set of the 0th layer propagation path of the query entity in the query information; S3. The entity set of the propagation path of the query entity at layer 0 is processed layer by layer through the multi-head attention mechanism to obtain the entity set of each layer. Then, the message and edge weight are calculated through the multi-head attention mechanism based on the entity, relationship and position information in the entity set of each layer. The message is then updated according to the edge weight to obtain the updated message of each triple. Finally, the updated messages are concatenated and linearly mapped to obtain the target entity. S4. Define a scoring function to calculate the score of the target entity. Calculate the probability of generating the target entity based on its score. Use Gumbel-Top-K sampling technique to sample the top K highest-scoring entities from the probability distribution of the target entities and merge them with the entity set of the previous layer to form the entity set of the current layer. S5. Repeat steps S3-S4 until the propagation path depth reaches... At that time, the first The complete entity set of each layer, calculate the first layer respectively. The target entity with the highest score in the layer entity set is selected as the final target entity to be queried.
2. The graph neural network knowledge graph reasoning learning method based on attention mechanism according to claim 1, characterized in that, S1 is as follows: Load the knowledge graph dataset. Each data point in the knowledge graph dataset is represented as a triple. A set of triplets is constructed based on each triplet. ,in, Indicates the head entity. Indicates the tail entity. express and Relationship, Describe the set of triples The set of head and tail entities. Describe the set of triples China-US relations A set; According to the set of triples Generate a set of reverse triples , , express and Relationship, Represents the set of reverse triples China-US relations The set, , Describe the set of triples The number of relations between China and the United States exist The triple represents the head entity. exist The triple represents the tail entity.
3. The graph neural network knowledge graph reasoning learning method based on an attention mechanism according to claim 2, characterized in that, S2 is as follows: S2.1, The set of triples formed by the data in the knowledge graph dataset. and based on triplet sets The generated set of reverse triples Generate a set of triples for training. , ; Based on the set of triples used for training Building knowledge graphs The specific process is as follows: Set the triplet set In and Mapped to knowledge graph The nodes in the set of triples The relationships in the graph are mapped to a knowledge graph. From the edges in the graph, we obtain the knowledge graph. , , Representation of knowledge graph A set of relationships; S2.2, Read query information , , Indicates the entity being queried. Indicates the target entity to be queried. This indicates the relationship between the queried entity and the entity to be queried, for the queried entity The set of entities in the propagation path at layer 0. Perform initialization.
4. The graph neural network knowledge graph reasoning learning method based on attention mechanism according to claim 3, characterized in that, S3 Specifically as follows: The multi-head attention mechanism employs a multi-layered cascading structure to process the entity set along the 0th layer propagation path of the query entity, layer by layer. During message passing at each layer, according to the first layer Layer Entity Collection From knowledge graphs The set of direct neighbor entities queried in the middle Calculate the new target entity set ,from Remove from The entity in, and guarantee The calculation formula is as follows: ; from The triplet information obtained from the data is represented as follows: , , The number of target entities is The number of triples is , No. Layer query entities Indicates the first The target entity of the layer, the first target entity of layer Indicates the first The query entity of the layer, Indicates the first Layer query entities and the target entity of layer The relationship between them; for Each triple in the set will Embedded feature vector representation , Embedded feature vector representation , Embedded feature vector representation and location information The message is obtained by performing calculations using a multi-head attention mechanism. Location information , The initial value is 1, where , , This represents the embedding dimension of the entity vector. The embedding dimension of the resulting message vector is represented by the following: , , and The inputs are fed into the multi-head attention mechanism module to calculate the edge weights. By edge weight Regarding the message The update is performed, and then all the updated messages are concatenated to obtain the target entity based on the concatenated messages.
5. The graph neural network knowledge graph reasoning learning method based on an attention mechanism according to claim 4, characterized in that, The specific computational process of the multi-head attention mechanism is as follows: (1) , , respectively with By splicing the components, the features are obtained separately. , , , Then, a linear transformation is performed to obtain the query matrix of the current layer's multi-head attention. Key matrix Sum matrix The calculation formula is as follows: , , , , , , in, This indicates a splicing operation. , , These represent the trainable weight matrices used for the query matrix, key matrix, and value matrix, respectively. , , Let represent the trainable bias vectors used for the query matrix, key matrix, and value matrix, respectively. , Indicates the number of long positions; (2) , , Divided into Size, Attention weights are calculated for each head, specifically through matrix multiplication and... Function to calculate attention weights , The calculation formula is as follows: , in, Indicates transpose; Next and Perform a weighted summation to obtain the weighted output. , Then through a fully connected layer Restored to the original dimensions, the calculation formula in the fully connected layer is as follows: , in, Indicates the relationships between nodes. , This represents the operation of expanding a matrix into row vectors. This represents a trainable weight matrix. , This represents a trainable bias vector. ; (3) Through a linear layer pair Projection, and through The function normalizes its values to (0,1) and then uses a gating mechanism to fuse them. and The message indicated The calculation formula is as follows: , , in, Represents the weight vector. express function, Represents element-wise product; (4) , , and The inputs are processed together using multi-head attention. The calculation method for each input is the same as in steps (1)-(3), but the computational parameters are independent of each other. After multi-head attention calculation, the inputs are processed by ReLU activation function, linear mapping, and... The edge weights are obtained after the function processes the edge layer by layer. edge weight It is a one-dimensional vector with edge weights. The calculation formula is as follows: , in, This represents multi-head attention computation; (5) Based on edge weights Update message Generate updated information The calculation formula is as follows: ; For the l Layer Repeat steps (1)-(4) to obtain three triples. One reason Received message Then, the message is updated, and subsequently, the updated messages of all candidate entities in that layer are concatenated to obtain the messages of all target entities in that layer. The formula is as follows: , Messages pointing to the same target entity are aggregated and, after a linear mapping, the target entity is obtained. Feature representation The calculation formula is as follows: , in, , express l The number of target entities in the layer. This represents the trainable weight matrix, and ACG represents the summation operation. Indicates query information The target entity to be queried .
6. The graph neural network knowledge graph reasoning learning method based on an attention mechanism according to claim 5, characterized in that, S4 is as follows: For target entity Define the scoring function The calculation formula is as follows: , in, Indicates trainable parameters, Representation device; Based on score Calculate and generate target entity probability distribution The calculation formula is as follows: , in, This represents the temperature parameter used to adjust the smoothness of the distribution. , express l All candidate target entities in the layer; Using Gumbel-Top-K technology to analyze the target entity probability distribution The K highest-scoring target entities before sampling are denoted as And the sampled entities are compared with the historical entity set. Merge, form l The complete set of entities in a layer is calculated using the following formula: , Update the feature representation of the target entity using the GRU gating mechanism. The updated data is then passed to the next layer of the neural network in the multi-head attention mechanism.
7. The graph neural network knowledge graph reasoning learning method based on an attention mechanism according to claim 6, characterized in that, S5 is detailed below: Repeat steps S3-S4 until the propagation path depth reaches At that time, the first The complete entity collection of a layer Through the first The complete entity collection of a layer Feature representation of target entities Predict a score for each target entity, and select the target entity with the highest score as the final target entity to be queried. The formula for calculating the target entity score is as follows: , in, This represents a trainable weight matrix.