A Knowledge Graph Prediction Method Based on Attention Mechanism
By introducing a prediction method based on attention mechanism in the knowledge graph completion method, the graph attention network and the Transformer relationship aggregator are combined, and the completion problem of low-frequency relationships and unusual entities is solved, and the accuracy and robustness of knowledge graph completion are improved.
Patent Information
- Application Number
- CN202310721138.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-19
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-06-19
AI Technical Summary
The existing knowledge graph completion method is not effective when dealing with low-frequency relationships and uncommon entities, and is prone to amplification of noise in sparse neighborhood scenarios, reducing completion accuracy.
Using the knowledge graph prediction method based on attention mechanism, the one-hop neighborhood information of the fusion entity is integrated through the graph attention network, a small sample relationship representation of the fusion neighborhood information is obtained using the Transformer relation aggregator, and the relationship representation is updated through MTransH to reduce the noise impact caused by sparse neighborhoods.
Effectively capturing the representation of long-tail relationships in the knowledge graph improves the expression accuracy of small sample relationships, reduces the noise impact caused by sparse neighbors, and improves the accuracy of knowledge graph completion.
Smart Images

Figure CN116680414B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of knowledge graphs, and particularly relates to a knowledge graph prediction method based on an attention mechanism. Background Art
[0002] A knowledge graph (KG) was initially developed based on natural language processing (NLP) and can be used to query complex associated information more efficiently. A knowledge graph is very powerful in describing data. Compared with machine learning algorithms, it can fill the gap in data description of machine learning because machine learning algorithms are very strong in prediction but weak in description.
[0003] A knowledge graph is a semantic network used to describe various entities and concepts existing in the real world and the relationships between them. It describes the concepts, entities and the relationships between them in the objective world in a structured form. The information expression method is closer to the form of the human cognitive world and has the ability to better organize and manage massive information.
[0004] A knowledge graph is generally represented in the form of triples (head entity, relation, tail entity) and presented in a graph structure connected by nodes. Among them, the nodes represent entities and the edges represent relationships. Ideally, a knowledge graph can be used to enhance many other intelligent-dependent systems, such as information retrieval systems, question answering systems, recommendation systems, and reading comprehension, etc.
[0005] Generally speaking, real-world knowledge graphs are sparse and incomplete, and new instances will continuously emerge over time. However, the cost of manually finding all valid triples and adding new entity relationships is very high. Therefore, predicting the missing links in the generated knowledge graph has become a key research direction. This problem is also known as the Knowledge Graph Completion (KGC) task. The goal of the KGC task is to predict the missing part of an incomplete knowledge triple. Through analysis and training with a large amount of data, the tail entity is predicted based on the extracted relevant information to complete the triple, and this process is the knowledge graph completion work. The KG completion task can be divided into two types according to different triple completion targets: given the head entity and query relationship to predict the tail entity, or given the head entity and tail entity to predict the invisible relationship between them. During the KGC research process, a common problem quickly emerged in existing knowledge graphs. Most relationships in the knowledge graph have very few related triples. Generally, the trend is that the frequency of a relationship is inversely proportional to the proportion of uncommon entities associated with it, namely the so-called long-tail relationships. Previous work on completing knowledge graphs has focused on the most frequent entities and relationships, ignoring low-frequency relationships and uncommon entities. In fact, high-frequency relationships and entities are usually already sufficiently complete in the knowledge graph. Instead, those neglected low-frequency relationships and uncommon entities are more in need of completion to improve and enrich the knowledge network, which has inspired the research on Few-shot Knowledge Graph Completion (FKGC).
[0006] Since using previous knowledge graph completion methods to model and infer complex relationships, the model complexity is high, requiring a large number of training instances, and the relationships in the knowledge graph generally exhibit the phenomenon of long-tail distribution, which brings no small difficulty to knowledge graph completion. Therefore, the FKGC task has received increasing attention. Predicting the missing entity based on the incomplete triple where the task relationship of a small number of corresponding entity pairs is located not only relaxes the requirement for the amount of data but also solves the problem of the long-tail distribution of knowledge graph relationships. When the algorithms for FKGC learn the feature representation of task relationships using the information within the neighborhood, most of them equally treat the contributions of entities and neighbor relationships within the neighborhood. However, the fact is that the influences of different neighbor entities and corresponding neighbor relationships on entity and relationship representations are strong or weak, and the processing method with the same weight will surely affect the correctness of triple completion. In addition, in the face of the scenario of neighborhood sparsity, the model that uses multi-hop neighbor information to enhance the semantic representation of few-shot relationships may amplify the noise in the neighbor information, which will also reduce the accuracy of knowledge graph completion.
[0007] Most of the existing KGC models independently process the triples in KGs and rarely utilize the inherent and valuable information from the neighborhoods of entities. There are also some models that consider the impact of multi-hop neighbor information on few-shot relation representation and enhance the semantic representation of few-shot relations by capturing the hidden information in the neighborhood. However, when the neighborhood of a relation is too sparse or even has no neighbors, it is difficult to mine the hidden information and it may amplify the noise in the neighbor information. How to more effectively capture the useful information, especially the neighborhood information, in the knowledge graph is the research focus. Summary of the Invention
[0008] The purpose of the embodiments of the present invention is to provide a knowledge graph prediction method based on an attention mechanism, aiming to solve the problems raised in the above background technology.
[0009] The embodiments of the present invention are implemented as follows. A knowledge graph prediction method based on an attention mechanism includes a knowledge graph relation prediction model based on an attention mechanism (Knowledge Graph Relation Prediction ModelBased on Attention Mechanism, KGRPA). The model includes three modules: a neighborhood encoder, a relation aggregator, and a matching processor. The method includes the following steps:
[0010] Step 1: Load the training set triple data and obtain the few-shot relations and their corresponding head and tail entity pairs.
[0011] Step 2: Generate the neighbor sets of the head and tail entities of the few-shot relations.
[0012] Step 3: Use the graph attention network to aggregate the semi-neighborhood representations of the few-shot relations.
[0013] Step 4: Fuse the neighborhood information through Transformer to generate the few-shot relation representation.
[0014] Step 5: To judge the effectiveness of the few-shot relation representation, establish a scoring function formula based on TransH. The scoring function formula is as follows:
[0015] h ri = h i - h i W r h i , t ri = t i - t i W r t i
[0016] E(h i , r, t i ) = ||hri +r′ - t ri || L1 / L2
[0017] Among them, h i and t i are the head and tail embeddings learned during pre - training, W r is the normal vector of the hyperplane related to the relation r, and r′ is the general representation of the few - shot relation derived from the formula r = h - t;
[0018] Step 6, calculate the loss function, and the formula of the loss function is as follows:
[0019]
[0020] Among them, S r ={(h i , r, t i )} represents the support set of the few - shot relation r, and S′ r ={(h i , r, t′ i )} is the negative sample set generated by the tail entities that contaminate S r ;
[0021] Step 7, minimize the loss function through the gradient - descent method to train the model parameters, and repeat the above steps until the model effect reaches the optimal or the training reaches the maximum number of iterations.
[0022] For the further technical solution, in the neighborhood encoder part of the KFRPA model, for the triples of the few - shot relations in the support set, the designed gated and graph - attention - based neighbor aggregator first encodes the head entity and the tail entity with their neighbors to generate the neighbor representations of the head entity and the tail entity regarding the few - shot relation; when entering the relation aggregator, the neighborhood representations of all support sets are integrated by Transformer to learn the general representation of the few - shot relation; the neighborhood encoder and the relation aggregator can learn a good initialization of the relation representation from the background knowledge graph, and the matching processor uses the MTransH method to locally adjust the representation. Update the relation representation and hyperparameters. Finally, all parameters can be learned by transferring the updated relation representation and hyperparameters from the support set to the query set.
[0023] For the further technical solution, since KGRPA is a model based on deep - learning methods, after being trained by the above method, it is necessary to test whether the final actual effect reaches the expectation. The specific process includes:
[0024] First, load the test dataset and the trained model parameters;
[0025] Then, the effects of the MRR and Hit@N evaluation models are judged. The higher the values of both, the better the model effect.
[0026] A knowledge graph prediction method based on an attention mechanism provided by an embodiment of the present invention has the following beneficial effects:
[0027] (1) The one-hop neighborhood information of entities is fused through a graph attention network to obtain the neighborhood representation of entities, and an effective representation of long-tail relationships is obtained from the neighborhood representations of related entities. At the same time, a gating mechanism is introduced to reduce the noise impact caused by sparse neighborhoods.
[0028] (2) A Transformer relation aggregator is used to obtain a small-sample relation representation that fuses neighborhood information, and the matching processor updates the relation representation using MTransH to achieve a more accurate expression of small-sample relations. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a basic framework diagram of the KGRPA model in a knowledge graph prediction method based on an attention mechanism provided by an embodiment of the present invention;
[0030] Figure 2 It is a training flow chart of the KGRPA model in a knowledge graph prediction method based on an attention mechanism provided by an embodiment of the present invention;
[0031] Figure 3 It is an overall architecture diagram of the KGRPA model in a knowledge graph prediction method based on an attention mechanism provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0033] The following describes the specific implementation of the present invention in detail with reference to specific embodiments.
[0034] As Figure 2 shown, a knowledge graph prediction method based on an attention mechanism provided by an embodiment of the present invention includes a knowledge graph prediction model based on an attention mechanism (Knowledge Graph Relation Prediction Model Based on Attention Mechanism, KGRPA). The model includes three modules: a neighborhood encoder, a relation aggregator, and a matching processor. Its basic architecture is as Figure 1 shown, where N eDenote the neighborhood of entity e, Sr and Q r Denote the support set and query set of the few-shot relation r, r m Denote the vector representation of the few-shot relation r; The method includes the following steps:
[0035] Step 1, Load the training set triple data, and obtain the few-shot relations and their corresponding head and tail entity pairs;
[0036] Step 2, Generate the neighbor sets of the head and tail entities of the few-shot relations;
[0037] Step 3, Use the graph attention network to aggregate the semi-neighborhood representations of the few-shot relations;
[0038] Step 4, Fuse the neighborhood information through Transformer to generate the few-shot relation representation;
[0039] Step 5, To judge the effectiveness of the few-shot relation representation, establish a scoring function formula based on TransH, and the scoring function formula is as follows:
[0040] h ri =h i -h i W r h i ,t ri =t i -t i W r t i
[0041] E(h i ,r,t i ) = ||h ri +r′-t ri || L1 / L2
[0042] where h i and t i are the head and tail embeddings learned in the pre-training, W r is the normal vector of the hyperplane related to the relation r, and r′ is the general representation of the few-shot relation derived from the formula r = h - t;
[0043] Step 6, Calculate the loss function, and the loss function formula is as follows:
[0044]
[0045] where, S r ={(h i ,r,t i )} denotes the support set of the few-shot relation r, S r ′={(h i,r,t i ′)} is the negative sample set generated by the tail entity that pollutes S r .
[0046] Step 7: Train the model parameters by minimizing the loss function through the gradient descent method, and repeat the above steps until the model effect reaches the optimal or the training reaches the maximum number of iterations.
[0047] In the embodiment of the present invention, the specific structure of the overall framework of the KFRPA model is as Figure 3 shown. In the figure, (a), (b), and (c) respectively correspond to the neighborhood encoder, the relationship aggregator, and the matching processor. In the neighborhood encoder part, for the triples of small-sample relationships in the support set, the designed gated and graph attention-based neighbor aggregator first encodes the head entity and the tail entity with their neighbors to generate the neighbor representations of the head entity and the tail entity for the small-sample relationships. When entering the relationship aggregator, the neighborhood representations of all support sets are integrated through Transformer to learn the general representation of the small-sample relationships. The neighborhood encoder and the relationship aggregator can learn a good initialization of the relationship representation from the background knowledge graph, and the matching processor uses the MTransH method to locally adjust the representation. For the updated relationship representation and hyperparameters, finally, all parameters can be learned by transferring the updated relationship representation and hyperparameters from the support set to the query set.
[0048] As a preferred embodiment of the present invention, since KGRPA is a model based on deep learning methods, after being trained by the above method, it is necessary to test whether the final actual effect reaches the expectation.
[0049] The specific process includes:
[0050] First, load the test data set and the trained model parameters.
[0051] Then, evaluate the effect of the model through MRR (the mean of the reciprocals of the ranks of the correct triples in link prediction) and Hit@N (the average proportion of triples with ranks less than n in link prediction). The higher the values of the two, the better the model effect.
[0052] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A knowledge graph prediction method based on the attention mechanism, characterized in that, it includes a knowledge graph prediction model based on the attention mechanism, and the model includes three modules: a neighborhood encoder, a relation aggregator, and a matching processor; the method includes the following steps: Step 1, load the training set triple data, and obtain the small-sample relations and their corresponding head and tail entity pairs; Step 2, generate the neighbor sets of the head and tail entities of the small-sample relations; Step 3, use the graph attention network to aggregate the semi-neighborhood representations of the small-sample relations; Step 4, fuse the neighborhood information through Transformer to generate the small-sample relation representation; Step 5. To determine the effectiveness of the small sample relationship representation, a scoring function formula based on TransH is established, and the scoring function formula is as follows: ; ; Among them, h i and t i are the head and tail embeddings learned during pre-training, W r is the normal vector of the hyperplane related to the relationship r and is the general representation of the few-shot relationship derived from the formula r = h - t ; Step 6, calculate the loss function, and the formula of the loss function is as follows: ; Among them, represents the support set of the small sample relationship r , is the negative sample set generated by the tail entity of the contamination S r ; Step 7, minimize the loss function through the gradient descent method to train the model parameters, and repeat the above steps to optimize the model effect; In the neighborhood encoder part, for the triples of the small-sample relations in the support set, the designed gated and graph attention-based neighbor aggregator first encodes the head entity and the tail entity with their neighbors to generate the neighbor representations of the head entity and the tail entity of the small-sample relation; when entering the relation aggregator, the neighborhood representations of all support sets are integrated through Transformer to learn the general representation of the small-sample relation; the matching processor uses the MTransH method to locally adjust the representation, update the relation representation and hyperparameters, and transfer the updated relation representation and hyperparameters from the support set to the query set to learn all parameters.
2. The knowledge graph prediction method based on the attention mechanism according to claim 1, characterized in that, test whether the actual effect of the trained model reaches the expectation, and perform the following tests: First, load the test data set and the trained model parameters; Then, evaluate the effect of the model through MRR and Hit@N. The higher the values of the two, the better the model effect.
Citation Information
Patent Citations
Metalearning-based small sample knowledge graph completion method
CN115438192A
Risk prediction method and apparatus, and device and storage medium
WO2023065545A1