An Open World Knowledge Graph Completion Method and Device

By introducing Word2Vec module, attention module, Transformer module and CNN network into knowledge graph completion, the problems of poor results and high training costs in the open world knowledge graph completion in the existing technology are solved, and efficient knowledge graph completion and feature extraction are achieved.

CN114444694BActive Publication Date: 2025-06-13CHENGDU JINHAO BUILDING MATERIALS CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210070660.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-21
Publication Date
2025-06-13
Estimated Expiration
2042-01-21

AI Technical Summary

Technical Problem

The existing knowledge graph completion methods are mainly limited to the closed world and cannot fully utilize the resources of the open world, resulting in limited completion effects, high training costs and high requirements for the quality of the original data.

Method used

A open-world knowledge graph completion method is proposed, using Word2Vec module, attention module, Transformer module and CNN network, through word embedding, relationship-aware representation, global feature extraction and feature fusion, a knowledge graph completion model can be built, and knowledge can be obtained from open-world resources and complete the knowledge graph.

Benefits of technology

This method can effectively complete the missing data in the knowledge graph, improve triplet accuracy, reduce training costs, and perform good feature extraction while reducing the embedding size.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114444694B_ABST
    Figure CN114444694B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of open-world knowledge graph completion, and specifically relates to an open-world knowledge graph completion method and device, including obtaining triple data and performing word embedding, obtaining a relationship perception representation by an attention module and connecting it with a head entity vector, and obtaining a vector representation of the connection result by a Transformer; inputting a fusion result of the encoding problem vector and the vector representation of the connection result and the candidate vector representation obtained by the Transformer into a CNN network respectively; scoring the output of the CNN network, and taking the candidate tail entity with the highest score as the tail entity; training the model using a cross entropy loss function; obtaining a knowledge graph to be completed and inputting it into a trained model for completion. The present invention uses an attention mechanism and a Transformer network framework to make full use of feature information in a text description of an entity, reduce the cost of model training, and shorten the time of model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of open-world knowledge graph completion, and specifically relates to an open-world knowledge graph completion method and device. Background Art

[0002] There are two main storage methods for knowledge graphs: RDF and graph databases; the RDF language is a very simple language, which is essentially a triple composed of a subject, a predicate, and an object. The RDF language represents the associations between various things. Drawing out this association results in a very large graph, which is then transformed into a graph database. Google and Microsoft both have their own graph databases. Knowledge graphs have been applied in fields such as web search, link prediction, recommendation, natural language processing, and entity linking. However, most knowledge graphs are still imperfect. Denis Krompa conducted statistics on some open-source large knowledge bases. In Freebase, 71% of the entity's "place of birth" attribute values are missing for people, while this value is 66% in DBpedia. As the underlying database for many tasks and applications, the lack of data will seriously affect the performance of upper-layer applications.

[0003] To solve these problems, knowledge graph completion has been proposed to improve the knowledge graph by filling in the missing connections. Given a knowledge graph G=(E, R, T), where E represents the set of entities, R represents the set of relationships, and T represents the set of triples. Knowledge graph completion can be divided into two types: closed-world and open-world. The closed-world assumes that the knowledge graph is fixed and uses the topological structure of the graph to discover new relationships between existing entities and add new triples. The commonly used methods can be divided into three categories. The first category is the logic rule-based model, which infers new rules based on the existing triples through defined rules. The second category is the relationship path information-based model, which is a method of fusing the path information of the knowledge graph for path reasoning. Relationship path reasoning aims to utilize the path information in the knowledge graph structure and can improve the performance of the knowledge representation learning model. The third category is the embedding-based model, which maps entity vectors to the space determined by the relationship and then infers the missing relationship through vector operations.

[0004] However, the information that can be obtained by the methods of closed-world knowledge graph completion is limited, and more and more methods tend to obtain knowledge from open-world resources. To solve the problem of open-world knowledge graph completion, researchers have proposed models such as the ConMask model and the OWE model. The ConMask model proposed by Baoxu Shi first uses relation-based content masking to filter text information, delete irrelevant information, and only leave the content related to the task. Then, it uses a fully convolutional neural network to extract the embedding of the target entity from the relevant text. Finally, it compares this target entity embedding with the existing target candidate tail entities in the graph to generate a sorted list; however, this model does not fully utilize the rich feature information in the entity text description. Haseeb Shah et al. proposed the OWE model, which combines the conventional link prediction model learned from the knowledge graph and the word embedding learned from the text corpus. After independent training, it learns a transformation to map the embeddings of the entity name and description to the graph-based embedding space. This model utilizes the complete knowledge graph and does not depend on long texts, and has high scalability. However, the training cost of this model is high, and it has high quality requirements for the original data. Summary of the Invention

[0005] To solve the above problems, the present invention provides an open-world knowledge graph completion method and device.

[0006] An open-world knowledge graph completion method constructs a knowledge graph completion model, which includes a Word2Vec module, an attention module, and a scoring module. The open-world knowledge graph completion method includes the following steps:

[0007] S1. Obtain triple data, where each triple in the triple data includes a head entity description, a head entity name, a relation name, a candidate tail entity description, and a candidate tail entity;

[0008] S2. Use the Word2Vec module to perform word embedding on the head entity description and the candidate tail entity description to obtain a head entity vector and a candidate tail entity vector. Regard the text connection of the head entity name and the relation name as a question, and use the Word2Vec module to perform word embedding on the question to obtain a question vector;

[0009] S3. Use the attention module to calculate the head entity vector and the question vector to obtain a relation-aware representation;

[0010] S4. Connect the head entity vector with the relation awareness, and use Transformer to extract the global features of the connection result to obtain a vector representation of the connection result;

[0011] S5. Encode the question vector using a GRU network, fuse the encoded question vector with the vector representation of the connection result through a gating mechanism, and input the fusion result into a CNN network to obtain the first CNN output;

[0012] S6. Use Transformer to extract the global features of the candidate tail entity vector, obtain the candidate vector representation and input it into a CNN network to obtain the second CNN output;

[0013] S7. Score the first CNN output and the second CNN output through a scoring module and output the score;

[0014] S8. Calculate the loss value of the score using the cross-entropy loss function, and use the Adam optimization algorithm to train the parameters of the knowledge graph completion model until the model parameters converge;

[0015] S9. Obtain the knowledge graph to be completed and input it into the trained knowledge graph completion model for completion.

[0016] Furthermore, the triple data is obtained from the DBpedia50k dataset and the DBpedia500k dataset, and the triple data is divided into a training set, a validation set, and a test set in a ratio of 8:1:1.

[0017] Furthermore, add a label y to the triple data * indicating the correctness of the triple, that is, the correct triple label is 1, the wrong triple label is 0, and the label is represented as y * ∈{0,1}.

[0018] Furthermore, the attention function used in the attention module is:

[0019]

[0020]

[0021] where is the attention score, x represents the input word, Y represents the text, y i represents the i-th word in the text, m is the length of the text, w is a matrix, and α(·) is the ReLU non-linear activation function.

[0022] Furthermore, obtain the relation-aware representation corresponding to the head entity vector according to the attention function in the attention module

[0023]

[0024] where is the i-th word embedding in the head entity vector, is a set of problem vectors, and att(·) represents an attention function.

[0025] Furthermore, the fusion result of the encoded problem vector and the vector representation of the connection result in step S5 is expressed as:

[0026]

[0027] where σ is the sigmoid function, is the encoded problem vector, is the vector representation after extracting the global features of the connection result of the head entity vector and the relation perception using Transformer.

[0028] Furthermore, the scoring function adopted by the scoring module is expressed as:

[0029]

[0030] where, is the output of the first CNN, is the output of the second CNN, and W s is the transformation matrix to be trained.

[0031] Furthermore, the cross-entropy loss function is expressed as:

[0032]

[0033] where y i is the label value of the i-th triple, and y′ i represents the score of the i-th candidate tail entity output by the model, and m represents the total number of triples.

[0034] An open-world knowledge graph completion device includes:

[0035] An acquisition module for acquiring knowledge graph data to be completed;

[0036] A Word2Vec module for performing word embedding on the knowledge graph data in the acquisition module to obtain a head entity vector, a candidate tail entity vector, and a problem vector;

[0037] An attention module for calculating the head entity vector and the problem vector to obtain a relation perception representation;

[0038] A Transformer module for extracting the global features of the connection result of the head entity vector and the relation perception to obtain the vector representation of the connection result, and extracting the global features of the candidate tail entity vector to obtain the candidate vector representation;

[0039] A fusion module for fusing the encoded question vector with the vector representation of the connection result output by the Transformer module through a gating mechanism;

[0040] A CNN network for extracting features from the fusion result of the fusion module and the candidate vector representation of the Transformer module;

[0041] A scoring module for scoring the output result of the CNN network and selecting the triple corresponding to the highest score as the new triple to be supplemented into the knowledge graph.

[0042] Advantages of the present invention:

[0043] The present invention provides a method for open-world knowledge graph completion, which does not limit that the entities of the triples to be completed are all in the entity set of the knowledge graph to be completed, but obtains knowledge from open-world resources, such as online encyclopedias, and can complete various large-scale knowledge graphs to solve the problem of missing data in the knowledge graph.

[0044] The present invention mainly uses the Transformer network framework and the CNN network. Among them, the Transformer can well capture the global features of entity descriptions. The CNN network structure consists of 2 convolutional operations and 1 pooling operation, which can reduce the network training cost and can perform good feature extraction while reducing the embedding size, helping to improve the triple accuracy of knowledge graph completion. At the same time, the attention mechanism used can make full use of the information of text descriptions, and the GRU network used for questions can also improve the training efficiency while encoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is a flowchart of the method of the present invention;

[0046] Figure 2 It is a model structure diagram of the present invention;

[0047] Figure 3 It is a schematic diagram of the CNN network structure of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0049] An open-world knowledge graph completion method based on the attention mechanism and Transformer, as Figure 1, 2 As shown in 2 , a knowledge graph completion model is constructed. The model includes a Word2Vec module, an attention module, and a scoring module, and comprises the following steps:

[0050] S1. Obtain triple data. Each triple in the triple data includes a head entity description, a head entity name (also known as the head entity), a relation name (also called the relation), a candidate tail entity description, and a candidate tail entity;

[0051] S2. Use the Word2Vec module to perform word embedding on the head entity description and the candidate tail entity description to obtain a head entity vector and a candidate tail entity vector. Concatenate the text of the head entity name and the relation name as a question, and use the Word2Vec module to perform word embedding on the question to obtain a question vector;

[0052] S3. Use the attention module to calculate the head entity vector and the question vector to obtain a relation-aware representation;

[0053] S4. Concatenate the head entity vector and the relation awareness, and use Transformer to extract the global features of the concatenation result to obtain a vector representation of the concatenation result;

[0054] S5. Use a GRU network to encode the question vector, fuse the encoded question vector with the vector representation of the concatenation result through a gating mechanism, and input the fusion result into a CNN network to obtain a first CNN output;

[0055] S6. Use Transformer to extract the global features of the candidate tail entity vector to obtain a candidate vector representation and input it into a CNN network to obtain a second CNN output;

[0056] S7. Score the first CNN output and the second CNN output through the scoring module and output a score;

[0057] S8. Use a cross-entropy loss function to calculate the loss value of the score, and use the Adam optimization algorithm to train the parameters of the knowledge graph completion model until the model parameters converge;

[0058] S9. Obtain the knowledge graph to be completed and input it into the trained knowledge graph completion model for completion.

[0059] The Freebase 15K dataset is widely used in knowledge graph completion. However, FB15K is full of a large number of reverse triples or synonym triples and does not provide enough text information for text description-based knowledge graph completion methods.

[0060] In this embodiment, due to the limited text content and redundancy in the FB15K dataset, two new datasets, DBPedia50k and DBPedia500k, are used for open-world knowledge graph completion; the DBPedia50k dataset contains 49,900 entities, with an average description length of 454 words per entity and 654 relations. The DBPedia500k dataset contains 517,475 entities and 654 relations. The complete triples in the obtained datasets are divided into training set, validation set, and test set datasets at a ratio of 8:1:1.

[0061] Word2Vec is a word vector representation that converts text into a set of vectors with the help of a dictionary;

[0062] Use the Word2Vec module to perform word embedding on the head entity description and candidate tail entity description to obtain the head entity vector and candidate tail entity vector, denoted as:

[0063]

[0064]

[0065] where h i is the i-th word in the head entity description, is the i-th word embedding in the head entity vector, |M h | is the length of the head entity description, t n is the n-th word in the candidate tail entity description, is the n-th word embedding in the candidate tail entity vector, |Z t | is the length of the candidate tail entity description.

[0066] Regard the text connection of the head entity name and the relation name as a question, and use the Word2Vec module to perform word embedding on the question to obtain the question vector, denoted as:

[0067]

[0068] where r j is the j-th word in the question, is the j-th word embedding in the question vector, |Q r | is the length of the question.

[0069] The head entity name is a word, and the head entity description is a piece of text containing the name. The representations of each word in the head entity description are not equally important. For the relationship and each word in the head entity description text, there are words closely related to the word "relationship" in the head entity description, as well as many irrelevant words. Therefore, the attention mechanism is used for the head entity vector and the question vector to emphasize the information related to the relationship in the head entity description, obtaining a relationship-aware representation of the words in the head entity description, which is equivalent to reducing the representations of those irrelevant words and removing noise.

[0070] Preferably, define the attention function adopted in the attention mechanism as:

[0071]

[0072] Wherein, is the attention score, given the input word x and the text m is the length of the text;

[0073]

[0074] Where the attention score captures the similarity between the given input word x and each word y in the text Y, w is a matrix, and α(·) is the ReLU non-linear activation function. According to the defined attention function, the relationship-aware representation i corresponding to the head entity vector can be obtained. The formula is: The formula is:

[0075]

[0076] Connect the head entity vector without attention operation with the relationship-aware representation obtained through attention operation to get a new head entity vector

[0077] To better capture long-term dependency relationships and extract global features, input it into the Transformer encoder for encoding to obtain

[0078] Then input the candidate tail entity vector into the Transformer encoder as well to obtain

[0079] GRU is a type of recurrent neural network, which can solve problems such as long-term memory and gradients in backpropagation, and is easier to train compared to LSTM, which can greatly improve the training efficiency. Input the question vector​ Encoding context information in relevant texts using a GRU network to obtain

[0080] To fuse the head entity vector with the question vector A gating mechanism is used for fusion to obtain the target entity embedding R s , and the formula is:

[0081]

[0082] where σ is the sigmoid function.

[0083] Convolutional neural networks can solve the overfitting problem of deep structures. At the same time, CNN networks are also often used in the field of knowledge graph completion and have achieved good performance. The CNN network adopted in the present invention is as Figure 3 shown, consisting of two convolutional layers, a pooling layer, and a fully connected layer. Its network structure is to perform a max pooling operation after two 3×3 convolutional operations and then connect to the fully connected layer. Specifically, an input image of 572×572 is fed into the CNN network. After the first 3×3 convolution, a first feature map of 570×570 is obtained. The first feature map is fed into the second 3×3 convolution to obtain a second feature map of 568×568. The second feature map is subjected to max pooling to obtain a third feature map of 284×284, and finally fed into the fully connected layer.

[0084] Using the CNN network as the target entity fusion structure, the target entity embedding R s and are respectively input into the CNN network to obtain and

[0085]

[0086]

[0087] In the scoring module, a scoring function is used to score and , and its scoring function is expressed as:

[0088]

[0089] where W s is the transformation matrix to be trained, and · T represents the transpose operation. After passing through the Score(·) function, each candidate tail entity has its corresponding score s i , and the candidate tail entity corresponding to the highest score is used as the correct tail entity.

[0090] In this embodiment, during the training phase of the model, a score s corresponding to each candidate tail entity needs to be output i , and the output of the set model score is expressed as:

[0091] y′ = softmax([s 1 ; s 2 ; …; s m );

[0092] Preferably, during the training process of the knowledge graph completion model, labels y * are added to the training data. The label represents the correctness of the triples in the training data, that is, the correct triple label is 1, and the wrong triple label is 0. The cross-entropy loss function is used to minimize the gap between the predicted triples and the correct triples. The formula of the cross-entropy loss function is as follows:

[0093]

[0094] where y i is the one-hot encoding of the label, y i is the label value of the i-th triple, and y′ i represents the score of the i-th candidate tail entity output by the model.

[0095] Preferably, the Adam algorithm is used to minimize the loss function to optimize the model. Adam is a first-order optimization algorithm that can replace the traditional stochastic gradient descent process. It can iteratively update the neural network weights based on the training data. Adam designs independent adaptive learning rates for different parameters by calculating the first-order moment estimate and second-order moment estimate of the gradient. The main calculation formula is as follows:

[0096]

[0097] where represents the corrected first-order moment estimate and second-order moment estimate, and ∈, η are parameters to be adjusted during the training process.

[0098] Preferably, after the training of the knowledge graph completion model is completed, the model is evaluated. The evaluation indicators of the model are MRR, MR, Hits@1, Hits@3, Hits@10. For each test triple, the tail entity is predicted. By scoring all the candidate tail entity descriptions and then arranging these scores in ascending order. Hits@10 is the probability that the correct triple ranks in the top 10. Similarly, Hits@3 is the probability of ranking in the top 3, and Hits@1 is the probability of ranking first.

[0099] MR is the average rank, that is, the average of the ranks of the correct triples.

[0100]

[0101] where t i is the true ranking of the i-th triple.

[0102] MRR is the mean reciprocal rank. That is, if the correct triple is ranked at the k-th position, then MRR is

[0103]

[0104] k i is the correct ranking of the i-th triple.

[0105] An open-world knowledge graph completion device, comprising:

[0106] An acquisition module, configured to acquire knowledge graph data to be completed;

[0107] A Word2Vec module, configured to perform word embedding on the knowledge graph data in the acquisition module to obtain a head entity vector, a candidate tail entity vector, and a question vector;

[0108] An attention module, configured to calculate the head entity vector and the question vector to obtain a relationship-aware representation;

[0109] A Transformer module, configured to extract global features of the connection result of the head entity vector and the relationship-aware connection result to obtain a vector representation of the connection result, and extract global features of the candidate tail entity vector to obtain a candidate vector representation;

[0110] A fusion module, configured to fuse the encoded question vector with the vector representation of the connection result output by the Transformer module through a gating mechanism;

[0111] A CNN network, configured to perform feature extraction on the fusion result of the fusion module and the candidate vector representation of the Transformer module;

[0112] A scoring module, configured to score the result output by the CNN network, and select the triple corresponding to the highest score as a new triple to be supplemented into the knowledge graph.

[0113] Specifically, the knowledge graph data to be completed acquired by the acquisition module is the knowledge graph G = {E, R, F}, where E represents the set of all entities, R represents the set of all relationships, F is the set of all triples, each triple includes a head entity description, a head entity name, a relationship name, a candidate tail entity description, a candidate tail entity, and the text connection of the head entity name and the relationship name is regarded as a question. Applying this model in the knowledge graph to be completed, the triple corresponding to the candidate tail entity with the highest score is selected as the correct triple, and one group is completed each time.

[0114] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An open-world knowledge graph completion method, characterized in that, a knowledge graph completion model is constructed, and the model includes a Word2Vec module, an attention module, and a scoring module. The open-world knowledge graph completion method includes the following steps: S1. Obtain triple data, where each triple in the triple data includes a head entity description, a head entity name, a relationship name, a candidate tail entity description, and a candidate tail entity; S2. Use the Word2Vec module to perform word embedding on the head entity description and the candidate tail entity description to obtain a head entity vector and a candidate tail entity vector. Concatenate the text of the head entity name and the relationship name as a question, and use the Word2Vec module to perform word embedding on the question to obtain a question vector; S3. Use the attention module to calculate the head entity vector and the question vector to obtain a relationship-aware representation; S4. Concatenate the head entity vector with the relationship awareness, and use Transformer to extract the global features of the concatenation result to obtain a vector representation of the concatenation result; S5. Use a GRU network to encode the question vector, fuse the encoded question vector with the vector representation of the concatenation result through a gating mechanism, and input the fusion result into a CNN network to obtain a first CNN output; S6. Use Transformer to extract the global features of the candidate tail entity vector to obtain a candidate vector representation and input it into a CNN network to obtain a second CNN output; S7. Score the first CNN output and the second CNN output through a scoring module and output a score; S8. Use a cross-entropy loss function to calculate the loss value of the score, and use the Adam optimization algorithm to train the parameters of the knowledge graph completion model until the model parameters converge; S9. Obtain the knowledge graph to be completed and input it into the trained knowledge graph completion model for completion.

2. The open-world knowledge graph completion method according to claim 1, characterized in that, the triple data is obtained from the DBpedia50k dataset and the DBpedia500k dataset, and the triple data is divided into a training set, a validation set, and a test set dataset in a ratio of 8:1:

1.

3. The open-world knowledge graph completion method according to claim 2, characterized in that, Add the label y to the triple data * Indicates the correctness of the triple, that is, the correct triple label is 1, the wrong triple label is 0, and the label is represented as y * ∈{0,1}.

4. The open-world knowledge graph completion method according to claim 1, characterized in that, the attention function used in the attention module is: Among them, is the attention score, x represents the input word, Y represents the text, and y i represents the i-th word in the text, m is the text length, w is a weight matrix, and α(·) is the ReLU non-linear activation function.

5. The open-world knowledge graph completion method according to claim 4, characterized in that, Obtain the relation-aware representation corresponding to the head entity vector according to the attention function in the attention module Among them, is the i-th word embedding in the head entity vector, is the set of problem vectors, and att(·) represents the attention function.

6. The open-world knowledge graph completion method according to claim 1, characterized in that, the fusion result of the encoded question vector and the vector representation of the concatenation result in step S5 is expressed as: where σ is the sigmoid function, is the encoded question vector, is the vector representation after extracting global features using Transformer for the connection result of the head entity vector and the relation perception.

7. The open-world knowledge graph completion method according to claim 1, characterized in that, the scoring module uses a bilinear scoring function expressed as: Among them, is the first CNN output, is the second CNN output, and W s is the transformation matrix to be trained.

8. The open-world knowledge graph completion method according to claim 1, characterized in that, the cross-entropy loss function is expressed as: where y i is the label value of the i-th triple, and y' i represents the score of the i-th candidate tail entity output by the model, and m represents the total number of triples.

9. An open-world knowledge graph completion device, characterized in that, comprising: An acquisition module for acquiring knowledge graph data to be completed; A Word2Vec module for performing word embedding on the knowledge graph data in the acquisition module to obtain a head entity vector, a candidate tail entity vector, and a question vector; An attention module for calculating the head entity vector and the question vector to obtain a relation-aware representation; A Transformer module for extracting the global features of the connection result of the head entity vector and the relation-aware connection result to obtain a vector representation of the connection result, and extracting the global features of the candidate tail entity vector to obtain a candidate vector representation; A fusion module for fusing the encoded question vector with the vector representation of the connection result output by the Transformer module through a gating mechanism; A CNN network for extracting features from the fusion result of the fusion module and the candidate vector representation of the Transformer module; A scoring module for scoring the result output by the CNN network and selecting the triple corresponding to the highest score as a new triple to be supplemented into the knowledge graph.