A Knowledge Graph Completion Method and System Based on Unstructured Information

Through the cooperation between graph neural networks and adversarial neural networks, unstructured information is used to generate free text data, and through adversarial training and graph neural network scoring, data noise and sparse problems in knowledge graph completion tasks are solved, improving the efficiency and accuracy of completion.

CN113934847BActive Publication Date: 2025-06-13DAREWAY SOFTWARE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111226461.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-21
Publication Date
2025-06-13
Estimated Expiration
2041-10-21

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively use unstructured information to complete the knowledge graph completion task, resulting in increased data noise and decreased training accuracy. The free text containing the target entity in the Internet is sparse, making it more difficult to complete.

Method used

The method of cooperation between graph neural networks and adversarial neural networks is adopted to obtain missing triple data to be completed, identify entity nodes, generate free text data, and complete the knowledge graph through adversarial training between generators and discriminators, combined with graph neural network scoring.

Benefits of technology

The efficiency and accuracy of knowledge graph completion are improved, the calculation efficiency is improved through pruning processing, and the accuracy of entity node scores is improved through secondary verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113934847B_ABST
    Figure CN113934847B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for knowledge graph completion based on unstructured information, which identifies entity nodes in missing triple data; obtains sentences associated with the entity nodes, identifies entity triples in the sentences, and simultaneously inputs the obtained sentences into a generator to generate free text data; combines the free text data and structured text data for training the generator, and a discriminator discriminates the entity triple prediction result of the generator according to the entity triples in the sentences to perform adversarial training between the generator and the discriminator; when the discriminator passes the discrimination, adds the entity triples in the sentences to the knowledge graph, scores using a graph neural network, and combines the scoring result and the information of the previous node of the known entity nodes of the missing triples to obtain the ranking result of the entity triples, thereby completing the completion of the knowledge graph; the present invention improves the efficiency and accuracy of knowledge graph completion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graph completion processing, and particularly relates to a method and system for knowledge graph completion based on unstructured information. Background Art

[0002] The statements in this part merely provide background art related to the present invention and do not necessarily constitute prior art.

[0003] The essence of a knowledge graph is a semantic web, and in recent years, knowledge graphs have developed rapidly and have a wide range of applications in real life, such as fact answering systems based on large knowledge bases such as DBpedia, YAGO, Freebase, etc. However, in real life, knowledge is always linked one by one, so a lot of knowledge is not available in the original knowledge base, or can be said to be incomplete. Therefore, the task of knowledge graph completion has gradually attracted people's attention. Generally speaking, a knowledge graph consists of three parts: a head entity, a relationship entity, and a tail entity (h, r, t), and knowledge graph completion is when one of the head and tail entities is known along with the relationship entity, and then the other entity is completed according to a certain relationship or information.

[0004] Now many projects or texts contain many entities in the knowledge base, and some texts in the real world contain a lot of information. How to effectively utilize this information and how to use this unstructured information to complete the knowledge graph completion task are the current research focuses and hotspots.

[0005] The inventors found that if some unstructured data and some structured data collected are simply put together, it may not be helpful for the knowledge graph completion task, but may instead increase data noise and reduce its training accuracy. Therefore, it is not feasible to directly fuse some free texts on the Internet with some structured data; and on the Internet, there are not many free texts containing the target entities required, or the structured data itself is not much, which is relatively sparse, making it difficult to complete the knowledge graph completion. Summary of the Invention

[0006] In order to solve the deficiencies of the prior art, the present invention provides a method and system for knowledge graph completion based on unstructured information, which uses a graph neural network and an adversarial neural network to cooperate with each other to complete the knowledge graph completion task, improving the efficiency and accuracy of knowledge graph completion.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] The first aspect of the present invention provides a method for knowledge graph completion based on unstructured information.

[0009] A method for knowledge graph completion based on unstructured information, comprising:

[0010] Obtaining missing triple data to be completed;

[0011] Identifying entity nodes in the missing triple data;

[0012] Obtaining sentences associated with the entity nodes, identifying entity triples in the sentences, and at the same time inputting the obtained sentences into a generator to generate free text data;

[0013] Combining the free text data and the structured text data to train the generator, and the discriminator discriminates the entity triple prediction result of the generator according to the entity triples in the sentences, and performs adversarial training between the generator and the discriminator;

[0014] When the discriminator passes the discrimination, adding the entity triples in the sentences to the knowledge graph, scoring using a graph neural network, and combining the scoring result and the information of the previous node of the known entity node of the missing triple to obtain the ranking result of the entity triples, and completing the completion of the knowledge graph.

[0015] The second aspect of the present invention provides a knowledge graph completion system based on unstructured information.

[0016] A knowledge graph completion system based on unstructured information, comprising:

[0017] A data acquisition module, configured to: obtain missing triple data to be completed;

[0018] An entity node identification module, configured to: identify entity nodes in the missing triple data;

[0019] A sentence acquisition module, configured to: obtain sentences associated with the entity nodes, identify entity triples in the sentences, and at the same time input the obtained sentences into a generator to generate free text data;

[0020] An adversarial training module, configured to: combine the free text data and the structured text data to train the generator, and the discriminator discriminates the entity triple prediction result of the generator according to the entity triples in the sentences, and performs adversarial training between the generator and the discriminator;

[0021] A knowledge graph completion module, configured to: when the discriminator passes the discrimination, add the entity triples in the sentences to the knowledge graph, score using a graph neural network, and combine the scoring result and the information of the previous node of the known entity node of the missing triple to obtain the ranking result of the entity triples, and complete the completion of the knowledge graph.

[0022] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the steps in the method for completing a knowledge graph based on unstructured information as described in the first aspect of the present invention are implemented.

[0023] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, the steps in the method for completing a knowledge graph based on unstructured information as described in the first aspect of the present invention are implemented.

[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0025] 1. The present invention innovatively proposes a method and system for completing a knowledge graph based on unstructured information, and uses a graph neural network and an adversarial neural network to cooperate with each other to complete the knowledge graph completion task, improving the efficiency and accuracy of knowledge graph completion.

[0026] 2. The present invention innovatively proposes a method and system for completing a knowledge graph based on unstructured information. Through pruning processing, the computing efficiency is improved, and through secondary verification, the accuracy of entity node scoring is improved.

[0027] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0029] Figure 1 It is a schematic flow chart of the method for completing a knowledge graph based on unstructured information provided in Embodiment 1 of the present invention.

[0030] Figure 2 It is a schematic diagram showing the change of the accuracy of the model provided in Embodiment 1 of the present invention with the increase of epochs in different sparse free text data sets. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0032] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0033] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0034] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0035] Embodiment 1:

[0036] Embodiment 1 of the present invention provides a method for knowledge graph completion based on unstructured information, which uses an adversarial neural network and a graph neural network to cooperate with each other to complete the final knowledge graph completion task.

[0037] As described in the background art, at present, there are many data that are different from the data in the knowledge graph. However, there is a large amount of information stored in a lot of free texts on the Internet, which can not only help the human brain make accurate judgments, but also help machines make some judgments if properly utilized. For example, for knowledge graph completion. However, if such free texts are directly put into the dataset for the knowledge graph completion task, it is very likely that they will be of no help and may even increase a certain amount of noise. Considering the characteristics of heterogeneity and the like of various data, this embodiment designs an algorithm that combines adversarial learning and graph neural networks to use free texts to complete the knowledge graph completion task.

[0038] In this model, several components with different purposes are set for the knowledge graph completion task, one is an entity prediction component (i.e., a generator), an evaluator component (i.e., a discriminator), and an extractor for free texts.

[0039] Among them, the adversarial network has two functions. First, after the generator is trained and missing entity data is input, it will generate an entity as its prediction result, and the function of the discriminator is to evaluate the credibility of the answer generated by the generator. These two components, the generator and the discriminator, promote and supervise each other. The purpose is to make the generator and the discriminator trained to be powerful enough. Only in this way can the discriminator better enable the generator to use the knowledge graph information to produce more credible answers.

[0040] The second function is to generate some free texts. Because when extracting information from unstructured texts, not too much text can be extracted. However, if not enough text can be extracted, the reasoner will not learn enough features to perform reasoning. Therefore, it is necessary for the adversarial neural network to generate some free texts with relatively high authenticity to help the machine train.

[0041] The overall process of the entire technology is as follows. After the missing triples are input into the model, the known entities and relationships of the triples are respectively input into the knowledge graph and unstructured text. For the unstructured text, it will collect three types of sentences and identify the entity nodes in the sentences. While identifying the nodes, the collected sentences will be input into the generator, which will generate some data according to the input sentences to solve the problem of data sparsity. After solving the data sparsity problem, the known structured data will be input to retrain the discriminator and the generator. Then, the entity triples identified from the collected sentences will be input into the discriminator for discrimination. If the discriminator passes the discrimination, the triples will be input into the knowledge graph, scored using a graph neural network, rank the entities, and combine the information of the previous node of the known nodes. Finally, an entity result ranking is obtained to complete the knowledge graph completion task.

[0042] After collecting the text, it is necessary to perform a structured modeling on the entire collected free text, that is, construct these texts into a knowledge graph, and then proceed to the next step of the entire knowledge graph completion task.

[0043] Entity - Node: This sub - graph inherits the entities in the sentences collected from unstructured information as nodes. Or the entity nodes of other corpora can also be used. Or the known entities of the triples that need to be completed in the original knowledge graph can be used as nodes.

[0044] Collection of Free Text - Edge: When given a knowledge graph completion task, it is actually to query the missing entity in (s, r,?). First, search for sentences in which another entity and relationship appear simultaneously in the unstructured text (of course, other corpora like ClueWeb can also be used). Then, find the word that co - occurs most frequently with this entity s and relationship r, and regard this word as the candidate answer for the missing entity.

[0045] However, if only one such sentence is selected, it is very difficult to ensure the accuracy of this answer. And in real life, relying solely on one sentence cannot determine the true answer. But blindly searching will cause noise. Therefore, TagMe will be applied to the search page of unstructured text to find another entity mentioned in the free text, and select and collect some texts that meet the following conditions:

[0046] (1) Search for sentences containing r on the page where s appears in the unstructured text.

[0047] (2) Search for sentences containing s on the page where r appears in the unstructured text.

[0048] (3) Sentences containing both s and r appear simultaneously in the unstructured text page.

[0049] The function of the method mentioned in this embodiment is to complete the knowledge graph, but it does not supplement the relationships. Assuming that in the knowledge graph or triples, the default relationships are known, such as (h, r,?), the task of this embodiment is to complete the "?".

[0050] After collecting the sentences, entity recognition will be performed based on the collected sentences to identify the entities in the sentences and label them.

[0051] Completion problem handling: So far, the general components of the graph have been roughly described. Next is to handle the entity completion problem. When a question (s, r,?) is posed to this subgraph, an answer needs to be found in this graph and added to the list of candidate entities, and finally scored.

[0052] First, mark the entities in the incoming question Mq as Xq = {Vs|Vs ∈ Mq}. Assuming that the entity s or the relationship r appears in the free text, TagMe will be used to identify and mark the entity nodes in the sentence.

[0053] The ultimate goal is to find the candidate entity nodes. In the constructed free text knowledge graph, the entity nodes related to s and r are the candidate entity nodes to be found.

[0054] Assume that entities that appear within five words related to s, r, or both s and r in the collected free text will be added to the candidate list.

[0055] Secondary verification: This embodiment will collect three types of sentences on the web page. The alternative entities identified in these three sentences are marked as Em = {Pm|Pm ∈ Ps ∪ Pr ∪ Psr}. For example, if the entity Es is marked, it means that this entity appears on the page related to S. Then this entity will be listed as an alternative entity and compared with the related sentences on Pr and Psr to see if there are common entities. If it only appears on one page, the alternative score Se of this entity will increase accordingly:

[0056]

[0057] The denominator +1 is to prevent the denominator from being 0. If the same entity appears on the remaining two pages, its score will be increased accordingly. Conversely, if it does not appear on both, its ranking may be relatively lower. Of course, this is not absolute.

[0058] Graph Pruning: The finally generated graph must ensure high coverage of all texts. Because if there is little free text, it is easy to have insufficient number of nodes, resulting in too high error rate in the case of sparse texts. Therefore, this embodiment introduces an adversarial neural network to solve this problem.

[0059] However, in many cases, the number of free texts is sufficient. If the number of nodes is too large and the sub-graph is too big, then too much time will be consumed in the training, inference or verification part. Therefore, a simple pruning is needed. First, filter out some nodes with too low scores, which will improve the computing efficiency.

[0060] Regarding the filtering of all candidate answer entity nodes as a sorting problem, after given candidate nodes and the above-mentioned secondary verification, each node will get a relevance score. During the whole inference process, the top twenty nodes with the highest scores (if there are not enough alternative entity nodes, then the maximum value will be taken) will be retained, and the rest will be pruned.

[0061] When given a missing triple problem (e, r,?), entities in the free text are expected to be collected. To achieve this goal, a reverse aggregation process is also needed, as follows:

[0062]

[0063] where, γ ne ={(ne, r, nk)| dne = dnk + 1, (ne, r, nk) ∈ Q} represents the previous node of the entity node in the original knowledge graph. Because in the actual inference process, sometimes the previous entity node related to the entity node also contains a lot of information.

[0064] When the entire network updates information in the feed-forward network layer, there will be a score for several words before and after each candidate entity collected from the free text. The higher this score, the greater the influence of this word on the selection of the candidate entity answer.

[0065] The representation of the entity is not directly used to score the candidate entity nodes, but another entity node h of the missing triple vq and each entity node h in the sentence e are concatenated together, and then fed through its feed-forward network, and then weighted by the entity scores:

[0066]

[0067] Therefore, the update of each candidate entity node combines the nodes in the upper layer of the knowledge graph, that is, the information of the other known entity of the missing triple, and in addition to this information, it also includes the information in the collected free text:

[0068]

[0069] After the entire model has undergone several iterations and loop training, each candidate node then contains the information of the previous node and the information in the free text. Then, the two pieces of information are aggregated, and finally the answer is obtained.

[0070] Entity discriminator: In the method proposed in this embodiment, the main function of the entity discriminator is to be able to distinguish true and false answers in a given query situation. Compared with the previous methods for knowledge graph completion tasks based on adversarial neural networks, the difference in this embodiment is to utilize the information in the text and then combine the entity representation information in the free text to enhance the discrimination ability of the entity discriminator.

[0071] Discriminator D(t|h, r; G, θ D ) will evaluate whether this entity node can be the entity node answer to the question (e, r,?) through the score of the following formula:

[0072]

[0073] In this score evaluation formula, s(·) is a scoring function for evaluating the completed triple (e, r, o) to evaluate whether this triple is true. In TransE and DistMult, they both have their own unique ways to instantiate this function. In the formula of this embodiment, in order to ensure the speed of training, testing, and verification, its general form is used.

[0074] To distinguish from general adversarial neural networks, this embodiment combines the entity information in the free text to enhance the representation, thereby improving the discrimination ability of the entity discriminator:

[0075]

[0076] Among them, w 1 、w 2 and b 1 、b 2 are parameter matrices and bias vectors, and tanh(·) is incorporated as a non-linear transformation function. In this way, the influence of the candidate entity in the sentence on the entity in the original knowledge graph is input into the discriminator. Finally, when the discriminator makes an entity judgment, it will comprehensively consider various aspects of information, then make a judgment, and then feedback to the generator.

[0077] To optimize the entity discriminator, it is necessary to consider its loss function accordingly. First, the discriminator should recognize the true answer o of the triple query (e, r,?) on the original knowledge graph as positive, and then judge the entity G(h, r, θ G ) generated by the generator, rather than making a direct judgment. The overall expression is as follows:

[0078]

[0079] Among them, λ D > 0 is set for two purposes. The first purpose is to prevent overfitting during training. The second purpose is that when a triple question query is given, the entities in the original knowledge graph should be judged as correct, while the nodes generated by the generator will be judged as wrong. Then, as the samples generated by the generator get better and better, the discriminator will also optimize its parameters according to the samples generated by the generator to make its discrimination ability more powerful.

[0080] Entity generator: The generator has three functions in this model. The first function is that the generator is trained when receiving incoming structured data, and then receives the missing triples, generates another entity for the missing entity or the newly added entity in the triples. The generated entity is used as the prediction result of the generator and then passed into the entity discriminator for scoring, and then updated according to the feedback of the discriminator, and then generates new entities again, iterating in this way, which can not only promote the training of the discriminator, but also promote the ability of the generator itself.

[0081] The second function is that when the model collects text, the number of free texts collected is not sufficient, and there is not enough data volume for the machine to identify features, resulting in overfitting. When free texts are collected, according to the known information in the missing triples, and then passed into the generator, the generator will generate some free texts according to the received known information, so that the model can learn more features and complete the overall work.

[0082] The third function of the entity generator is to improve the ability of the discriminator, which is also its most important function. Because when the ability of the discriminator is high enough, it can complete a preliminary judgment and conduct a preliminary screening process for most wrong nodes, which will greatly save the time of training and inference.

[0083] For each triple missing problem (e, r,?), this embodiment will set a candidate entity answer set C e,r, the generator then defines its candidate entity set and sample distribution. When given a triple query problem Q=(e, r,?), the representation V of this problem is calculated q as:

[0084]

[0085] where, v e and v r are the representations of the initialized nodes and relationships:

[0086]

[0087] where, n e and n k both represent entity nodes on the graph. F ne represents the set of entity nodes linked to the previous entity node in the knowledge graph. Finally, all vectors are concatenated and then fed into a multi-layer perceptron (MLP), and then activated by the LeakyReLU function. In the entity generator model of this embodiment, the probability distribution for entity sampling of the candidate entity answer set C e,r is defined as:

[0088]

[0089] Then, using this distribution, the candidate answer set is sampled and then input into the discriminator. Then, the discriminator makes a judgment, the generator receives the reward value feedback from the discriminator, and then changes its strategy and next action according to the reward value.

[0090] When optimizing the loss of the generator, the strategy gradient method is used to learn the parameters. The most important thing is how to set a reasonable reward function. The entity generator in this embodiment uses the feedback signal of the discriminator as the reward of the generator, and then promotes the generator to perform the next step of learning, and then generates more realistic data:

[0091]

[0092] where, the scoring function s(·) has been defined, and the bias of uniform sampling is introduced as a reference. Then, when the probability that the generated sample is accepted by the entity discriminator is greater than the average value, a higher-scoring reward will be assigned to it. In this way, the generator and the discriminator will start to compete with each other. The generator wants to change its strategy by obtaining a higher reward based on the reward score given by the discriminator to itself, while the discriminator tries its best to distinguish the false data of the incoming data to improve its discrimination ability.

[0093] Knowledge Graph Completion: When the model receives a question, the generator will first train based on the collected sentences to solve the problem of data sparsity, and then train based on structured data to enhance the discriminative ability of the discriminator. When the discriminator receives the triples extracted from unstructured information, it will make judgments one by one. After passing through the discriminator, the triples will be passed into the knowledge graph, and then a graph neural network will be used for entity ranking. Then, the information of the adjacent nodes of the known nodes of the triples will be combined to finally obtain the ranking of entity nodes, completing the knowledge graph completion task.

[0094] In the entire dataset setting, the data of the knowledge graph and free text need to be in a state of mutual dependence. In the entire experiment, the KB4Rec dataset was used to construct the evaluation dataset.

[0095] KB4Rec is a dataset obtained by connecting various other datasets. It not only includes the datasets of the knowledge bases of Freebase and YAGO, but also some text datasets including AMAZON book, MOVIELENS movie, and LFM-1B music. These datasets not only include the information and attributes of each entity, but also a large amount of free text such as comments. However, due to the excessive size of these datasets, subsets were taken for several of them. For LFM-1B music, the subset of 2011 was used in this embodiment. For MovieLens, the subset of 2005 - 2012 was used in this embodiment.

[0096] In the real world, there are many data for which it is impossible to search for enough evidence for the machine to learn enough features for reasoning or completing the knowledge graph completion task. Therefore, after obtaining these datasets, it is necessary to perform some sparsity processing on the data in the original knowledge graph, that is, structured data. The following formula is used to evaluate the sparsity of each entity:

[0097]

[0098] Spa(e) represents the number of times or frequency of entity e appearing in all structured data. The value range is defined as [0, 1]. The higher the spa value, the higher the sparsity of this entity. When spa(e) = 1, it means the entity has the highest sparsity, otherwise the sparsity is lower. Then, by changing the validation and test datasets, sparse versions of these data are generated, and these datasets are used for evaluation. In the valid triples (s, r, o) in the test and validation datasets, if one or both of the two entities s or o are sparse entities, they will be retained, otherwise they will be filtered out to ensure data sparsity.

[0099] The statistical data of the processed dataset are summarized in Table 1. The sparsified datasets are defined as Spa-Movie, Spa-Music, and Spa-Book respectively. Then, the data of each different entry will be randomly divided into a training set, a validation set, and a test set.

[0100] Table 1: Statistical data of the processed dataset.

[0101]

[0102] Experimental results: TransE, Dismult, Conve, ConvTransE, R-GCN, KBGAN, CoFM, and KGAT were respectively experimented and compared on these three datasets, and the experimental results are given in Table 2.

[0103] Table 2: Comparative experiments of the model in this embodiment with other various methods. Except for MRR, other results are percentages (%).

[0104]

[0105] Table 2 shows that the method and model proposed in this embodiment are significantly superior to other models, and the following points can be observed:

[0106] (1) The performance of TransE is the worst among these models because the method it adopts is only a distance function. Then, based on semantic matching, methods such as DistMult, ConvE, and ConvTransE perform better than TransE because they use matching functions to model the semantics of triples. The performance of the GNN-based method R-GCN is also much higher than that of TransE, but it is similar to other several baseline methods.

[0107] (2) KBGAN is a GAN-based method whose main purpose is to generate more high-quality negative samples. The results also show that its performance is indeed better than that of TransE, which indicates that the adversarial network is beneficial to the whole model. However, here KBGAN only uses the information of triples, and the improvement of the results is limited. For the other several models such as DistMult, ConvE, and ConvTransE, the results do not differ much.

[0108] (3) CoFM is a translation-based method, a multi-task collaborative decomposition model for optimizing the KGC task. KGAT is a knowledge graph attention network that learns the embeddings of heterogeneous nodes. However, due to the data sparsity, although their performance is stronger than that of TransE, their overall performance is average.

[0109] (4) Finally, the method proposed in this embodiment is compared with all other methods, and the result is obviously better than other models. Because the model in this embodiment uses the method of combining GAN and then uses the data in the KG and extracts the information in the free text to complete this KGC task. And this embodiment takes optimizing the KGC task as the only goal, and the model in this embodiment combines the graph neural network and the adversarial neural network or assists each other to complete the KGC task. And when the overall data is relatively sparse, by training the discriminator and the generator, some information with relatively high authenticity can be generated, which can help the machine learn more features.

[0110] To verify the accuracy of the model proposed in this embodiment, structured data and free text are respectively added to GAN, and then another data set does not enable GAN, and then compared with the method proposed in this embodiment and DisMult. Because DisMult, this baseline method, is the best-performing method among all baseline methods in terms of comprehensive performance.

[0111] The results are as Figure 2 shown. If GAN is added only after structured data or in free text, whether it is the change in accuracy or the change in amplitude with the increase of epochs is better than the DisMult baseline method, but it is worse than the case where both are added to the GAN network. Because in sparse data, the information and features contained in itself are less, and moreover, the free text of the current data will also add some information. Only when the quantities of these two types of data are large enough can the machine extract more features, so as to summarize more information to complete the knowledge graph completion task.

[0112] Embodiment 2:

[0113] Embodiment 2 of the present invention provides a knowledge graph completion system based on unstructured information, including:

[0114] A data acquisition module, configured to: acquire missing triple data to be completed;

[0115] An entity node recognition module, configured to: recognize entity nodes in the missing triple data;

[0116] A sentence collection module, configured to: acquire sentences associated with the entity nodes, recognize entity triples in the sentences, and at the same time input the obtained sentences into a generator to generate free text data;

[0117] An adversarial training module, configured to: train a generator by combining free text data and structured text data, and a discriminator discriminates the entity triple prediction result of the generator according to the entity triples in a sentence, and perform adversarial training between the generator and the discriminator;

[0118] A knowledge graph completion module, configured to: when the discriminator passes the discrimination, add the entity triples in the sentence to the knowledge graph, score using a graph neural network, and combine the scoring result and the information of the previous node of the known entity node of the missing triple to obtain a ranking result of the entity triples, and complete the completion of the knowledge graph.

[0119] The working method of the system is the same as the knowledge graph completion method based on unstructured information provided in Embodiment 1, and will not be elaborated here.

[0120] Embodiment 3:

[0121] Embodiment 3 of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the knowledge graph completion method based on unstructured information as described in Embodiment 1 of the present invention.

[0122] Embodiment 4:

[0123] Embodiment 4 of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the knowledge graph completion method based on unstructured information as described in Embodiment 1 of the present invention.

[0124] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program code.

[0125] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocksFigure 1 a device for the functions specified in one or more boxes

[0126] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions specified in one or more processes Figure 1 one or more processes and / or boxes Figure 1 a device for the functions specified in one or more boxes

[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes Figure 1 one or more processes and / or boxes Figure 1 a device for the functions specified in one or more boxes

[0128] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0129] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for knowledge graph completion based on unstructured information, characterized in that: It includes: Obtain the missing triple data to be completed; Identify the entity nodes in the missing triple data; Obtain the sentences associated with the entity nodes, identify the entity triples in the sentences, and at the same time input the obtained sentences into the generator to generate free text data; Obtain the sentences associated with the entity nodes. When the missing triple data includes the first entity and the entity relationship and the second entity is missing, the sentences include: Sentences containing the entity relationship on the page where the first entity appears in the unstructured text; Sentences containing the first entity on the page where the entity relationship appears in the unstructured text; Sentences that simultaneously appear containing the first entity and the entity relationship in the unstructured text page; According to the three types of sentences collected on the web page, perform secondary verification of the entity nodes according to the number of occurrences of the entity on each page; The secondary verification is specifically as follows: Collect three types of sentences on the web page. The alternative entities identified in these three sentences are marked as Em = {Pm|Pm∈Ps∪Pr∪Psr}. If the entity Es is marked, it means that this entity appears on the page related to S, and then this entity will be included in the alternative entities and then compared with the related sentences on Pr and Psr to see if there are common entities appearing. If it only appears on one page, the alternative score Se of this entity will be increased accordingly: The +1 in the denominator is to prevent the denominator from being 0; According to the results of the secondary verification, obtain the relevance score of the entity nodes, perform pruning processing on the obtained entity nodes, and retain the preset number of entity nodes in descending order of the score; Combine the free text data and the structured text data for training the generator. The discriminator discriminates the entity triple prediction result of the generator according to the entity triples in the sentence, and performs adversarial training between the generator and the discriminator; The discriminator identifies the true answer queried according to the entity triples as affirmative, and then judges the entity predicted by the generator; When the discriminator passes the discrimination, add the entity triples in the sentence to the knowledge graph, use the graph neural network for scoring, and combine the scoring result and the previous node information of the known entity nodes of the missing triples to obtain the ranking result of the entity triples, and complete the completion of the knowledge graph.

2. The method for knowledge graph completion based on unstructured information according to claim 1, characterized in that: Each candidate entity node includes the node information of the previous layer in the knowledge graph and the information in the free text.

3. The method for knowledge graph completion based on unstructured information according to claim 1, characterized in that: When optimizing the loss of the generator, the policy gradient method is used to learn the parameters, and the entity generator uses the feedback signal of the discriminator as the reward of the generator.

4. A knowledge graph completion system based on unstructured information, characterized in that: It includes: A data acquisition module, configured to: obtain the missing triple data to be completed; An entity node identification module, configured to: identify the entity nodes in the missing triple data; The sentence collection module is configured to: obtain sentences associated with entity nodes, identify entity triples in the sentences, and at the same time input the obtained sentences into a generator to generate free text data; Obtain sentences associated with entity nodes. When the missing triple data includes the first entity and the entity relationship and the second entity is missing, the sentences include: Sentences containing the entity relationship on the page where the first entity appears in the unstructured text; Sentences containing the first entity on the page where the entity relationship appears in the unstructured text; Sentences containing both the first entity and the entity relationship on the unstructured text page; According to the three types of sentences collected on the web page, perform secondary verification of the entity nodes according to the number of occurrences of the entity on each page; The secondary verification is specifically: collect three types of sentences on the web page. The alternative entities identified in these three sentences are marked as Em = {Pm|Pm∈Ps∪Pr∪Psr}. If the entity Es is marked, it means that this entity appears on the page related to S, and then this entity will be included in the alternative entities and then compared with the relevant sentences on Pr and Psr to see if there are common entities appearing. If it only appears on one page, the alternative score Se of this entity will increase accordingly: The denominator +1 is to prevent the denominator from being 0; According to the result of the secondary verification, obtain the relevance score of the entity node, perform pruning processing on the obtained entity node, and retain a preset number of entity nodes in descending order of the score; The adversarial training module is configured to: combine the free text data and the structured text data to train the generator. The discriminator discriminates the entity triple prediction result of the generator according to the entity triples in the sentence, and perform adversarial training between the generator and the discriminator; The discriminator identifies the true answer queried according to the entity triple as affirmative, and then judges the entity predicted by the generator; The knowledge graph completion module is configured to: when the discriminator passes the discrimination, add the entity triples in the sentence to the knowledge graph, use the graph neural network for scoring, and combine the scoring result and the previous node information of the known entity node of the missing triple to obtain the ranking result of the entity triple, and complete the completion of the knowledge graph.

5. A computer-readable storage medium, on which a program is stored, characterized in that, When the program is executed by a processor, it implements the steps in the method for completing a knowledge graph based on unstructured information according to any one of claims 1-3.

6. An electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the method for completing a knowledge graph based on unstructured information according to any one of claims 1-3.

Citation Information

Patent Citations

  • Sample knowledge graph relationship learning method and system based on adversarial attention mechanism

    CN111046187A

  • Knowledge graph embedding model based on entity-relation association graph

    CN113220897A

  • Knowledge graph specific relation completion method based on Internet open world

    CN113360675A