Apparatus and method for determining a knowledge graph
Patent Information
- Application Number
- JP2021076200
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-04-30
- Filing Date
- 2021-04-28
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2041-04-28
AI Technical Summary
Conventional methods lack an efficient means to automatically populate and structure knowledge graphs with entities and relationships.
A device and method for determining a knowledge graph using a model that predicts triplets of entities and relationships based on input data, employing probability calculations and gradient descent to refine parameters, and utilizing cross-entropy and Kullback-Leibler divergence to optimize the model's performance.
Enables the automated construction of knowledge graphs with improved accuracy by determining the likelihood of entity relationships and their descriptions, enhancing the structure and metadata for better knowledge representation.
Smart Images

Figure 00000010_0000 
Figure 00000011_0000 
Figure 00000012_0000
Abstract
Description
[Technology Field]
[0001] Conventional technology The present invention is based on an apparatus and method for determining a knowledge graph. [Background technology]
[0002] A knowledge graph is understood as a way of structuring and storing knowledge in a graph format within a knowledge-based system. A knowledge graph contains multiple entities and reflects the relationships between them. Entities define nodes in the knowledge graph. Relationships are defined as edges between two nodes. [Overview of the project] [Problems that the invention aims to solve]
[0003] It is desirable to provide a means to automatically fill the knowledge graph. [Means for solving the problem]
[0004] Disclosure of the invention This is achieved by the apparatus and method for determining a knowledge graph as described in each independent claim. The knowledge graph includes multiple entities and relationships. The knowledge graph is defined by multiple triplets of the form <entity 1, entity 2, relationship>, where the relationship of one triplet defines the relationship between entity 1 and entity 2 of that triplet. To determine the knowledge graph, a classification decision is made by a model for entity 1, entity 2 and relationship, determining whether a triplet of the form <entity 1, entity 2, relationship> exists and whether it should be written into the knowledge graph.
[0005] A method for determining a knowledge graph includes the steps of forming a first entity of the knowledge graph, forming a text body, forming input data for a model defined depending on the text body and the first entity of the knowledge graph, the model determining predictions for a second entity, predictions for the relationships of a triplet in the knowledge graph, and predictions for the explanation of the triplet, depending on the input data, the model determining a first probability assigned to a triplet and a second probability assigned to the explanation prediction, the model determining the classification of the triplet depending on the first and second probabilities, and, if the classification satisfies predetermined conditions, determining the explanation depending on the explanation prediction, and determining the triplet in the knowledge graph depending on the first entity, the predictions for the second entity, and the predictions for the relationships, wherein a function is defined that depends on a particularly weighted sum of the first and second probabilities, and at least one parameter of the model is trained depending on this function. The first probability represents the probability that a triplet containing a first entity, a predicted second entity, and a predicted relationship exists. The second probability represents the probability that a predicted explanation applies to that triplet. Thus, in this embodiment, a triplet is entered into the knowledge graph only if the triplet exists based on the first probability and the explanation applies based on the second probability. Explanations can be entered or output along with triplets to better understand the structure of the knowledge graph. This function is used in training, for example, to determine at least one parameter that minimizes the function in gradient descent.
[0006] In one embodiment, the method includes the steps of forming a second entity and relation; determining a first cross-entropy between the prediction of the second entity and the second entity; determining a second cross-entropy between the prediction of the relation and the relation; and determining a particularly weighted third cross-entropy between the prediction of the explanation and the explanation, where the function is defined depending on the first cross-entropy, the second cross-entropy, the particularly weighted third cross-entropy, and the particularly weighted sum of the first and second probabilities. The function is, for example, a loss function to be minimized in gradient descent to determine at least one parameter that minimizes the loss function.
[0007] Instead of cross-entropy, other measures can also be used here to characterize the difference between two probability distributions, such as the Kullback-Leibra divergence or other f-divergences. Favorably, the first measure characterizing the difference between the prediction of a second entity and the second entity, and the second measure characterizing the difference between the prediction of a relation and the relation, are given by the same measure.
[0008] For training, training data is formed that includes multiple pairs of triplets and explanations assigned to those triplets, and a classifier is trained to depend on the training data to determine predictions of relationships and explanations about the triplets for a first entity from the triplets.
[0009] In one embodiment, for at least one word or at least one sentence in the text body, a vector representation is determined that defines at least a portion of the input data, in particular depending on at least one other word or at least one other sentence. For example, for each word and each sentence, a context-dependent vector representation is determined that depends on both other sentences in the text body and the first entity.
[0010] Preferably, a first vector is assigned to a first word from a sentence in the text body, a second vector is assigned to a second word from the same sentence in the text body, and the vector representation is calculated as a weighted sum of the first vector and the second vector.
[0011] Preferably, the first output side of the model outputs an output containing a triplet. In this embodiment, the output is a triplet containing a given first entity, a prediction of a second entity, and the predicted relationship between these entities.
[0012] In one embodiment, the second output side of the model outputs an output that defines the beginning and end of at least one region within the text body. The description in this embodiment is actually a portion of the text.
[0013] Preferably, the prediction of a second entity, a relationship, or an explanation is defined by a single value in a value distribution across multiple vectors. The model maps the input data to values representing the degree of fit to a knowledge graph or explanation decision for each vector.
[0014] Preferably, the metadata assigned to the knowledge graph triplets is determined based on the prediction or explanation of the explanation. The metadata is particularly suitable as an explanation of the reason for the triplets being retrieved.
[0015] Preferably, the classification satisfies the above conditions if the first probability exceeds the first threshold and the second probability exceeds the second threshold. The device for determining the knowledge graph is configured to carry out the above method.
[0016] Further advantageous embodiments can be obtained from the following description and drawings. The drawings show the following: [Brief explanation of the drawing]
[0017] [Figure 1] This is a schematic diagram showing the device used to determine the knowledge graph. [Figure 2] This diagram shows the steps involved in determining the knowledge graph. [Figure 3] This diagram shows the steps involved in training a model to determine the knowledge graph. [Modes for carrying out the invention]
[0018] Figure 1 shows a schematic representation of the knowledge graph 100. The knowledge graph 100 can be defined by multiple entities. The first entity is a 1t and the second entity a 2t This is schematically shown in Figure 1.
[0019] The knowledge graph 100 is decidable depending on the model 102. A text body 104 is formed for the determination of the knowledge graph 100. The input data 106 for the model 102 is formed by the device 108 that determines the knowledge graph 100. In this embodiment, the text body 104 is a collection of text or a collection of documents. Based on the text body 104, the device forms the embeddings 110 of individual words or sentences, for example, as vectors. In this embodiment, the input data 106 is the embeddings 110 of the text body and the embeddings 110 of the first entity. In this embodiment, the vector is an embedding. The embeddings of entities or the relationship between the knowledge graph 100 and the text body 104 represent, for example, a representation of multidimensional entities in a lower-dimensional vector space.
[0020] The device 108 includes one or more processors and at least one memory for instructions, and is configured to carry out the method described below. Model 102 in this embodiment includes a first entity, a second entity a 1t and its relationship a 2t A triplet of knowledge graph 100 including t 12It is configured to determine this.
[0021] In relation to Figure 2, the steps in determining the knowledge graph are described below.
[0022] In step 202, the first entity of the knowledge graph 100 is formed. The first entity can be selected from several entities of the knowledge graph 100 that have already been defined. The first entity can be configured by the user through input.
[0023] In step 204, the text body 104 is formed. The text body 104 is read from, for example, a database.
[0024] In step 206, input data 106 for model 102 is formed, which is defined depending on the text body 104 and the first entity of the knowledge graph 100. In this embodiment, input data 106 for model 102 is defined by the text body 104, in particular by the embedding of the document collection or text collection, and by the embedding of the first entity.
[0025] The first entity and text body 104 are represented, for example, by word vectors as embeddings.
[0026] Each word from the first entity and the text body 104 is assigned, for example, a word vector in an n-dimensional vector space.
[0027] Each sentence in the text body 104 is assigned, for example, a sentence vector in an m-dimensional vector space. The dimensions of the vector spaces may be considered identical.
[0028] For example, for each word and / or each sentence of the text body 104, a context-dependent vector representation that depends on other words of the text body 104 is calculated. The context-dependent word representation is determined, for example, by a model that calculates one word representation as a weighted sum of the representations of surrounding words.
[0029] In step 208, the prediction a of the second entity a 1t is determined by the model 102 depending on the input data 106. 1p
[0030] In step 208, the prediction a of the relationship a between the first entity and the second entity a 1t is determined by the model 102 depending on the input data 106. 2t 2p
[0031] For example, the word vectors of the text body 104 are classified in the sense of whether or not the word vector represents the second entity a 1t or is part of it. The prediction a of the second entity a 1t defines, for example, the determined word vector or the determined word. The model 102 determines, for example, one value from an n-dimensional vector space on the output side for the word vector. The value of the word vector forms a value distribution over the word vectors, that is, the word group of the text body 104. The value distribution can be mapped to a probability distribution by a softmax function. The prediction a 1p defines, for example, a word having a higher value than other predictions of other word vectors, that is, as part of the second entity a 1p that is, the triplet 112. A plurality of word vectors or word groups can be assigned to the second entity a 1t where the word vectors or word groups are determined as the part of the second entity a 1t whose predicted value exceeds a predetermined threshold. 1t
[0032] For example, multiple word vectors from a sentence in text body 104 are such that the sentence is a first entity and a second entity a 1t They are classified in the sense of which of the relationships between them they include. For example, relationship a 2t Prediction a 2p This defines the following. Model 102 determines the values of possible relationships, for example, on the output side. These values form a value distribution across possible relationships. The value distribution can be mapped to a probability distribution using the softmax function. Relationship a 2t Prediction a 2p This is defined, for example, by a relation that has a higher value than other possible relations. Relation a in this embodiment 2t Prediction a 2p This is used as part of the triplet 112. It can also be used to determine relationships where the value exceeds a predetermined threshold.
[0033] The first entity and the second entity a in this embodiment 1t and relation a 2t As described below, triplet 112 of the knowledge graph 100 is defined when it is confirmed that triplet 112 and its corresponding description exist.
[0034] Step 208 explains t predictions p However, according to Model 102, this is determined depending on the input data 106.
[0035] For example, the sentence vector from text body 104 is described by the sentence vector s t They are classified in the sense of whether they are important or not. Model 102 determines a single value from an m-dimensional vector space on the output side for the sentence vector. This value forms a value distribution across the sentence vector, i.e., a set of sentences in the text body 104. The value distribution can be mapped to a probability distribution using the softmax function. Prediction s p For example, a statement vector, i.e., a description of a triplet. tDefines a sentence that has a higher value than other predictions of other sentence vectors. This describes multiple sentence vectors, i.e., groups of sentences. t It can be assigned to, where the sentence vector, i.e., the group of sentences, is explained by s t Of these, the portion where the predicted value exceeds a predetermined threshold is determined.
[0036] Descriptions t predictions p or explanation s t This allows defining metadata that can be assigned to the triplet 112 of the knowledge graph 100. The metadata may identify areas of the text body 104, or may include copies or parts of areas of the text body 104.
[0037] The first output side of Model 102 in this embodiment has a triplet 112, i.e., a first entity, a second entity a 1t Prediction a 1p , and relation a 2t Prediction a 2p Output containing this will be output.
[0038] The second output side of Model 102 in this embodiment outputs an output that defines the start and end of at least one region of the text body 104. The output in this embodiment is described in s t predictions p Defined by. Explanations t predictions p This defines, for example, the start and end portions of at least one region of the text body 104. In this embodiment, prediction s p This defines the offsets for the start and end of the region.
[0039] The text body 104 is represented, for example, by a matrix. The columns of the matrix represent, for example, word vectors. The word vectors are arranged in the matrix in the same order as, for example, the group of words in the text. The index of the column in the matrix uniquely identifies the word in this embodiment. A second output is, for example, a start offset and an end offset. The start offset is, for example, the index value in the matrix that uniquely indicates, for example, the position of the word in the text where the description begins. The end offset is, for example, the index value in the matrix that uniquely indicates the position of the word in the text where the description ends. The description is defined in the model, for example, as a vector or submatrix, i.e., as an embedding of a region. The start and end are, for example, integer values for each offset in the text.
[0040] In step 210, the first probability p is used to assign model 102 to the triplet 112. correct_answer This was decided, that is, p correct_answer =softspheric(a 1p )*softmax(a 2p ) You can obtain this.
[0041] The first probability p correct_answer is the second entity a 1t Prediction a 1p Value and relationship a 2t Prediction a 2p It may depend on the product with the value of . In this embodiment, the second entity a 1t Prediction a 1p The probability and relationship a 2t Prediction a 2p The product of the probabilities is determined.
[0042] In step 210, Model 102 is explained. t predictions p The second probability p assigned to gt_explanation This was decided, that is, [Math 1] You can obtain this.
[0043] Second probability p gt_explanation Model 102 describes the triplet 112. t predictions p The product of the values determined for, in this embodiment, can be determined depending on the product of the probability values. In this embodiment, the second probability p gt_explanation This is an explanation. t predictions p This is determined for the sentence vector, which is the part of the sentence.
[0044] In step 212, the classification of the triplet 112 is given by the first probability p. correct_answer and the second probability p gt_explanation It is determined by the following:
[0045] In this embodiment, the triplet 112 of the knowledge graph 100 becomes important when the classification meets predetermined conditions.
[0046] This classification is, for example, based on the first probability p. correct_answer The first threshold is exceeded, and the second probability p gt_explanation The condition is met when it exceeds the second threshold. For the combination of output of triplet 112 and explanation, for that classification, - Triplet 112 is true, and the explanation is true. - Triplet 112 is true, but the explanation is false. - Triplet 112 is false, but the explanation is true (true here means true for the correct output). -Triplet 112 is false, and the explanation is also false. There are four such cases. In this embodiment, the first and second thresholds are defined by probability values in the range of 0 to 1, for example, 0.8 or 0.9. The first and second thresholds may be defined by other values. The first and second thresholds may also be defined by mutually different values.
[0047] The first probability p correct_answerp is a measure of whether the output is a triplet 112, which is true, and the second probability p gt_explanation This is a measure of whether the explanation of the output is true. In the first case, the classification satisfies the condition. In the following three cases, the classification does not satisfy the condition.
[0048] In step 214, if the classification satisfies the conditions, the explanation is the prediction of the explanation. p It is determined by the triplet 112 of the knowledge graph 100, and the second entity a 1t Prediction a 1p and relation a 2t Prediction a 2p This is determined by the following. In this embodiment, if the classification satisfies the conditions, an entry to the knowledge graph 100, including the triplet 112, is determined.
[0049] In this embodiment, the first probability p correct_answer The first threshold is exceeded, and the second probability p gt_explanation If the value exceeds the second threshold, triplet 112 is input. Otherwise, triplet 112 is discarded in this embodiment. Subsequently, step 202 can be performed for the same or other first entities.
[0050] This means that the knowledge graph is built through iteration.
[0051] In relation to Figure 3, the steps of training the model 102 that determines the knowledge graph 100 are described below.
[0052] In step 302, the first entity of the knowledge graph 100 is formed. In step 302, the second entity a of the knowledge graph 100 is formed. 1t In this embodiment, the second entity a 1t is that relationship a 2t These are training data that are mutually known. In step 302, explain s t This is formed. This is, for example, the relevant explanationst is the metadata of
[0053] In step 304, a text body 104 is formed. The text body is advantageously the relationship a between the first entity and the second entity a 1t with 2t corresponding description s t of the text body 104 whose metadata is known.
[0054] In step 306, input data 106 of the model 102 is formed. This is done, for example, as described in step 206.
[0055] In step 308, the prediction a 1t of the second entity a 1p is determined by the model 102 depending on the input data 106.
[0056] In step 308, the prediction a 1t of the relationship a 2t between the first entity and the second entity a 2p is determined by the model 102 depending on the input data 106.
[0057] In step 308, the prediction s t of the description s p is determined by the model 102 depending on the input data 106.
[0058] For this, in this embodiment, it is done as described in step 208.
[0059] In step 310, the first probability p correct_answer assigned by the model 102 to the true triplet 112 known in training is determined.
[0060] For this, in this embodiment, p correct_answer = softmax(a 1t)*softmax(a 2t ) However, as explained in step 210, Model 102 has a second entity a that has become known during training. 1t and relationship a 2t The first probability p represents the probability assigned to the combination that is true. correct_answer This will be decided.
[0061] In step 310, Model 102 is used to explain the explanations that have become known during training. t The second probability p assigned to gt_explanation This will be decided.
[0062] For this reason, in this embodiment, the second probability p gt_explanation This was decided, and here, [Math 2] This represents the probability that Model 102, described in step 210, assigned to all related explanations under the assumption that all related explanations are independent of each other.
[0063] In step 312, predict the second entity a 1p and the second entity a 1t The first cross-entropy CE1 between and is determined. In step 312, the prediction of the relationship a 2p and relationship a 2t The second cross-entropy CE2 between is determined. In step 312, in particular, the coefficient λ sp Weighted explanations t predictions p and explained t The third cross-entropy CE3 between and is determined.
[0064] In step 314, at least one parameter of model 102 is determined such that function J satisfies the conditions. For example, multiple values of function J are determined depending on multiple parameters, where function J satisfies the conditions for a value that is an extremum compared to multiple other values of the multiple parameters, in particular, the smallest of these values. Function J is given by a first probability pcorrect_answer and the second probability p gt_explanation This is a loss function defined depending on the following: In this embodiment, the loss function J is defined by the first cross-entropy CE1, the second cross-entropy CE2, and especially the weighted λ. sp The third cross-entropy CE3 and the first probability p are obtained. correct_answer and the second probability p gt_explanation In particular, weighted λ CC It is defined depending on the sum that is performed.
[0065] In this embodiment, the loss function is the objective function J, which includes other hyperparameters c1, c2, c3 for the loss function J. con Defined by, this is, J=CE1(a 1p ,a 1t )+CE2(a 2p ,a 2t )+λ sp CE3(s p ,s t )+λ CC J con It is possible to optimize in this way, and here, [Math 3] That is the case.
[0066] For training purposes, steps 302 through 314 are repeated with training data.
[0067] In particular, training data is formed, and here, the training data consists of a triplet 112 and descriptions assigned to the triplet 112. t Including multiple pairs, Model 102 has a relationship a with respect to the first entity from the triplet 112. 2t Prediction a 2p and explanations t predictions pThe system includes a classifier that is trained dependent on training data to determine a certain outcome. The classifier may be an artificial neural network, particularly a deep neural network. The artificial neural network includes, for example, an input layer for input data 106 and output layers for a first output and a second output. One or more hidden layers may be placed between the input and output layers. In this embodiment, the parameters of these layers are defined by a set of parameters such that the function J satisfies the conditions during training.
[0068] The intended use is, for example, within the category of material allocation, and aims to build a knowledge database containing comprehensive information about materials and their relationships. This information can be extracted from text, and in addition to relationship information as a description, the relevant sentence portions that led to the extraction of the information will also be extracted.
Claims
1. A method for determining a knowledge graph (100), comprising: forming (202, 302) a first entity of a knowledge graph (100); forming (204, 304) a text body (104); forming (205, 306) input data (106) for a model (102) defined in dependence on the body of text (104) and the first entity of the knowledge graph (100); The second entity (a 1t ) prediction (a 2p ), the relationship (a) of the triplet (112) of the knowledge graph (100) 2t ) prediction (a 2p ), and the description of the triplet (112) (s t ) prediction (s p ) by said model (102) depending on said input data (106); a first probability that the model (102) assigned to the triplet (112) and a second probability that the model (102) assigned to the explanation (s t ) prediction (s p determining (210, 310) a second probability assigned to determining (212, 312) a classification of the triplet (112) depending on the first probability and the second probability; If the classification satisfies the predetermined conditions, the description (s t ) prediction (s p ) depending on the above explanation (s t ) and determining the first entity, the second entity (a 1t ) prediction (a 1p ) and the above relationship (a 2t ) prediction (a 2p determining (214, 314) the triplets (112) of the knowledge graph (100) depending on In a method comprising: a function is defined depending on the weighted sum of the first probability and the second probability, and at least one parameter of the model (102) is trained depending on the function; A method characterized by:
2. The second entity (a 1t ) and the above relationship (a 2t ) forming (302); The second entity (a) characterizes the difference between the two probability distributions. 1t ) prediction (a 1p ) and the second entity (a 1t determining (312) a first measure (CE1) between The relationship (a 2t ) prediction (a 2p ) and the above relationship (a 2t determining (312) a second measure (CE2) between The above explanation (s) characterizes the difference between two probability distributions. t ) prediction (s p ) and the above explanation (s t ) and the weighting (λ sp determining (312) a third measure (CE3), which is the third cross-entropy obtained by Including, The function is a function of the first measure (CE1), the second measure (CE2), the third measure (CE3), and a weighting (λ) of the first probability and the second probability. CC ) is defined as a sum of The method of claim 1.
3. The function is also defined as a function of the sum of the first measure (CE1), the second measure (CE2) and the third measure (CE3). The method of claim 2.
4. said first measure (CE1) characterizing the difference between two probability distributions is at least one of cross-entropy, Kullback-Leibler divergence and f-divergence; and / or the second measure (CE2) characterizing the difference between two probability distributions is at least one of the cross-entropy, the Kullback-Leibler divergence, and the f-divergence; and / or The third measure (CE3) that characterizes the difference between two probability distributions is the cross-entropy, a weighting (λ sp ) is at least one of the cross entropy, Kullback-Leibler divergence, and f-divergence, The method of claim 1.
5. Triplet (112) and the description (s) assigned to the triplet (112) t ) and training data is formed containing multiple pairs of The model (102) defines the relationship (a) for the triplet (112) for the first entity from the triplet (112). 2t ) prediction (a 2p ) and the above explanation (s t ) prediction (s p ), including a classifier trained in dependence on training data to determine 5. The method according to any one of claims 2 to 4.
6. for at least one word or at least one sentence of said text body (104), a vector representation is determined (206, 306) that defines at least a portion of said input data (106), in particular depending on at least one other word or at least one other sentence; 5. The method according to any one of claims 2 to 4.
7. a first vector is assigned to a first word from a sentence in the body of text (104), a second vector is assigned to a second word from the sentence in the body of text (104), and the vector representation is calculated (206, 306) as a weighted sum of the first vector and the second vector; The method of claim 6.
8. an output (212) including the triplet (112) is output to a first output side of the model (102); 5. The method according to any one of claims 2 to 4.
9. At a second output of the model (102), an output is output defining the start and end of at least one region within the text body (104).
5. The method according to any one of claims 2 to 4.
10. The second entity (a 1t ) prediction (a 1p ), the above relationship (a 2t ) prediction (a 2p ), or the above description (s t ) prediction (s p ) is defined by one value from a value distribution across multiple vectors (208,308), 5. The method according to any one of claims 2 to 4.
11. The above explanation (s t ) prediction (s p ) or the above description (s t ) metadata assigned to the triplet (112) of the knowledge graph is determined (208, 308); 5. The method according to any one of claims 2 to 4.
12. The classification satisfies the condition if the first probability is greater than a first threshold and the second probability is greater than a second threshold.
5. The method according to any one of claims 2 to 4.
13. In an apparatus (108) for determining a knowledge graph (100), An apparatus (108) characterized in that the apparatus is configured to carry out the method according to any one of claims 1 to 12.
14. A computer program comprising machine-readable instructions for carrying out the method according to any one of claims 1 to 12 when the computer program is executed on a computer.