Apparatus and method for determining a knowledge graph

CN113590833BActive Publication Date: 2026-09-29ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110471472.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-30
Filing Date
2021-04-29
Publication Date
2026-09-29
Estimated Expiration
2041-04-29

Smart Images

  • Figure CN113590833B_ABST
    Figure CN113590833B_ABST
Patent Text Reader

Abstract

Device and method for determining a knowledge graph, comprising: providing a first entity for the knowledge graph, providing a text body, providing input data for a model defined depending on the text body and the first entity of the knowledge graph, determining, with the model, for a triple of the knowledge graph, a prediction for a second entity, a prediction for a relation and determining a prediction for an explanation for the triple, determining a first probability assigned by the model to the triple and a second probability assigned by the model to the prediction for the explanation, determining a classification for the triple depending on the first probability and the second probability, and if the classification fulfills a condition: determining an explanation depending on the prediction for the explanation and determining a triple for the knowledge graph depending on the first entity, the prediction for the second entity and the prediction for the relation, wherein a function is defined depending on a, in particular weighted, sum of the first probability and the second probability, and wherein at least one parameter for the model is trained depending on the function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention is based on a device and method for determining knowledge graphs. Background Technology

[0002] In knowledge-based systems, a knowledge graph is understood as a structured storage of knowledge in the form of a graph. A knowledge graph includes entities and represents the relationships between them. Entities define the nodes of the knowledge graph. Relationships are defined as edges between two nodes.

[0003] The goal is to enable the possibility of automatically populating knowledge graphs. Summary of the Invention

[0004] This is achieved by the apparatus and method for determining a knowledge graph according to the independent claim. The knowledge graph includes entities and relations. The knowledge graph is defined, for example, by multiple triples of the form <entity 1, entity 2, relation>, where the relation of the triple defines the relationship between entity 1 and entity 2 of the triple. To determine the knowledge graph, a classification decision is made using a model for entity 1, entity 2, and relation regarding whether a triple of the form <entity 1, entity 2, relation> exists and whether the triple should be included in the knowledge graph.

[0005] A method for determining a knowledge graph includes the following steps: providing a first entity for the knowledge graph, providing a text body, providing input data for a model, the input data being defined based on the text body and the first entity of the knowledge graph; using the model to determine predictions for a second entity, a relation, and an interpretation for the triples for the knowledge graph based on the input data; determining a first probability assigned by the model to the triples and a second probability assigned by the model to the prediction for the interpretation; determining a classification for the triples based on the first probability and the second probability; and if the classification satisfies the condition: determining the interpretation based on the prediction for the interpretation and determining the triples for the knowledge graph based on the first entity, the prediction for the second entity, and the prediction for the relation, wherein a function is defined based on a weighted sum of the first probability and the second probability, and wherein at least one parameter for the model is trained based on the function. The first probability indicates the probability that a triple including the first entity, the predicted second entity, and the predicted relation exists. The second probability specifies how likely the predicted interpretation of the triple is correct. Therefore, in this example, the triple is eintragenized into the knowledge graph only if the triple exists based on the first probability and the interpretation is correct based on the second probability. To better understand the structure of the knowledge graph, the interpretation can be eintragenized or output along with the triple. The function is used during training to determine, for example, at least one parameter that minimizes the function in a gradient descent method.

[0006] In one aspect, the method includes: providing a second entity and the relation; determining a first cross-entropy between a prediction for the second entity and the second entity; determining a second cross-entropy between a prediction for the relation and the relation; and determining a third cross-entropy, particularly a weighted one, between a prediction for the interpretation and the interpretation, wherein the function is defined based on the sum of: the first cross-entropy, the second cross-entropy, particularly the weighted third cross-entropy, and particularly a weighted sum of the first probability and the second probability. The function is a loss function, for example, minimized in a gradient descent method to determine at least one parameter that minimizes the loss function.

[0007] Instead of cross-entropy, other measures characterizing the difference between two probability distributions can be used here and below, such as Kullback-Leibler divergence (Kullback-Leibler-Divergenz) or other f-divergence (f-Divergenz). Advantageously, a first measure characterizing the difference between the prediction for the second entity and the first entity, and a second measure characterizing the difference between the prediction for the relation and the relation, are given by the same measure.

[0008] Training data can be provided for training, wherein the training data comprises a large number of pairs consisting of triples and explanations assigned to the triples, wherein the model includes a classifier trained on the training data to determine a prediction for the relation for a first entity from the triples and to determine a prediction for the explanations used for the triples.

[0009] In one aspect, a vector representation is determined for at least one word or at least one sentence of the text body, particularly based on at least one other word or at least one other sentence, wherein the vector representation defines at least a portion of the input data. For example, a context-dependent vector representation is determined for each word and each sentence, which depends on both other sentences of the text body and the first entity.

[0010] Preferably, a first vector is assigned to a first word in a sentence from the text body, and a second vector is assigned to a second word in the sentence from the text body, wherein the vector representation is calculated as a weighted sum of the first vector and the second vector.

[0011] Preferably, the output includes the triples at the first output of the model. In this example, the output defines a triple that includes: a given first entity, a prediction for the second entity, and the predicted relationship between the given first entity and the prediction for the second entity.

[0012] In one aspect, the model outputs at its second output terminal an explanation defining the start and end points of at least one region within the text body. In this example, the explanation is actually derived from a fragment of the text.

[0013] Preferably, the prediction for the second entity, the prediction for the relationship, or the prediction for the interpretation is defined by values ​​about the distribution of values ​​of a large number of vectors. The model maps the input data to values ​​that, for each vector, indicate its suitability for determining the knowledge graph or for the interpretation.

[0014] Preferably, metadata assigned to the triples in the knowledge graph is determined based on predictions made in response to the explanation or based on the explanation itself. Metadata is particularly suitable as an explanation of the cause for the obtained triples.

[0015] Preferably, the classification satisfies the condition if the first probability exceeds a first threshold and the second probability exceeds a second threshold.

[0016] The device used to determine the knowledge graph is configured to perform the method. Attached Figure Description

[0017] Other advantageous embodiments are derived from the following description and accompanying drawings. In the drawings: Figure 1 A schematic diagram of a device for determining a knowledge graph is shown. Figure 2 The steps in the method for determining the knowledge graph are shown. Figure 3 The steps in a method for training a model to determine a knowledge graph are shown. Detailed Implementation

[0018] exist Figure 1 Knowledge graph 100 is schematically illustrated. Knowledge graph 100 can be defined using a large number of entities. Figure 1 The first entity is schematically shown in the middle. Second Entity .

[0019] Knowledge graph 100 can be determined based on model 102. To determine knowledge graph 100, a text body 104 is provided. Input data 106 for model 102 is provided by device 108 for determining knowledge graph 100. In this example, text body 104 is a collection of texts or documents. Starting from text body 104, the device generates embeddings 110 for each word or sentence, for example, as vectors. In this example, input data 106 includes embeddings 110 for the text body and for a first entity. In this example, the vector is an embedding. The embeddings of the text body 104 and the entities or relations of knowledge graph 100 are representations of multidimensional entities in a vector space, for example, a lower dimension.

[0020] Device 108 includes one or more processors and at least one memory for instructions, and is configured to perform the methods described below. In this example, model 102 is configured to determine triples t for knowledge graph 100. 12 The triple includes the first entity and the second entity. and their relationship .

[0021] refer to Figure 2 The steps in the method for determining the knowledge graph are described.

[0022] In step 202, a first entity of knowledge graph 100 is provided. This first entity can be selected from a large number of entities from the already defined knowledge graph 100. The user can pre-define the first entity through input.

[0023] In step 204, text body 104 is provided. For example, text body 104 is read from a database.

[0024] In step 206, input data 106 for model 102 is provided, which is defined based on text body 104 and a first entity of knowledge graph 100. In this example, the input data 106 for model 102 is defined by the embedding of text body 104 (specifically a collection of documents or texts) and by the embedding of the first entity.

[0025] The first entity and text body 104 are represented, for example, by means of embedded word vectors.

[0026] For example, assign a word vector in an n-dimensional vector space to each word from the first entity and the text body 104.

[0027] For example, assign a sentence vector in an m-dimensional vector space to each sentence from the text body 104. The dimensions of the vector space can also be the same.

[0028] For example, a context-dependent vector representation is computed for each word and / or each sentence of the text body 104, which depends on the other words in the text body 104. The context-dependent word representation is determined, for example, by a model that computes the word representation as a weighted sum of the representations of the surrounding (umgeben) words.

[0029] In step 208, model 102 is used to determine the target for the second entity based on input data 106. Prediction .

[0030] In step 208, model 102 is used to determine the first entity and the second entity based on input data 106. Relationship Prediction .

[0031] For example, the second entity is represented by the word vector for the text body 104. Or a second entity A portion is used to classify the word vector. For the second entity... Prediction Define, for example, a specific word vector, i.e., a specific word. Model 102, for example, determines values ​​at the output for word vectors from an n-dimensional vector space. These values ​​of the word vectors form a value distribution with respect to the word vectors, i.e., the words of the text body 104. This value distribution can be mapped to a probability distribution using a softmax function. Prediction For example, the word vector (i.e., the word) that has the highest value compared to other predictions for other word vectors is defined as the second entity. This is defined as part of triple 112. Multiple word vectors (i.e., words) can be assigned to the second entity. The following word vectors (i.e., words) are identified as the second entity. As a part, the predicted values ​​of these word vectors exceed the threshold.

[0032] For example, based on the sentence from text body 104 in the first entity and the second entity The relationships between the words are classified using a large number of word vectors for the sentence. For example, a relationship is defined for each word. Prediction Model 102, for example, determines the values ​​of possible relationships at the output. These values ​​form a value distribution about the possible relationships. This value distribution can be mapped to a probability distribution using the softmax function. For relationships... Prediction For example, a relation can be defined by having the highest value compared to other possible relations. In this example, this relation is used as part of triple 112. Relations whose values ​​exceed a threshold can also be identified.

[0033] In this example, if it is determined that triple 112 exists and the correct interpretation of that triple is as described below, then the first entity and the second entity... and relationships Triple 112 is defined for knowledge graph 100.

[0034] In step 208, model 102 is used to determine the interpretation s based on input data 106. t Predictions p .

[0035] For example, depending on whether the sentence vector from text body 104 is used as an explanation s tThe relevant approach is to classify the sentence vector. Model 102 determines values ​​for the sentence vector from the m-dimensional vector space at the output. These values ​​form a value distribution with respect to the sentence vector (i.e., the sentence in the text body 104). This value distribution can be mapped to a probability distribution using the softmax function. Predictions s p The sentence vector (i.e., the sentence) that has the highest value compared to other predictions of other sentence vectors is defined as the explanation s for the triple. t Multiple sentence vectors (i.e., sentences) can be assigned to the interpretations. t The following sentence vectors (i.e., sentences) are identified as the explanations s. t As a part of it, the predicted values ​​of these sentence vectors exceed the threshold.

[0036] Regarding the explanation of s t Predictions p Or explain s t Metadata can be defined and assigned to triples 112 in knowledge graph 100. This metadata can identify regions of text 104, copies of those regions, or copies of portions of those regions.

[0037] In this example, the output at the first output terminal of model 102 includes a triple 112, which is: the first entity, and for the second entity Prediction and targeting relationship Prediction .

[0038] In this example, the second output of model 102 defines the start and end of at least one region in the text body 104. In this example, by targeting the interpretation s t Predictions p Define the output. For the interpretation s t Predictions p For example, the start and end of at least one region described in text body 104 are defined. In this example, predictions s p The offsets for the start and end of the region are defined.

[0039] The text body 104 is represented, for example, by a matrix. For instance, the columns of this matrix represent word vectors. These word vectors are arranged in the matrix, for example, in the same order as the words in the text. In this example, the indices of the columns in the matrix explicitly identify the words. The second output is, for example, a start offset and an end offset. The start offset is, for example, a value for the following index in the matrix, which explicitly indicates the position in the text where the explanation begins. The end offset is, for example, a value for the following index in the matrix, which explicitly indicates the position in the text where the explanation ends. Within the model, the explanation is defined, for example, as a vector or submatrix, i.e., as an embedding of the region. The start and end are, for example, integer values ​​of corresponding offsets in the text.

[0040] In step 210, the first probability is determined. Model 102 assigns the first probability to triple 112: .

[0041] First probability It can depend on the second entity Prediction Values ​​and Relationships Prediction The product of the values. In this example, determine the product for the second entity. Prediction The probability value and the relationship Prediction The product of probability values.

[0042] Determine the second probability in step 210 Model 102 assigns the second probability to the explanation s t Predictions p : .

[0043] The second probability can be determined by the product of the following values. Model 102 is for the interpretation of triple 112. t Predictions p These values ​​have been determined; in this example, they are probability values. In this example, is used as the basis for interpreting s. t Predictions p The second probability is determined by a portion of the sentence vector. .

[0044] In step 212, according to the first probability Second probability Determine the classification of triple 112.

[0045] In this example, if the classification satisfies the condition, then triple 112 is relevant for knowledge graph 100.

[0046] For example, if the first probability Exceeding the first threshold, and the second probability If the second threshold is exceeded, the classification satisfies the condition. For the combination of the output for triple 112 and the interpretation, there are four cases for the classification: - Triplet 112 is correct, and the interpretation is correct. - Triplet 112 is correct, but the interpretation is wrong. - Triplet 112 is incorrect, but the interpretation is correct (correct in this context means correct for the correct output). - Triplet 112 is incorrect, and its interpretation is wrong. In this example, the first and second thresholds are defined for probability values ​​in the range of 0 to 1, such as 0.8 or 0.9. The first and second thresholds can be defined by other values. The first and second thresholds can be defined by values ​​that are different from each other.

[0047] First probability The measure of the output being a correct triplet 112 is the second probability. This is a measure of whether the interpretation used for the output is correct. In the first case, the classification satisfies the condition. In the latter three cases, the classification does not satisfy the condition.

[0048] In step 214, if the classification satisfies the condition, then based on the prediction s for the explanation... p To determine the interpretation, and based on the first entity, for the second entity Prediction and targeting relationship Prediction Triples 112 for the knowledge graph 100 are determined. In this example, if the classification satisfies the conditions, an entry (Eintrag) including triples 112 is determined in the knowledge graph 100.

[0049] In this example, if the first probability Exceeding the first threshold and the second probability If the threshold is exceeded, triplet 112 is entered. Otherwise, triplet 112 is discarded in this example. Step 202 can then be performed for the same or different first entities.

[0050] The knowledge graph is thus constructed iteratively.

[0051] refer to Figure 3 This describes the steps in a method for training a model 102 for determining a knowledge graph 100.

[0052] In step 302, the first entity of knowledge graph 100 is provided. In step 302, the second entity of knowledge graph 100 is provided. In this example, this is the training data, showing the relationships between them. It is known. An explanation is provided in step 302. t This is, for example, a correct interpretation of s. t Metadata.

[0053] In step 304, a text body 104 is provided. Advantageously, this text body 104 is a text body that, for the first entity and the second entity, provides... Relationship The correct explanation of s t The metadata is known.

[0054] In step 306, input data 106 for model 102 is provided. This is done, for example, as described in step 206.

[0055] In step 308, model 102 is used to determine the target entity based on input data 106. Prediction .

[0056] In step 308, model 102 is used to determine the first entity and the second entity based on input data 106. Relationship Prediction .

[0057] In step 308, model 102 is used to determine the interpretation s based on input data 106. t Predictions p .

[0058] Therefore, in this example, it is processed as described in step 208 (verfahren).

[0059] Determine the first probability in step 310 Model 102 assigns the first probability to the correct triplet 112 known during training.

[0060] Therefore, the first probability is determined in the example. ,in = softmax( )*softmax( This indicates that model 102, as described in step 210, is trained by the second entity known during training. and relationships The probability of the correct combination being assigned.

[0061] In step 310, the model 102 is determined to be based on the known interpretations s from training. t The second probability of allocation .

[0062] Therefore, in this example, the second probability is determined. ,in =∏softmax(s t ) represents the probability that model 102 assigns to all relevant explanations as described in step 210, where it is assumed that these explanations are independent of each other.

[0063] In step 312, the prediction for the second entity is determined. With the second entity The first cross-entropy CE1 between them. The prediction for the relationship is determined in step 312. With Relationship The second cross-entropy CE2 between them. In step 312, the cross-entropy for explanation s is determined. t Predictions p With explanations t In particular, the factor λ sp Weighted third cross-entropy CE3.

[0064] In step 314, at least one parameter for model 102 is determined, wherein for said at least one parameter, function J satisfies a condition. For example, a large number of values ​​for function J are determined based on a large number of parameters, wherein function J satisfies a condition for values ​​of the large number of parameters that are extreme values ​​compared to other values ​​among the large number of values, particularly the minimum value among these values. Function J is a loss function, which is based on a first probability. Second probability And defined. In this example, the loss function J is defined based on the sum of the following: the first cross-entropy CE1, the second cross-entropy CE2, and in particular, the weighted λ. sp The third cross-entropy CE3 and the first probability With the second probability In particular, the weighted λ cc The sum of .

[0065] In this example, the loss function is obtained through the objective function J. con Define the objective function J. con Other hyperparameters c1, c2, and c3 for the loss function J can be optimized: , in For the training, steps 302 to 314 are repeated using the training data.

[0066] Specifically, training data is provided, wherein the training data includes a large amount of data consisting of triples 112 and explanations assigned to triples 112. t The pair consists of a model 102, which includes a classifier trained on the training data to determine the relation for the first entity from the triple 112. Prediction And determine the interpretation s for the triple 112 t Predictions p The classifier can be an artificial neural network, particularly a deep artificial neural network. The artificial neural network includes, for example, an input layer for input data 106 and an output layer for a first output and a second output. One hidden layer or multiple hidden layers can be arranged between the input layer and the output layer. In this example, the parameters of these layers are defined by a large number of parameters, with respect to which the function J satisfies conditions during training.

[0067] Applications, for example, fall within the scope of material allocation, and these applications aim to construct a knowledge database containing all information about materials and their relationships. This information can be extracted from text, wherein, in addition to information about relationships, relevant sentence segments are also extracted as interpretations, leading to the extraction of said information.

Claims

1. A method for determining a knowledge graph (100), characterized in that... The following steps are: providing a first entity for the knowledge graph (100), providing a text body (104), wherein the text body includes a text set or a document set, providing input data (106) for the model (102), the input data being defined based on the text body (104) and the first entity of the knowledge graph (100), and using the model (102) to determine for the second entity (…) of the triples (112) for the knowledge graph (100) based on the input data (106). ) prediction ( ), targeting relationships ( ) prediction ( ) and determine the interpretation (s) for the triple (112). t ) prediction (s p ), determine the first probability assigned by the model (102) to the triple (112) and the probability assigned by the model (102) to the explanation (s t The prediction (s) p The second probability of ), the classification for the triple (112) is determined based on the first probability and the second probability, and if the classification satisfies the condition: based on the interpretation (s) t The prediction (s) p Determine the explanation (s) t And according to the first entity, for the second entity ( The prediction () ) and for the relationship ( The prediction () The triples (112) for the knowledge graph (100) are determined, wherein a function is defined based on the sum of the first probability and the second probability, and wherein at least one parameter for the model (102) is trained based on the function.

2. The method according to claim 1, characterized in that, Provide the second entity ( ) and the relationship ( ), determine for the second entity ( The prediction () ) and the second entity ( The first metric (CE1) characterizing the difference between two probability distributions is determined for the relationship ( The prediction () ) and the relationship ( A second metric (CE2) characterizing the difference between two probability distributions is determined for the interpretation (s) t The prediction (s) p ) and the explanation (s) t The weighted average between (λ) sp The third cross-entropy is a third metric (CE3) representing the difference between two probability distributions, wherein the function is weighted according to the first metric (CE1), the second metric (CE2), and the third metric (CE3), and the first probability and the second probability, by (λ). cc Defined by the sum of ).

3. The method according to claim 2, wherein, The function is defined based on the sum of the first metric (CE1), the second metric (CE2), and the third metric (CE3).

4. The method according to claim 2, wherein, The first metric (CE1) characterizing the difference between two probability distributions is at least one of the following: cross-entropy, Kullback-Leibler divergence, and f-divergence. And / or the second metric (CE2) characterizing the difference between the two probability distributions is at least one of the following: cross-entropy, Kullback-Leibler divergence, and f-divergence. And / or the third metric (CE3) characterizing the difference between the two probability distributions is at least one of the following: cross-entropy, weighted (λ) sp The cross-entropy, Kullback-Leibler divergence, and f-divergence of ).

5. The method according to any one of claims 2 to 4, characterized in that, Provide training data, wherein the training data includes multiple triplets (112) and interpretations (s) assigned to the triplets (112). t The pair consisting of the model (102), wherein the model (102) includes a classifier trained on the training data to determine the relation ( ) for the first entity from the triple (112). The prediction () ) and determine the interpretation (s) for the triple (112). t The prediction (s) p ).

6. The method according to any one of claims 2 to 4, characterized in that, Determine a vector representation for at least one word or at least one sentence of the text body (104), wherein the vector representation defines at least a portion of the input data (106).

7. The method according to claim 6, characterized in that, A first vector is assigned to a first word in a sentence from the text body (104), and a second vector is assigned to a second word in the sentence from the text body (104), wherein the vector representation is calculated as a weighted sum of the first vector and the second vector.

8. The method according to any one of claims 2 to 4, characterized in that, The output including the triplet (112) is output at the first output terminal of the model (102).

9. The method according to any one of claims 2 to 4, characterized in that, The model (102) outputs an output at its second output end that defines the start and end of at least one region in the text body (104).

10. The method according to any one of claims 2 to 4, characterized in that, For the second entity ( The prediction () ), for the relationship ( The prediction () ) or regarding the explanation (s) t The prediction (s) p It is defined by the distribution of values ​​of multiple vectors.

11. The method according to any one of claims 2 to 4, wherein, According to the explanation (s) t The prediction (s) p ) or according to the explanation (s t ) to determine the metadata assigned to the triple (112) in the knowledge graph.

12. The method according to any one of claims 2 to 4, characterized in that, If the first probability exceeds the first threshold and the second probability exceeds the second threshold, then the classification satisfies the condition.

13. The method according to claim 1, characterized in that, The function is defined based on the weighted sum of the first probability and the second probability.

14. The method according to claim 6, characterized in that, For at least one word or at least one sentence of the text body (104), the vector representation is determined based on at least one other word or based on at least one other sentence.

15. A device (108) for determining a knowledge graph (100), characterized in that, The device is configured to perform the method according to any one of claims 1 to 14.

16. A computer program product, characterized in that, The computer program product includes machine-readable instructions that, when executed on a computer, perform the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Relationship prediction method based on knowledge map

    CN108694469A

  • Method and system for automatically constructing knowledge maps for mass unstructured texts

    CN108875051A