Method, apparatus, device, storage medium and program product for knowledge graph completion
By selecting models and answer inference models with associated information, using pre-trained language models to determine the target sub-graphs and triples in the knowledge graph, the problem of the knowledge graph completion method occupying a large amount of storage resources is solved, and efficient knowledge graph completion is achieved.
Patent Information
- Application Number
- CN202210288155.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-03-22
AI Technical Summary
The existing knowledge graph completion method occupies a large amount of storage resources, resulting in inefficient storage.
The correlation information selection model and answer inference model are used to obtain query information, and the target sub-graph and target triplets associated with query information are determined in the knowledge graph, and the entity confidence calculation is performed using a pre-trained language model such as the BERT network, and the answer corresponding to the query information is determined and added to the knowledge graph.
Reduce dependence on inference relationship tables, save storage resources, and improve processing efficiency and accuracy.
Smart Images

Figure CN114547343B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of knowledge graphs, and particularly relates to a method, apparatus, device, storage medium, and program product for knowledge graph completion. Background Art
[0002] A knowledge graph is a graph that uses a network structure to indicate the relationships between information. Knowledge graphs are widely used in scenarios such as semantic search, intelligent question answering, and decision-making assistance. A knowledge graph consists of a vast number of triples, which can be considered as the information units that make up the knowledge graph. A triple consists of a head entity, relationship information, and a tail entity. For example, in the triple (Mr. A, spouse, Ms. B), the head entity is "Mr. A", the relationship is "spouse", and the tail entity is "Ms. B", and this triple means that "the spouse of Mr. A is Ms. B". Users can use the knowledge graph to query the required information. However, the information contained in the existing knowledge graph may not meet the query needs of users. Therefore, knowledge graph completion is required.
[0003] Knowledge graph completion is to infer the corresponding answer based on the given query information through the information in the existing knowledge graph. For example, if the user's query information is "Who is the father-in-law of Mr. A", and the triple (Mr. A, father-in-law,?) is missing in the existing knowledge graph, but the two triples (Mr. A, spouse, Ms. B) and (Ms. B, father, Mr. C) are included, then the information that "the father-in-law of Mr. A is Mr. C" can be inferred through these two triples, and then fed back to the user, and the triple (Mr. A, father-in-law, Mr. C) is added to the knowledge graph.
[0004] Currently, the commonly used knowledge graph completion methods generally establish an inference relationship table in advance, and the inference relationship table is used to record how to infer new relationship information using the existing relationship information. For example, if it is recorded in the inference relationship table that the new relationship information "father-in-law" is inferred through the two existing relationship information "spouse" and "father", then the new triple (Mr. A, father-in-law, Mr. C) can be inferred through the two existing triples (Mr. A, spouse, Ms. B) and (Ms. B, father, Mr. C) in the knowledge graph, and the tail entity of the new triple is used as the answer.
[0005] However, a large number of inference relationships are recorded in the inference relationship table, resulting in a large amount of storage resources being occupied. Summary of the Invention
[0006] This application provides a method, apparatus, device, storage medium, and program product for knowledge graph completion, which can solve the problem of occupying a large amount of storage resources in the prior art.
[0007] In a first aspect, a method for knowledge graph completion is provided. The method includes: obtaining query information, where the query information includes a first entity and relationship information; selecting a model based on existing triples and association information in the knowledge graph, and determining a target subgraph in the knowledge graph associated with the query information; determining target triples included in the target subgraph; based on the query information, the target triples, and an answer inference model, determining a second entity corresponding to the query information, using the second entity to provide feedback on the query information, and adding a triple composed of the first entity, the relationship information, and the second entity to the knowledge graph.
[0008] In a possible implementation, the association information selection model includes a primary selection submodel and a refined selection submodel. The step of selecting a model based on existing triples and association information in the knowledge graph and determining a target subgraph in the knowledge graph associated with the query information includes: based on the existing triples in the knowledge graph and the primary selection submodel, determining a candidate subgraph in the knowledge graph associated with the query information; based on the candidate subgraph and the refined selection submodel, determining a target subgraph in the knowledge graph associated with the query information.
[0009] In a possible implementation, the step of based on the query information, the target triples, and an answer inference model, determining a second entity corresponding to the query information includes: based on the query information, the target triples, and an answer inference model, determining the confidence levels of the entities in the target triples; determining the entity with the highest confidence level among the entities in the target triples as the second entity corresponding to the query information.
[0010] In a possible implementation, the step of based on the query information, the target triples, and an answer inference model, determining the confidence levels of the entities in the target triples includes: based on the query information and the target triples, establishing an input string, where in the input string, the front part of the substring corresponding to each entity of each target triple has a first identifier and the rear part has a second identifier; based on the input string and the answer inference model, determining the confidence levels of the first identifier and the second identifier corresponding to each entity in the target triples; based on the confidence levels of the first identifier and the second identifier of each entity, determining the confidence levels of the entities in the target triples.
[0011] In a possible implementation, the pre-trained language model network with an attention structure is a bidirectional encoder representations from transformers (BERT) network based on a translator.
[0012] In a possible implementation, before obtaining the query information, it further includes: obtaining sample triples, where the sample triples are triples inferred based on the existing triples in the knowledge graph, and the sample triples include a sample first entity, sample relationship information, and a sample second entity; selecting a model based on the existing triples in the knowledge graph and the model for selecting associated information to be trained, and determining a prediction subgraph in the knowledge graph associated with the sample query information, where the sample query information includes the sample first entity and the sample relationship information; determining the prediction triples included in the prediction subgraph; based on the sample query information, the prediction triples, and the answer inference model to be trained, determining a predicted second entity corresponding to the sample query information; and based on the sample second entity and the predicted second entity, adjusting the model parameters of the model for selecting associated information to be trained and the answer inference model to be trained.
[0013] In a second aspect, a device for knowledge graph completion is provided. The device includes: an obtaining module, configured to obtain query information, where the query information includes a first entity and relationship information; a determining module, configured to determine, based on the existing triples in the knowledge graph and the model for selecting associated information, a target subgraph in the knowledge graph associated with the query information; determining the target triples included in the target subgraph; based on the query information, the target triples, and the answer inference model, determining a second entity corresponding to the query information, using the second entity to feedback on the query information, and adding the triple formed by the first entity, the relationship information, and the second entity to the knowledge graph.
[0014] In a possible implementation, the model for selecting associated information includes a primary selection sub-model and a refined selection sub-model. The determining module is configured to determine, based on the existing triples in the knowledge graph and the primary selection sub-model, a candidate subgraph in the knowledge graph associated with the query information; and based on the candidate subgraph and the refined selection sub-model, determine a target subgraph in the knowledge graph associated with the query information.
[0015] In a possible implementation, the determining module is configured to determine the confidence levels of the entities in the target triples based on the query information, the target triples, and the answer inference model; and determine the entity with the highest confidence level among the entities in the target triples as the second entity corresponding to the query information.
[0016] In a possible implementation manner, the determining module is configured to establish an input string based on the query information and the target triple. In the input string, a first identifier is provided at the front of the substring corresponding to each entity of each target triple, and a second identifier is provided at the rear; determine the confidence levels of the first identifier and the second identifier corresponding to each entity in the target triple based on the input string and the answer inference model; and determine the confidence levels of each entity in the target triple based on the confidence levels of the first identifier and the second identifier of each entity.
[0017] In a possible implementation manner, the pre-trained language model network with an attention structure is a bidirectional encoder representation from Transformer (BERT) network.
[0018] In a possible implementation manner, the apparatus is further configured to obtain sample triples, where the sample triples are triples inferred based on the existing triples in the knowledge graph. The sample triples include a sample first entity, sample relationship information, and a sample second entity; determine a predicted subgraph associated with the sample query information in the knowledge graph based on the existing triples in the knowledge graph and a model for selecting associated information to be trained, where the sample query information includes the sample first entity and the sample relationship information; determine the predicted triples included in the predicted subgraph; determine a predicted second entity corresponding to the sample query information based on the sample query information, the predicted triples, and the answer inference model to be trained; and adjust the model parameters of the model for selecting associated information to be trained and the answer inference model to be trained based on the sample second entity and the predicted second entity.
[0019] In a third aspect, a computer device is provided. The computer device includes a processor and a memory. The memory is configured to store computer instructions, and the processor is configured to execute the computer instructions stored in the memory, so that the computer device executes the method according to the first aspect and its possible implementation manners.
[0020] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer program code, and when the computer program code is executed by a computer device, the computer device executes the method according to the first aspect and its possible implementation manners.
[0021] In a fifth aspect, a computer program product is provided. The computer program product includes computer program code, and when the computer program code is executed by a computer device, the computer device executes the method according to the first aspect and its possible implementation manners.
[0022] In the embodiments of the present application, an associated information selection model is used to determine a target subgraph in the knowledge graph that is associated with the query information, and the target triples included in the target subgraph are determined. Based on the query information, the target triples, and an answer inference model, the answer corresponding to the query information is determined. In this way, only the associated information selection model and the answer inference model need to be stored, instead of storing a huge inference relation table. Generally, the data volume of the machine learning model is relatively small, thus saving a large amount of storage resources. Moreover, using the machine learning model avoids performing inference processing using a large number of inference relations, thereby improving the processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 is a schematic diagram of a knowledge graph provided by an embodiment of the present application;
[0025] Figure 2 is a schematic diagram of a triple provided by an embodiment of the present application;
[0026] Figure 3 is a schematic diagram of a path provided by an embodiment of the present application;
[0027] Figure 4 is a schematic diagram of a chain provided by an embodiment of the present application;
[0028] Figure 5 is a schematic diagram of the structure of a computer device provided by an embodiment of the present application;
[0029] Figure 6 is a flowchart of knowledge graph completion provided by an embodiment of the present application;
[0030] Figure 7 is a flowchart of knowledge graph completion provided by an embodiment of the present application;
[0031] Figure 8 is a schematic diagram of knowledge graph completion provided by an embodiment of the present application;
[0032] Figure 9 is a flowchart of determining the tail entity provided by an embodiment of the present application;
[0033] Figure 10 is a flowchart of training a machine learning model provided by an embodiment of the present application;
[0034] Figure 11 It is a flowchart for determining a predicted tail entity provided by an embodiment of the present application;
[0035] Figure 12 It is a schematic diagram of a device for knowledge graph completion provided by an embodiment of the present application;
[0036] Figure 13 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0037] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0038] First, explain the nouns involved in the embodiments:
[0039] Knowledge graph: A graph that uses a network structure to indicate the connections between information. The network includes nodes and edges connecting the nodes. Refer to Figure 1 , where the nodes in the graph represent the entities represented by the information (entities can be divided into head entities and tail entities), and the edges represent the relationships between the entities.
[0040] Triple: A set of information composed of a head entity, a relationship, and a tail entity, which can be expressed as (h, r, t). The elements within a triple are arranged in order. The first element within a triple represents the head entity, the second element represents the relationship information, and the third element represents the tail entity. A triple can also be expressed as . Refer to Figure 2 For the triple shown, the head entity is "Mr. A", the relationship information is "spouse", and the tail entity is "Ms. B". The meaning expressed by this triple can be "The spouse of Mr. A is Ms. B".
[0041] Subgraph: A part of the knowledge graph, which can be understood as a subset of the knowledge graph. A subgraph can be composed of paths and / or chains.
[0042] Path: Composed of multiple connected triples, which can be expressed as . Refer to Figure 3 , in a path, the tail entity of the previous triple is the head entity of the next triple, and each entity is connected to at most one head entity and / or one tail entity.
[0043] Chain: Composed of multiple connected triples, or it can be said to be a structure composed of multiple paths. Refer to Figure 4 , the difference from a path is that each entity in a chain can be connected to multiple head entities and / or tail entities.
[0044] Knowledge graphs can be applied to scenarios such as semantic search, intelligent question answering, and decision-making assistance to help users obtain the information they need. Taking the intelligent question answering scenario as an example, a user can enter question information in the question input field of the intelligent question answering, such as "Who is Mr. A's father-in-law?". After receiving the question information, the question information can be analyzed and processed to obtain the triple information (which can be called query information) of the question information. The query information can be (Mr. A, father-in-law,?). Here, "?" represents the content to be queried. Then, look for triples in the existing triples of the knowledge graph whose head entity and relationship information are the same as those in the query information (which can be called answer triples). If there are answer triples in the knowledge graph, such as (Mr. A, father-in-law, Mr. C), then the tail entity of this triple is fed back to the user. If there are no answer triples in the knowledge graph, then the knowledge graph can be completed based on this question information.
[0045] Based on the above application scenarios, the embodiments of the present application provide a method for completing a knowledge graph, and this method can be implemented by a computer device. This computer device can be a server or a terminal, etc. The terminal includes but is not limited to mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, aircraft, etc. The server can be a single server or a server group composed of multiple servers. The embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.
[0046] Figure 5 It is a schematic structural diagram of a computer device provided by the embodiments of the present application. From the perspective of hardware composition, the structure of the computer device 500 can be as Figure 5 shown, including a processor 501, a memory 502, and a communication component 503.
[0047] The processor 501 can be a central processing unit (CPU) or a system on chip (SoC), etc. The processor 501 can be used to obtain query information, can be used to select a model based on the existing triples and association information of the knowledge graph, determine the target sub-graph associated with the query information in the knowledge graph, can be used to determine the target triples included in the target sub-graph, and can also be used to determine the second entity corresponding to the query information based on the query information, the target triples, and the answer inference model, etc.
[0048] The memory 502 may include various volatile or non-volatile memories, such as a solid state disk (SSD), a dynamic random access memory (DRAM), etc. The memory 502 may be used for pre-stored data, intermediate data, and result data during the processing of knowledge graph completion. For example, query information, existing triples in the knowledge graph, target subgraphs, target triples, the second entity corresponding to the query information, and so on.
[0049] In addition to the processor 501 and the memory 502, the computer device 500 may further include a communication component 503.
[0050] The communication component 503 may be a wired network connector, a wireless fidelity (WiFi) module, a Bluetooth module, a cellular network communication module, etc. The communication component 503 may be used for data transmission with other devices, and the other devices may be servers or terminals, etc. For example, the computer device 500 may receive query information, may send the second entity corresponding to the query information, and the computer device 500 may also send a triple composed of the first entity and relationship information included in the query information and the second entity corresponding to the query information to the server for storage.
[0051] In the embodiments of the present application, the execution subject is taken as an example of a server for illustration. Figure 6 It is a flowchart of a method for knowledge graph completion provided by the embodiments of the present application. Refer to Figure 6 This process may include the following steps:
[0052] 601. Obtain query information.
[0053] Among them, the query information includes a first entity and relationship information.
[0054] 602. Based on the existing triples and association information selection model in the knowledge graph, determine the target subgraph in the knowledge graph associated with the query information.
[0055] Among them, the association information selection model may include a primary selection sub-model and a refined selection sub-model.
[0056] In implementation, the query information and the existing triples in the knowledge graph may be input into the association information selection model, and the association information selection model may output the target subgraph in the knowledge graph associated with the query information.
[0057] 603. Determine the target triples included in the target subgraph.
[0058] The server can determine the triples included in the target subgraph, and further determine the same triples included in the target subgraph. Then, the server can perform deduplication processing on the same triples, and only retain one of the multiple identical triples. Then, the server can determine the distinct triples included in the target subgraph as the target triples.
[0059] 604. Based on the query information, the target triples, and the answer inference model, determine the second entity corresponding to the query information, use the second entity to feedback on the query information, and add the triple composed of the first entity, the relationship information, and the second entity to the knowledge graph.
[0060] Among them, both the association information selection model and the answer inference model include a pre-trained language model PLM network with an attention structure. The first entity can be the head entity, and the second entity can be the tail entity; or, the first entity can be the tail entity, and the second entity can be the head entity.
[0061] In implementation, the query information and the target triples can be input into the answer inference model, and the answer inference model can output the determination of the second entity corresponding to the query information.
[0062] Figure 7 is a flowchart of a method for knowledge graph completion provided by an embodiment of the present application. Refer to Figure 7 , this process is described with the first entity as the head entity and the second entity as the tail entity, and may include the following steps:
[0063] 701. Obtain query information.
[0064] Among them, the query information includes the head entity and the relationship information.
[0065] 702. Based on the existing triples in the knowledge graph and the primary selection sub-model, determine the candidate subgraph associated with the query information in the knowledge graph.
[0066] Among them, the primary selection sub-model can be a machine learning model, and there can be multiple possible types of the primary selection sub-model. For example, the primary selection sub-model can be a MINERVA algorithm model, etc. The candidate subgraph can include an initial path and / or an initial chain.
[0067] In implementation, the query information and the existing triples in the knowledge graph can be input into the primary selection sub-model, and the primary selection sub-model can output at least one candidate path and / or at least one candidate chain, and the first confidence corresponding to these candidate paths and / or candidate chains. Then, these candidate paths and / or candidate chains can be used as the initial path and / or initial chain associated with the query information.
[0068] Alternatively, based on the first confidence level corresponding to the candidate path and / or candidate chain, the initial path and / or initial chain associated with the query information in the knowledge graph can be determined, and there can be multiple possible corresponding processing methods.
[0069] In Method 1, the candidate path and / or candidate chain with the first confidence level within the preset first confidence interval can be used as the initial path and / or initial chain associated with the query information.
[0070] For example, the first confidence interval can be (0.8, 1). If the confidence level of the candidate path or candidate chain is within (0.8, 1), then the candidate path or candidate chain can be used as the initial path or initial chain associated with the query information. There can be multiple candidate paths and / or candidate chains.
[0071] In Method 2, the first confidence levels can be sorted from high to low, and the candidate paths and / or candidate chains with the first few highest confidence levels can be selected as the initial paths and / or initial chains associated with the query information.
[0072] For example, the candidate paths and / or candidate chains with the top 5 first confidence levels can be selected as the initial paths and / or initial chains associated with the query information.
[0073] 703. Based on the candidate subgraph and the selected sub-model, determine the target subgraph in the knowledge graph that is associated with the query information.
[0074] Among them, the selected sub-model can be a machine learning model, and there can be multiple possible types of the selected sub-model. For example, the selected sub-model can include a pre-trained language model (PLM) network with an attention structure, such as a bidirectional encoder representation from transformers (BERT) model network, a RoBERTa model network, an ALBERT model network, etc. The target subgraph can include a target path and / or a target chain.
[0075] In implementation, the query information and the initial path / initial chain can be input into the selection sub-model, and the selection sub-model can output the second confidence corresponding to the initial path / initial chain. The selection sub-model can include an encoding module and a fully connected layer module. The encoding module can encode the input query information and the initial path / initial chain to obtain the vector corresponding to the initial path / initial chain. The encoding module can include a pre-trained language model network with an attention structure, such as a BERT model network, etc. Then, the vector corresponding to the initial path / initial chain can be input into the fully connected layer module to obtain the score corresponding to the initial path / initial chain. The calculation of the fully connected layer can be expressed as:
[0076]
[0077] where represents the score corresponding to the k-th initial path / initial chain, represents the network parameters of the encoding module, represents the vector corresponding to the k-th initial path / initial chain.
[0078] Furthermore, the score corresponding to the initial path / initial chain can be normalized to obtain the second confidence corresponding to the initial path / initial chain. The calculation of the normalization process can be expressed as:
[0079]
[0080] where represents the second confidence corresponding to the k-th initial path / initial chain, represents the network parameters of the encoding module, represents the vector corresponding to the k-th initial path / initial chain, represents the vector corresponding to the i-th initial path / initial chain, and there are a total of j initial paths and initial chains.
[0081] After determining the second confidence corresponding to the initial path / initial chain, based on the second confidence, the target path and / or target chain associated with the query information in the knowledge graph can be determined, and there can be multiple possible corresponding processing methods.
[0082] Method 3: The initial path and / or initial chain with the second confidence within the preset second confidence interval can be used as the target path and / or target chain associated with the query information. The related processing is similar to Method 1 in step 702, and the relevant description of Method 1 can be referred to.
[0083] In Method 4, the second confidence levels can be sorted from high to low, and the initial paths and / or initial chains among the top several in the second confidence level sorting are selected as the target paths and / or target chains associated with the query information. The relevant processing is similar to Method 2 in Step 702, and the relevant descriptions of Method 2 can be referred to.
[0084] The primary selection sub-model in 702 and the refined selection sub-model in 703 can be different sub-models in the same machine learning model (associated information selection model).
[0085] 704. Determine the target triples included in the target sub-graph.
[0086] The server can determine the target paths and / or target chains included in the target sub-graph, determine the triples included in the target paths and / or target chains, and then determine the same triples included in multiple target paths and / or target chains. Then, the server can perform deduplication processing on the same triples, and only retain one of the multiple identical triples. Then, the server can determine the different triples included in the target paths and / or target chains as the target triples.
[0087] 705. Based on the query information, the target triples, and the answer inference model, determine the tail entity corresponding to the query information.
[0088] Among them, the answer inference model can be a machine learning model, and there can be various types of the answer inference model. For example, the answer inference model can include a PLM network with an attention structure, such as a BERT model network, a RoBERTa model network, an ALBERT model network, etc. The attention structure can be used to establish the connection between each character in the input string, and the PLM network can be used to learn the context relationship between each character in the input string.
[0089] The corresponding processing of this step can be refined into the execution process as Figure 8 shown:
[0090] 7051. Based on the query information and the target triples, establish an input string.
[0091] In implementation, an input string including the string of the query information and the string of the target triples can be established. The input string can include the string of the query information and the string of the target triples, and the front part of the substring corresponding to each entity of each target triple has a start position identifier, and the rear part has an end position identifier.
[0092] Taking the answer inference model including the BERT model network as an example, the input string can be "[CLS] Mr. A's father-in-law [SEP][S] Mr. A [E] spouse [S] Ms. B [E]. [S] Ms. B [E] father [S] Mr. C [E]", and this input string contains the string of query information and the string of the target triple. The front part of the string of query information has the [CLS] identifier, and the rear part has the [SEP] identifier. The [CLS] identifier is the classification identifier corresponding to this input string, and the [SEP] identifier is the delimiter that separates the string of query information and the string of the target triple. The strings of each target triple are separated by the "." symbol. In the string of each target triple, the front part of the substring corresponding to each entity has the start position identifier [S], and the rear part has the end position identifier [E].
[0093] 7052. Based on the input string and the answer inference model, determine the confidence levels of the start position identifier and the end position identifier corresponding to each entity in the target triple.
[0094] In implementation, the input string can be input into the answer inference model, and the answer inference model can output the confidence levels of the start position identifier and the end position identifier corresponding to each entity.
[0095] Taking the answer inference model including the BERT model network as an example, the BERT model can generate the position embedding vector, paragraph vector, and word meaning vector corresponding to each character (including letters, symbols, and identifiers) in the input string. By adding the position embedding vector, paragraph vector, and word meaning vector corresponding to the character, the initial vector corresponding to this character can be obtained.
[0096] Optionally, the digital label of the position embedding vector corresponding to each triple in the target triple can be changed so that the actual length of the input string can exceed the 512-character length limit of the BERT network. Thus, more target triples can be used for inference.
[0097] The BERT model can process the initial vector corresponding to each character to obtain the output vector corresponding to each character. Then, the average of the output vectors corresponding to the start position identifier of the same entity can be calculated (that is, divide the sum of the output vectors by the number of output vectors) to obtain the average vector corresponding to the start position identifier of this entity. The average of the output vectors corresponding to the end position identifier of the same entity can be calculated to obtain the average vector corresponding to the end position identifier of this entity. The processing from step 703 to step 7052 can refer to Figure 9 the example of
[0098] The answer reasoning model may further include a fully connected layer. The average vector corresponding to the start position identifier of the entity can be input into the fully connected layer, and the fully connected layer can output the score of the start position identifier corresponding to the entity. The calculation of the fully connected layer can be expressed as:
[0099]
[0100] where represents the score of the start position identifier corresponding to the k-th entity, represents the PLM network parameters with an attention structure, represents the average vector corresponding to the start position identifier of the k-th entity.
[0101] By normalizing the score, the confidence of the start position identifier corresponding to the entity can be obtained. The calculation of normalization can be expressed as:
[0102]
[0103] where represents the confidence of the start position identifier corresponding to the k-th entity, represents the PLM network parameters with an attention structure, represents the average vector corresponding to the start position identifier of the k-th entity, represents the average vector corresponding to the start position identifier of the i-th entity, and there are a total of j entities.
[0104] Similarly, the average vector corresponding to the end position identifier of the entity can be input into the fully connected layer, and the fully connected layer can output the score of the end position identifier corresponding to the entity. By normalizing the score, the confidence of the end position identifier corresponding to the entity can be obtained.
[0105] 7053. Based on the confidence of the start position identifier and the end identifier of each entity, determine the confidence of each entity in the target triple.
[0106] In implementation, the average of the confidence of the start position identifier and the confidence of the end identifier of the entity can be calculated (that is, adding these two confidences and dividing by 2), and this average value is used as the confidence of the entity.
[0107] 7054. Among the entities in the target triple, determine the entity with the highest confidence as the tail entity corresponding to the query information.
[0108] The confidences of each entity can be compared, and the entity with the highest confidence is determined as the tail entity corresponding to the query information.
[0109] 706. Use the tail entity to feedback the query information, and add the triple composed of the head entity, relationship information, and tail entity to the knowledge graph.
[0110] After determining the tail entity corresponding to the query information, the tail entity can be fed back to the user, and the triple composed of the head entity, relationship information, and the tail entity included in the query information is added to the knowledge graph.
[0111] Some machine learning models are involved in the above processing process. The embodiments of the present application provide a training method for machine learning models, such as Figure 10 shown, this method may include the following steps:
[0112] 1001. Obtain sample triples.
[0113] Among them, the sample triples include sample head entities, sample relationship information, and sample tail entities.
[0114] The sample triples can be triples inferred from the existing triples in the knowledge graph. For example, if there are existing triples (Mr. A, spouse, Ms. B) and triples (Ms. B, father, Mr. C) in the knowledge graph, then the sample triples (Mr. A, father-in-law, Mr. C) can be inferred through these two triples.
[0115] 1002. Select a model based on the existing triples in the knowledge graph and the association information to be trained, and determine the prediction subgraph associated with the sample query information in the knowledge graph.
[0116] Among them, the sample query information may include sample head entities and sample relationship information. The association information selection model to be trained may include a primary selection submodel to be trained and a refined selection submodel to be trained. The prediction subgraph may include prediction paths and / or prediction chains.
[0117] In implementation, the query information and the existing triples in the knowledge graph can be input into the initial submodel to be trained. The initial submodel to be trained can output at least one predicted candidate path and / or at least one predicted candidate chain, and the third confidence corresponding to these predicted candidate paths and / or predicted candidate chains. These predicted candidate paths and / or predicted candidate chains can be used as the initial predicted paths and / or initial predicted chains associated with the query information. Or, based on the third confidence corresponding to the predicted candidate paths and / or predicted candidate chains, determine the initial predicted paths and / or initial predicted chains associated with the query information in the knowledge graph. The corresponding processing methods and steps are similar to those in Method 1 and Method 2 in step 702, and the relevant descriptions can be referred to.
[0118] Then, the query information and the initial prediction path / initial prediction chain can be input into the to-be-trained selected sub-model, and the to-be-trained selected sub-model can output the fourth confidence corresponding to the initial prediction path / initial prediction chain. Furthermore, based on the fourth confidence, the prediction path and / or prediction chain associated with the query information in the knowledge graph can be determined. The corresponding processing methods and steps 703 are similar to Method 3 and Method 4, and the relevant descriptions can be referred to.
[0119] 1003. Determine the prediction triples included in the prediction sub-graph.
[0120] The server can determine the prediction paths and / or prediction chains included in the prediction sub-graph, determine the triples included in the prediction paths and / or prediction chains, and then determine the same triples included in multiple prediction paths and / or prediction chains. Then, the server can perform deduplication processing on the same triples, and only retain one of the multiple identical triples. Then, the server can determine the different triples included in the prediction paths and / or prediction chains as the prediction triples.
[0121] 1004. Based on the sample query information, the prediction triples, and the to-be-trained answer inference model, determine the predicted tail entity corresponding to the sample query information.
[0122] Among them, the to-be-trained answer inference model can be a machine learning model, and there can be multiple possible types of the to-be-trained answer inference model. For example, the to-be-trained answer inference model can include a PLM network with an attention structure, such as a BERT model network, a RoBERTa model network, an ALBERT model network, etc.
[0123] The corresponding processing of this step can be refined into the execution process as Figure 11 shown:
[0124] 10041. Based on the sample query information and the prediction triples, establish an input string.
[0125] In implementation, an input string including the string of the sample query information and the string of the prediction triples can be established. The input string can include the string of the sample query information and the string of the prediction triples, and the front part of the substring corresponding to each entity of each prediction triple has a start position identifier, and the rear part has an end position identifier.
[0126] 10042. Based on the input string and the to-be-trained answer inference model, determine the confidence of the start position identifier and the end position identifier corresponding to each entity in the prediction triples.
[0127] In implementation, an input string can be input into an answer inference model to be trained, and the answer inference model to be trained can output the confidence levels of the start position identifiers and end position identifiers corresponding to each entity.
[0128] 10043, Based on the confidence levels of the start position identifiers and end identifiers of each entity, determine the confidence level of each entity in the predicted triple.
[0129] In implementation, the average of the confidence level of the start position identifier and the confidence level of the end identifier of each entity in the predicted triple can be calculated (that is, add these two confidence levels and divide by 2), and this average value is used as the confidence level of the entity.
[0130] 10044, Among the entities in the predicted triple, determine the entity with the highest confidence level as the predicted tail entity corresponding to the sample query information.
[0131] The confidence levels of each entity can be compared to determine the entity with the highest confidence level as the predicted tail entity corresponding to the query information.
[0132] 1005, Based on the sample tail entity and the predicted tail entity, adjust the model parameters of the association information selection model to be trained and the answer inference model to be trained.
[0133] The sample tail entity and the predicted tail entity can be input into a loss function to obtain the target loss value of the predicted tail entity. The loss function can be various types of loss functions, such as a quadratic loss function, an absolute loss function, and so on. Then, based on this target loss value, the model parameters of the association information selection model to be trained and the answer inference model to be trained can be adjusted.
[0134] After the model parameters are adjusted, the sample triple can be replaced, and the above training process can be repeated using the adjusted association information selection model and the answer inference model to be trained until the training end condition is met. The training end condition can be that the absolute value of the target loss value is less than a preset target loss value threshold, or it can also be that the number of training times reaches the training times threshold, and so on.
[0135] To test the beneficial effects of the knowledge graph completion method provided in this embodiment, the technical personnel set different processing methods and conducted performance tests separately. The test dataset used was FB15K-237. The technical personnel constructed a MINERVA model and a CoPER-MINERVA model for performance test comparison. In the first processing method, the MINERVA model was used for processing. The corresponding processing was to input the query information and the existing knowledge graph into the MINERVA model, and the MINERVA model could output the answer corresponding to the query information. In the second processing method, the CoPER-MINERVA model was used for processing. The corresponding processing was to output the query information and the existing knowledge graph into the CoPER-MINERVA model, and the CoPER-MINERVA model output the answer corresponding to the query information. The third processing method was the knowledge graph completion method provided in this embodiment. The performance test results are shown in the following table:
[0136]
[0137] Table 1
[0138] Among them, Hit@1 represents the frequency that the entity with the highest confidence obtained by this processing method is the sample tail entity; Hit@3 represents the frequency that among the top three entities with the highest confidence obtained by this processing method, there is an entity that is the same as the sample tail entity.
[0139] As can be seen from the above table, the knowledge graph completion method provided in this embodiment significantly improves the accuracy of knowledge graph completion.
[0140] Figure 12 This is a device for knowledge graph completion provided in an embodiment of the present application. The device includes: an acquisition module 1201, configured to acquire query information, where the query information includes a first entity and relationship information; a determination module 1202, configured to select a model based on the existing triples and association information of the knowledge graph, determine a target subgraph in the knowledge graph that is associated with the query information; determine the target triples included in the target subgraph; based on the query information, the target triples, and an answer inference model, determine a second entity corresponding to the query information, use the second entity to feedback the query information, and add the triple composed of the first entity, the relationship information, and the second entity to the knowledge graph.
[0141] In a possible implementation, the associated information selection model includes a primary selection sub-model and a refined selection sub-model. The determination module 1202 is configured to determine, based on the existing triples in the knowledge graph and the primary selection sub-model, a candidate sub-graph in the knowledge graph that is associated with the query information; and determine, based on the candidate sub-graph and the refined selection sub-model, a target sub-graph in the knowledge graph that is associated with the query information.
[0142] In a possible implementation, the determination module 1202 is configured to determine the confidence levels of the entities in the target triple based on the query information, the target triple, and the answer inference model; and determine the entity with the highest confidence level among the entities in the target triple as the second entity corresponding to the query information.
[0143] In a possible implementation, the determination module 1202 is configured to establish an input string based on the query information and the target triple. In the input string, the front part of the substring corresponding to each entity in each target triple has a first identifier, and the rear part has a second identifier; determine the confidence levels of the first identifier and the second identifier corresponding to each entity in the target triple based on the input string and the answer inference model; and determine the confidence levels of the entities in the target triple based on the confidence levels of the first identifier and the second identifier of each entity.
[0144] In a possible implementation, the pre-trained language model network with an attention structure is a bidirectional encoder representations from transformers (BERT) network based on a translator.
[0145] In a possible implementation, the device is further configured to obtain sample triples, where the sample triples are triples inferred based on the existing triples in the knowledge graph. The sample triples include a sample first entity, sample relationship information, and a sample second entity; determine a predicted sub-graph in the knowledge graph that is associated with the sample query information based on the existing triples in the knowledge graph and the to-be-trained associated information selection model, where the sample query information includes the sample first entity and the sample relationship information; determine the predicted triples included in the predicted sub-graph; determine a predicted second entity corresponding to the sample query information based on the sample query information, the predicted triples, and the to-be-trained answer inference model; and adjust the model parameters of the to-be-trained associated information selection model and the to-be-trained answer inference model based on the sample second entity and the predicted second entity.
[0146] Figure 13It is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 1300 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 1301 and one or more memories 1302. Among them, at least one instruction is stored in the memory 1302, and at least one instruction is loaded and executed by the processor 901 to implement the methods provided by the above various method embodiments. Of course, the computer device may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The computer device may also include other components for implementing device functions, which will not be elaborated here.
[0147] In an exemplary embodiment, a computer-readable storage medium is also provided. For example, a memory including instructions, and the above instructions can be executed by a processor in a terminal to complete the method for knowledge graph completion in the above embodiment. The computer-readable storage medium may be non-transitory. For example, the computer-readable storage medium may be a ROM (read-only memory), a RAM (random access memory), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0148] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0149] The above are only optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for knowledge graph completion, characterized in that, The method includes: Obtaining query information, where the query information includes a first entity and relationship information; Selecting a model based on existing triples and association information in the knowledge graph, and determining a target subgraph in the knowledge graph associated with the query information; Determining the target triples included in the target subgraph; Based on the query information and the target triples, an input string is established. In the input string, the front part of the substring corresponding to each entity of the target triples has a first identifier, and the rear part has a second identifier; Based on the input string and the answer inference model, determining the confidence of the first identifier and the confidence of the second identifier corresponding to each entity; Based on the confidence of the first identifier and the confidence of the second identifier corresponding to each entity, determining the confidence of each entity; Determining the entity with the highest confidence among each entity as the second entity corresponding to the query information, using the second entity to give feedback on the query information, and adding the triple composed of the first entity, the relationship information, and the second entity to the knowledge graph; Where each entity includes the same entity, and based on the input string and the answer inference model, determining the confidence of the first identifier and the confidence of the second identifier corresponding to each entity includes: Through the answer inference model, according to the input string, determining the output vector of the first identifier and the output vector of the second identifier corresponding to the same entity; For the first identifier or the second identifier corresponding to the same entity, averaging the corresponding output vectors to obtain an average vector, determining a score based on the average vector, and normalizing the score to obtain a confidence.
2. The method according to claim 1, wherein The association information selection model includes a primary selection sub-model and a refined selection sub-model. Based on the existing triples and the association information selection model in the knowledge graph, determining the target subgraph in the knowledge graph associated with the query information includes: Based on the existing triples in the knowledge graph and the primary selection sub-model, determining a candidate subgraph in the knowledge graph associated with the query information; Based on the candidate subgraph and the refined selection sub-model, determining the target subgraph in the knowledge graph associated with the query information.
3. The method according to claim 1, wherein Both the association information selection model and the answer inference model include a pre-trained language model network with an attention structure, and the pre-trained language model network with an attention structure is a bidirectional encoder representation BERT network based on a translator.
4. The method according to claim 1, wherein Before obtaining the query information, it further includes: Obtaining sample triples, where the sample triples are triples inferred based on the existing triples in the knowledge graph. The sample triples include a sample first entity, sample relationship information, and a sample second entity; Based on the existing triples in the knowledge graph and the association information selection model to be trained, determining a predicted subgraph in the knowledge graph associated with the sample query information, where the sample query information includes the sample first entity and the sample relationship information; Determining the predicted triples included in the predicted subgraph; Based on the sample query information, the predicted triple, and the answer inference model to be trained, determine the predicted second entity corresponding to the sample query information; Based on the sample second entity and the predicted second entity, adjust the model parameters of the association information selection model to be trained and the answer inference model to be trained.
5. An apparatus for knowledge graph completion, characterized in that The device includes: An acquisition module, configured to acquire query information, where the query information includes a first entity and relationship information; A determination module, configured to determine a target subgraph associated with the query information in the knowledge graph based on the existing triples in the knowledge graph and the association information selection model; determine the target triples included in the target subgraph; establish an input string based on the query information and the target triples, in the input string, the front part of the substring corresponding to each entity of the target triples has a first identifier, and the rear part has a second identifier; determine the confidence of the first identifier and the confidence of the second identifier corresponding to each entity based on the input string and the answer inference model; determine the confidence of each entity based on the confidence of the first identifier and the confidence of the second identifier corresponding to each entity; determine the entity with the highest confidence in each entity as the second entity corresponding to the query information, use the second entity to feedback the query information, and add the triple composed of the first entity, the relationship information, and the second entity to the knowledge graph; Wherein, each entity includes the same entity, and the determination module is configured to: determine the output vector of the first identifier and the output vector of the second identifier corresponding to the same entity according to the input string through the answer inference model; for the first identifier or the second identifier corresponding to the same entity, obtain an average vector by averaging the corresponding output vectors, determine a score based on the average vector, and normalize the score to obtain a confidence.
6. The device according to claim 5, characterized in that The association information selection model includes a primary selection sub-model and a refined selection sub-model; The determination module is configured to determine a candidate subgraph associated with the query information in the knowledge graph based on the existing triples in the knowledge graph and the primary selection sub-model; Based on the candidate subgraph and the refined selection sub-model, determine the target subgraph associated with the query information in the knowledge graph.
7. The device according to claim 5, characterized in that Both the association information selection model and the answer inference model include a pre-trained language model network with an attention structure, and the pre-trained language model network with an attention structure is a bidirectional encoding representation BERT network based on a translator.
8. The device according to claim 5, characterized in that, The device is further configured to: Acquire sample triples, where the sample triples are triples inferred based on the existing triples in the knowledge graph, and the sample triples include a sample first entity, sample relationship information, and a sample second entity; Based on the existing triples in the knowledge graph and the association information selection model to be trained, determine a predicted subgraph associated with the sample query information in the knowledge graph, where the sample query information includes the sample first entity and the sample relationship information; Determine the predictive triples included in the predictive sub-graph; Based on the sample query information, the predictive triples, and the answer inference model to be trained, determine the predictive second entity corresponding to the sample query information; Based on the sample second entity and the predictive second entity, adjust the model parameters of the association information selection model to be trained and the answer inference model to be trained.
9. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one instruction is stored in the memory. The at least one instruction is loaded and executed by the processor to implement the operations performed by the method for knowledge graph completion according to any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that, At least one instruction is stored in the storage medium. The at least one instruction is loaded and executed by a processor to implement the operations performed by the method for knowledge graph completion according to any one of claims 1 to 4.
11. A computer program product, characterized in that, The computer program product includes at least one instruction. The at least one instruction is loaded and executed by a processor to implement the operations performed by the method for knowledge graph completion according to any one of claims 1 to 4.
Citation Information
Patent Citations
United query network space knowledge graph reasoning method and device
CN111813949A
Knowledge graph reasoning method and device and storage medium
CN112084344A