Knowledge Graph Update Method, Device, Equipment, and Computer Readable Storage Medium
Through the methods of entity relationship inference and path relationship matching, the matching degree of new knowledge entities in the knowledge graph is improved, the problem of low matching degree in the existing technology is solved, and the accuracy of the knowledge graph is ensured.
Patent Information
- Application Number
- CN202210234408.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-10
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-03-10
AI Technical Summary
When the knowledge graph is completed and updated, the existing technology mines new knowledge entities based on the knowledge entity relationship, resulting in low matching and affecting the accuracy of the knowledge graph.
By obtaining the entity to be inferred, using the first target submodel to perform entity relationship inference, generating an entity relationship submap, and using the second target submodel to perform path relationship matching of the entity relationship submap, obtaining a set of matching scores, determining the largest target match score to determine the target result entity, and updating it to the knowledge graph.
Improve the matching degree of newly mined knowledge entity pairs, ensure the accuracy of the knowledge graph without the need for a large number of existing knowledge entity samples.
Smart Images

Figure CN114647740B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method, apparatus, device, and computer-readable storage medium for updating a knowledge graph. Background Art
[0002] A knowledge graph is a semantic network that reveals the relationships between knowledge entities, belonging to a way of using visualization or structuring to describe knowledge entities and entity relationships. This knowledge graph has the ability to organize and learn a large amount of structured and unstructured information. To improve the coverage of knowledge information in the knowledge graph, the knowledge graph can be completed and updated.
[0003] When the related art completes and updates the knowledge graph, it mainly models based on the entity relationship instance data of a large number of knowledge entities in the knowledge graph, and infers other entities based on the obtained model, so as to complete and update the knowledge graph according to the obtained other entities, so as to improve the coverage of knowledge information in the knowledge graph.
[0004] In the process of researching and practicing the prior art, the inventors of this application found that when the prior art completes and updates the knowledge graph, it mainly mines new knowledge entity pairs based on knowledge entity relationships. However, due to the limitations of knowledge entity relationships, it will affect the matching degree of the newly mined knowledge entity pairs, thereby affecting the accuracy of the knowledge graph. Summary of the Invention
[0005] Embodiments of this application provide a method, apparatus, device, and computer-readable storage medium for updating a knowledge graph, which can improve the matching degree of newly mined knowledge entity pairs and ensure the accuracy of the knowledge graph.
[0006] Embodiments of this application provide a method for updating a knowledge graph, including:
[0007] Obtain an entity to be inferred;
[0008] Perform entity relationship reasoning on the entity to be inferred through the first target sub-model to obtain an entity relationship subgraph between the entity to be inferred and the result entity, where the first target sub-model is obtained by performing entity relationship training on a preset inference entity sample and a sample entity relationship subgraph between the preset inference entity sample and the preset result entity sample;
[0009] Performing path relation matching on the entity relation subgraph through the second target submodel to obtain a set of matching scores, where the set of matching scores includes the matching scores of the path relations between each of the result entities and the entity to be inferred. The second target submodel is obtained by performing sample path relation training on the sample matching scores of the sample path relations between the sample entity relation subgraph, the preset inference entity samples, and the preset result entity samples;
[0010] Determining the target matching score with the largest score from the set of matching scores, and determining the target result entity corresponding to the entity to be inferred according to the target matching score;
[0011] Establishing the target entity relation between the entity to be inferred and the target result entity, and updating the target entity relation to the knowledge graph.
[0012] Correspondingly, an embodiment of the present application provides a knowledge graph updating device, including:
[0013] An acquisition unit, configured to acquire an entity to be inferred;
[0014] An inference unit, configured to perform entity relation inference on the entity to be inferred through the first target submodel to obtain an entity relation subgraph between the entity to be inferred and a result entity. The first target submodel is obtained by performing entity relation training on the preset inference entity samples and the sample entity relation subgraph between the preset inference entity samples and the preset result entity samples;
[0015] A matching unit, configured to perform path relation matching on the entity relation subgraph through the second target submodel to obtain a set of matching scores, where the set of matching scores includes the matching scores of the path relations between each of the result entities and the entity to be inferred. The second target submodel is obtained by performing sample path relation training on the sample matching scores of the sample path relations between the sample entity relation subgraph, the preset inference entity samples, and the preset result entity samples;
[0016] A determining unit, configured to determine the target matching score with the largest score from the set of matching scores, and determine the target result entity corresponding to the entity to be inferred according to the target matching score;
[0017] An updating unit, configured to establish the target entity relation between the entity to be inferred and the target result entity, and update the target entity relation to the knowledge graph.
[0018] In some embodiments, the inference unit is further configured to:
[0019] Extracting the preset entity relation subgraph of each preset entity pair in the knowledge graph through the first target submodel;
[0020] Based on the preset entity relationship sub-graph, reason about the entity to be inferred to obtain the entity relationship sub-graph between the entity to be inferred and the result entity.
[0021] In some embodiments, the inference unit is further configured to:
[0022] Obtain the preset path relationship of the preset entity pair in the preset entity relationship sub-graph;
[0023] Expand the entity to be inferred according to the preset path relationship to obtain each result entity corresponding to the entity to be inferred;
[0024] Generate an entity relationship sub-graph between the entity to be inferred and the result entity according to the entity to be inferred, the result entity and the preset path relationship.
[0025] In some embodiments, the matching unit is configured to:
[0026] Extract the path relationship between the entity to be inferred and each result entity in the entity relationship sub-graph through the second target sub-model, and encode each path relationship to obtain path relationship information corresponding to each path relationship;
[0027] Determine the matching score between the entity to be inferred and each result entity according to each path relationship information to obtain the matching score set.
[0028] In some embodiments, the matching unit is further configured to:
[0029] Extract each preset path relationship of the preset entity pair from the preset entity relationship sub-graph, and encode the preset path relationship information corresponding to each preset path relationship;
[0030] Calculate the interaction coefficient between the path relationship information and the preset path relationship information, and determine the interaction coefficient as the matching score between the entity to be inferred and the corresponding result entity.
[0031] In some embodiments, the matching unit is further configured to:
[0032] Generate a path relationship information matrix according to the path relationship information and the preset path relationship information;
[0033] Process the path relationship information matrix through a preset attention mechanism to obtain a processed path similarity matrix;
[0034] Perform max pooling processing on the path similarity matrix to obtain a one-dimensional matrix;
[0035] Perform mapping processing on the one-dimensional matrix to obtain a mapped multi-dimensional matrix;
[0036] Perform summation processing on the parameters in the multi-dimensional matrix column by column to obtain a multi-dimensional similarity matrix;
[0037] Perform multi-layer perception processing on the multi-dimensional similarity matrix to obtain the interaction coefficient between the path relationship information and the preset path relationship information.
[0038] In some embodiments, the knowledge graph updating device further includes a training unit for:
[0039] Obtain a preset inference entity and a preset result entity in the knowledge graph, and set a sample matching score between the preset inference entity and the preset result entity;
[0040] Input the preset inference entity into a preset model to obtain a predicted matching score between the preset inference entity and the preset result entity;
[0041] Obtain a matching difference value between the predicted matching score and the sample matching score;
[0042] Iteratively train the preset model based on the matching difference value until the matching difference value converges to obtain a trained target model, where the target model includes the first target sub-model and the second target sub-model.
[0043] In addition, an embodiment of the present application further provides a computer device, including a processor and a memory, where the memory stores a computer program, and the processor is configured to run the computer program in the memory to implement the steps in the knowledge graph updating method provided by the embodiment of the present application.
[0044] In addition, an embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in any one of the knowledge graph updating methods provided by the embodiment of the present application.
[0045] In addition, an embodiment of the present application further provides a computer program product, including computer instructions, where the computer instructions, when executed, implement the steps in any one of the knowledge graph updating methods provided by the embodiment of the present application.
[0046] Embodiments of the present application can obtain an entity to be inferred; perform entity relationship reasoning on the entity to be inferred through a first target sub-model to obtain an entity relationship sub-graph between the entity to be inferred and a result entity, where the first target sub-model is obtained by performing entity relationship training on a preset inference entity sample and a sample entity relationship sub-graph between the preset inference entity sample and a preset result entity sample; perform path relationship matching on the entity relationship sub-graph through a second target sub-model to obtain a set of matching scores, where the set of matching scores includes the matching scores of the path relationship between each result entity and the entity to be inferred, and the second target sub-model is obtained by performing sample path relationship training on the sample entity relationship sub-graph and the sample matching scores of the sample path relationship between the preset inference entity sample and the preset result entity sample; determine the target matching score with the largest score from the set of matching scores, and determine the target result entity corresponding to the entity to be inferred according to the target matching score; establish a target entity relationship between the entity to be inferred and the target result entity, and update the target entity relationship to the knowledge graph. Thus, it can be seen that this solution can infer the entity relationship sub-graph between the entity to be inferred and the result entity, and match the path relationship between the entity to be inferred and the result entity based on the entity relationship sub-graph to obtain a set of matching scores. Furthermore, select the target result entity with the largest matching score for the entity to be inferred and update it to the knowledge graph; in this way, by combining the path relationship and the entity relationship for knowledge entity mining, the matching degree when mining new knowledge entities is improved, without a large number of existing knowledge entity samples, ensuring the accuracy of the knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0048] Figure 1 is a schematic diagram of the scenario of the knowledge graph update system provided by the embodiments of the present application;
[0049] Figure 2 is a schematic flowchart of the steps of the knowledge graph update method provided by the embodiments of the present application;
[0050] Figure 3 is another schematic flowchart of the steps of the knowledge graph update method provided by the embodiments of the present application;
[0051] Figure 4 is a schematic diagram of the path relationship of the preset entity pair provided by the embodiments of the present application;
[0052] Figure 5 is a schematic diagram of the scenario of the knowledge graph update method provided by the embodiments of the application;
[0053] Figure 6 Schematic diagram of the interaction calculation scenario of the path relationship information of the knowledge graph update method provided for the application example;
[0054] Figure 7 Schematic diagram of the structure of the knowledge graph update device provided by the embodiments of the present application;
[0055] Figure 8 Schematic diagram of the structure of the computer device provided by the embodiments of the present application. Detailed implementation manners
[0056] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0057] The embodiments of the present application provide a knowledge graph update method, device, device, and computer-readable storage medium. Specifically, the embodiments of the present application will be described from the perspective of the knowledge graph update device. The knowledge graph update device can be specifically integrated in a computer device, which can be a server or a user terminal or other devices. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Among them, the user terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a smart home appliance, a vehicle terminal, a smart voice interaction device, an aircraft, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make limitations here.
[0058] The knowledge graph update method provided by the embodiments of the present application can be applied to various scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving. It will be specifically described through the following embodiments:
[0059] For example, refer to Figure 1 , which is a schematic diagram of the scenario of the knowledge graph update system provided by the embodiments of the present application. This scenario includes a terminal or a server.
[0060] A terminal or server can obtain an entity to be inferred; perform entity relationship inference on the entity to be inferred through a first target sub-model to obtain an entity relationship sub-graph between the entity to be inferred and a result entity, where the first target sub-model is obtained by performing entity relationship training on a preset inference entity sample and a sample entity relationship sub-graph between the preset inference entity sample and a preset result entity sample; perform path relationship matching on the entity relationship sub-graph through a second target sub-model to obtain a set of matching scores, where the set of matching scores includes the matching scores of the path relationship between each result entity and the entity to be inferred, and the second target sub-model is obtained by performing sample path relationship training on the sample entity relationship sub-graph and the sample matching scores of the sample path relationship between the preset inference entity sample and the preset result entity sample; determine the target matching score with the largest score from the set of matching scores, and determine the target result entity corresponding to the entity to be inferred according to the target matching score; establish a target entity relationship between the entity to be inferred and the target result entity, and update the target entity relationship to the knowledge graph.
[0061] Among them, the process of updating the knowledge graph may include: obtaining / receiving an entity to be inferred, inferring the entity relationship sub-graph of the entity to be inferred, determining the matching score between each result entity and the entity to be inferred, determining the final target result entity of the entity to be inferred, and updating it to the knowledge graph.
[0062] The following will be described in detail respectively. It should be noted that the order of the following embodiments does not limit the preferred order of the embodiments.
[0063] In the embodiments of the present application, a description will be made from the perspective of a knowledge graph update device, where the knowledge graph update device may be specifically integrated in a computer device such as a terminal or a server. Refer to Figure 2 , Figure 2 which is a schematic flowchart of the steps of a knowledge graph update method provided by the embodiments of the present application. Taking the case where the knowledge graph update device is specifically integrated on a server as an example, when the processor on the server executes the program corresponding to the information processing method, the specific process is as follows:
[0064] 101. Obtain an entity to be inferred.
[0065] Among them, the entity to be inferred may be the entity to be queried input by the user when performing knowledge entity query, or may be a knowledge entity that needs to re-establish an entity relationship. For example, taking the knowledge graph as an example, the entity to be inferred belongs to the entities in the knowledge graph, may belong to the knowledge entities with established entity relationships in the knowledge graph, or may belong to the instruction entities in the knowledge graph that have not established entity relationships with other entities.
[0066] Among them, a knowledge graph is a semantic network that describes entities and the relationships between entities. For example, the entities can be information entities such as people, places, organizations, concepts, etc., and the types of relationships can be relationships between people, relationships between people and organizations, relationships between concepts and a certain object, and so on. The above are only examples and are not limited here.
[0067] It should be noted that with the development of information data, the amount of knowledge information is also increasing, and the amount of knowledge entities contained or covered in the knowledge graph may be insufficient. In order to expand the knowledge graph, the knowledge graph can be updated or completed in real time or regularly, such as adding or completing the entity relationships between knowledge entities to expand the knowledge graph and improve the coverage of the knowledge graph.
[0068] In order to expand the knowledge graph subsequently, the knowledge graph can be updated and completed continuously or intermittently. For example, specifically complete the knowledge entities in the knowledge graph, which can specifically be to establish the entity relationship between this knowledge entity and other entities that can establish relationships, so as to complete or update the knowledge entity pair with the newly established entity relationship to the knowledge graph. Therefore, before updating and completing the knowledge graph, it is necessary to first determine the entity to be completed, that is, the entity to be inferred, in order to subsequently infer the knowledge entity that matches the entity to be inferred for use in updating and completing the knowledge graph.
[0069] Specifically, the way to obtain the entity to be inferred can be: receiving the text input by the user during information query, and this text can be text, fields, numbers, formulas, etc., and determining the entity to be inferred according to the received text. In addition, the set of knowledge entities contained in the knowledge graph can also be obtained, and the knowledge sub-entities to be established with entity relationships are screened out from the set of knowledge entities, and the knowledge sub-entities to be established with entity relationships are determined as the entities to be inferred.
[0070] Through the above methods, the entity to be inferred can be obtained, so as to subsequently infer the knowledge entity that matches the entity to be inferred to update and complete the knowledge graph.
[0071] It should be noted that after obtaining the entity to be inferred, it is necessary to determine the target result entity that matches the entity to be inferred, so as to subsequently establish the target entity relationship between the entity to be inferred and the target result entity and update and complete it to the knowledge graph. Among them, before determining the target result entity that matches the entity to be inferred, the matching score between the entity to be inferred and other result entities can be determined through the target model.
[0072] Specifically, the target model processes the entity to be inferred to obtain a set of matching scores. The set of matching scores includes the matching scores between the entity to be inferred and each result entity. Among them, the target model is trained by sample path relationships using preset inference entity samples, preset result entity samples, and sample matching scores. The matching score represents the matching degree between the path relationship and the sample path relationship, and the path relationship represents the path attribute between the entity to be inferred and the result entity.
[0073] Among them, the target model can be an independent model or a model jointly obtained by multiple sub-models. To determine that the target model can infer the result entity corresponding to the entity to be inferred and accurately score the matching degree between the entity to be inferred and the result entity at the same time, the embodiments of this application need to train the model so as to use the trained target model to perform entity inference and scoring on the entity to be inferred.
[0074] Specifically, before the step of "performing entity relationship inference on the entity to be inferred through the first target sub-model to obtain the entity relationship sub-graph between the entity to be inferred and the result entity", it may further include: obtaining the preset inference entity and the preset result entity in the knowledge graph, and setting the sample matching score between the preset inference entity and the preset result entity; inputting the preset inference entity into the preset model to obtain the predicted matching score between the preset inference entity and the preset result entity; obtaining the matching difference value between the predicted matching score and the sample matching score; and performing iterative training on the preset model based on the matching difference value until the matching difference value converges to obtain the trained target model.
[0075] Among them, the target model may include a first target sub-model and a second target sub-model. For the implementation processes of the first target sub-model and the second target sub-model, please refer to the following steps 102 and 103.
[0076] 102. Perform entity relationship inference on the entity to be inferred through the first target sub-model to obtain the entity relationship sub-graph between the entity to be inferred and the result entity.
[0077] Among them, the first target sub-model is trained by entity relationships using preset inference entity samples and the sample entity relationship sub-graph between the preset inference entity samples and the preset result entity samples. The trained first target sub-model is used to perform entity relationship inference on the entity to be inferred to determine the result entity that can establish a relevant entity relationship with the entity to be inferred, and thus, output the entity relationship sub-graph between the entity to be inferred and the result entity.
[0078] Among them, the result entity can refer to the knowledge entity corresponding to the entity to be inferred, which is inferred by the model based on the existing small number of entity relationships in the knowledge graph. This result entity can be understood as the final result entity to be expanded based on the entity to be inferred. It should be noted that when inferring the result entity, there is no need to consider the matching degree between the entity to be inferred and the result entity, as long as it conforms to the relevant entity relationships.
[0079] Among them, the entity relationship subgraph includes the entity relationship between the entity to be inferred and the corresponding result entity. In the embodiments of the present application, the entity relationship can be directly or indirectly represented by one or more path relationships between the entity to be inferred and the corresponding result entity; among them, the path relationship can be composed of one or more sub-path relationships. For example, the entity relationship between user A and user B refers to the legal spouse relationship in reality, the entity relationship between user A and user C refers to the mother-child relationship in reality, and the entity relationship between user B and user C refers to the legal stepfather-stepson relationship in reality. Based on this, taking the path of the "mother-child relationship" between user A and user C as the first sub-path relationship and the path of the "legal stepfather-stepson relationship" between user B and user C as the second sub-path relationship, the first sub-path relationship and the second sub-path relationship form the path relationship between user A and user B, which is used to represent the entity relationship (legal spouse) between user A and user B.
[0080] In order to obtain the entity relationship subgraph corresponding to the entity to be inferred, in the embodiments of the present application, after obtaining the entity to be inferred, the first target sub-model is used to perform entity relationship inference on the entity to be inferred, so as to infer the entity relationship subgraph between the entity to be inferred and the result entity, so as to facilitate subsequent evaluation of the matching degree or matching score between the entity to be inferred and the result entity based on this entity relationship subgraph.
[0081] In some embodiments, the step of "performing entity relationship inference on the entity to be inferred through the first target sub-model to obtain the entity relationship subgraph between the entity to be inferred and the result entity" may include:
[0082] (102.1) Extract the preset entity relationship subgraph of each preset entity pair in the knowledge graph through the first target sub-model.
[0083] (102.2) Based on the preset entity relationship subgraph, perform inference on the entity to be inferred to obtain the entity relationship subgraph between the entity to be inferred and the result entity.
[0084] Among them, the preset entity pair can be a knowledge entity pair with corresponding preset entity relationships established in the knowledge graph.
[0085] Among them, the preset entity relationship subgraph can be a feature graph or information graph representing the entity relationship of the preset entity pair.
[0086] In order to obtain an entity relationship sub-graph related to the entity to be inferred, the embodiments of the present application can infer the entity to be inferred based on the entity relationships of existing entity pairs in the knowledge graph. Specifically, one or more preset entity pairs in the current knowledge graph are extracted through a first target sub-model to identify the preset entity relationship sub-graph of the extracted preset entity pairs; furthermore, based on the preset entity relationship sub-graph, result entity inference is performed on the entity to be inferred to determine the result entities that can establish entity relationships with the entity to be inferred, and then, an entity relationship sub-graph between each result entity and the entity to be inferred is generated.
[0087] For example, taking a certain knowledge entity in the knowledge graph as the head entity, that is, the preset inference entity, path query and expansion are performed on the preset inference entity. For example, using two-side BFS to expand the path of the preset entity, different lengths of preset paths from the preset inference entity to each preset result entity are obtained. Then, the preset inference entity, the preset paths, and the preset result entities are merged to obtain a preset entity relationship sub-graph including the path relationships between each preset result entity and the preset inference entity. Further, based on the preset entity relationship sub-graph, inference is performed on the entity to be inferred to obtain the entity relationship sub-graph corresponding to the entity to be inferred.
[0088] In some embodiments, the step of "inferring the entity to be inferred based on the preset entity relationship sub-graph to obtain the entity relationship sub-graph between the entity to be inferred and the result entity" may include:
[0089] (102.2.1) Obtain the preset path relationship of the preset entity pairs in the preset entity relationship sub-graph.
[0090] Wherein, the preset path relationship is used to represent the path association relationship between the head entity (preset inference entity) and the tail entity (preset result entity) of the preset entity pair. For example, the entity relationship between user A and user B refers to the legal spouse relationship in reality, the entity relationship between user A and user C refers to the mother-child relationship in reality, and the entity relationship between user B and user C refers to the legal stepfather-stepson relationship in reality. Based on this, taking the path of the "mother-child relationship" between user A and user C as the first sub-path relationship and the path of the "legal stepfather-stepson relationship" between user B and user C as the second sub-path relationship, the first sub-path relationship and the second sub-path relationship form the path relationship between user A and user B, which is used to represent the entity relationship between user A and user B.
[0091] (102.2.2) Expand the entity to be inferred according to the preset path relationship to obtain each result entity corresponding to the entity to be inferred.
[0092] Specifically, the process of expanding the entity to be inferred is the main process of reasoning. Specifically, the entity to be inferred can be expanded according to the preset path relationship to obtain a result entity that has a similar preset path relationship with the entity to be inferred.
[0093] It should be noted that, in order to avoid the subsequent entity relationship subgraph from being too large, a preset expansion quantity threshold can be set during reasoning / expansion to limit the number of neighbor entities expanded per hop. Specifically, the entity to be inferred can be expanded in multiple hops according to the preset expansion quantity threshold to obtain the expanded result entity, so as to facilitate the subsequent generation of the relevant entity relationship subgraph. In this way, the problem that the subsequently synthesized entity relationship subgraph is too large can be effectively avoided.
[0094] Among them, when expanding the entity to be inferred in multiple hops according to the preset expansion quantity threshold, each hop of expansion can select the number of entities corresponding to the preset expansion quantity threshold for expansion. It should be noted that, in order to make the inferred entity relationship subgraph more accurate, when expanding the entity to be inferred in each hop, the sub-path relationship between the entity to be inferred and each surrounding neighbor entity can be determined respectively, and the similarity between the sub-path relationship and the corresponding preset sub-path relationship in the preset entity relationship subgraph can be calculated. Thus, multiple target neighbor entities with higher similarity are selected from the multiple neighbor entities around the entity to be inferred; until the multiple-hop expansion is completed, a result entity and an entity sequence from the entity to be inferred to the result entity are obtained, where the entity sequence can only include the entity to be inferred and the result entity; the entity sequence can also include the entity to be inferred, one or more intermediate entities, and the result entity, and the intermediate entity can be regarded as the result entity of the current hop of expansion. Through the above, the relationship expansion of the entity to be inferred is completed, so as to facilitate the subsequent generation of the corresponding entity relationship subgraph according to the entity to be inferred, the expanded result entity, and the preset path relationship.
[0095] (102.2.3) Generate an entity relationship subgraph between the entity to be inferred and the result entity according to the entity to be inferred, the result entity, and the preset path relationship.
[0096] It can be understood that the entity relationship subgraph includes the path relationship between the entity to be inferred and the result entity. It should be noted that the entity relationship subgraph can include the entity relationship between the entity to be inferred and one or more corresponding result entities.
[0097] In order to obtain the entity relationship subgraph between the entity to be inferred and the corresponding result entity, after obtaining the preset path relationship and one or more result entities corresponding to the entity to be inferred in the embodiments of the present application, the path between the entity to be inferred and the result entity can be established according to the path corresponding to the preset path relationship, so as to construct an entity relationship subgraph including the entity relationship between the entity to be inferred and the result entity.
[0098] In the above manner, entity relationship reasoning can be performed on the entity to be inferred to infer the corresponding result entity, so as to synthesize the entity relationship subgraph between the entity to be inferred and the result entity.
[0099] 103. Perform path relationship matching on the entity relationship subgraph through the second target submodel to obtain a set of matching scores.
[0100] Among them, the set of matching scores includes the matching scores of the path relationships between each result entity and the entity to be inferred.
[0101] Among them, the second target submodel is obtained by training the sample path relationship with the sample matching scores of the sample path relationships between the sample entity relationship subgraph, the preset inference entity sample, and the preset result entity sample. It should be noted that the trained second target submodel belongs to a scoring model, which is used to extract the path relationships between the entity to be inferred and each result entity in the entity relationship subgraph, and represent the path information of each path relationship, so as to determine the interaction coefficient between each path and the preset path according to the path information. Thus, the matching score of the path relationship between each result entity and the entity to be inferred is calculated.
[0102] Among them, the path relationship may include one or more sub-path relationships experienced when extending from the entity to be inferred to the result entity, and each sub-path relationship can represent or reflect the entity relationship between two adjacent knowledge entities during the extension process. Exemplarily, the entity relationship between user A and user B refers to the legal spouse relationship in reality, the entity relationship between user A and user C refers to the mother-child relationship in reality, and the entity relationship between user B and user C refers to the legal stepfather-stepson relationship in reality; based on this, the path relationship of the "mother-child relationship" between user A and user C belongs to the sub-path relationship, which can reflect the entity relationship between user A and user C; similarly, the path relationship of the "legal stepfather-stepson relationship" between user B and user C also belongs to the sub-path relationship, which can reflect the entity relationship between user B and user C.
[0103] In order to determine the matching degree between the entity to be inferred and the relevant result entity, in the embodiment of the present application, after obtaining the entity relationship subgraph between the entity to be inferred and the result entity, the information of each path in the entity relationship subgraph is represented, so as to determine the interaction coefficient between each path relationship and the preset path relationship according to the path information. The interaction coefficient can reflect the similarity between the path relationship from the entity to be inferred to the relevant result entity and the preset path relationship. Furthermore, the matching degree between the two entities, that is, the matching score, can be determined according to the interaction coefficient between each result entity and the entity to be inferred, and a set of matching scores is obtained. In this way, it is convenient to subsequently select the target result entity that best matches the entity to be inferred according to the set of matching scores and update it to the knowledge graph.
[0104] In some embodiments, the step of "performing path relationship matching on the entity relationship subgraph through the second target submodel to obtain a set of matching scores" may include:
[0105] (103.1) Extract the path relationship between the entity to be inferred and each result entity in the entity relationship subgraph through the second target submodel, and encode each path relationship to obtain the path relationship information corresponding to each path relationship.
[0106] The path relationship information may be composed of one or more sub-path relationship information from the entity to be inferred to the result entity. Specifically, when there is only one sub-path relationship information from the entity to be inferred to the result entity, use the unique sub-path relationship information as the path relationship information; when there are multiple sub-path relationship information from the entity to be inferred to the result entity, then according to the expansion time sequence of each expansion direction, generate a path relationship sequence information from the multiple sub-path relationship information, and use this relationship sequence information as the path relationship information. This path relationship sequence information is a kind of relationship sequence pattern. For example, taking the case of containing multiple sub-path relationship information, the entity relationship information between user A and user B refers to legal spouses in reality. Taking user A as the entity to be inferred and expanding, it needs to go through the first hop to expand to user C. The sub-path relationship information between user C and user A is "mother-son", and the second hop expands from user C to user B. The sub-path relationship information between user C and user B is "legal stepfather-stepson". Then, generate a path relationship sequence information from these two sub-path relationship information, which is "mother-son -> legal stepfather-stepson". This path relationship sequence information is used as the path relationship information between user A and user B to represent the entity relationship between the two.
[0107] In order to obtain the path relationship information between the entity to be inferred and each result entity, after obtaining the relevant entity relationship subgraph in the embodiments of the present application, the path relationship between the entity to be inferred and each result entity can be extracted to represent the information of each path relationship, so as to obtain the path relationship information corresponding to each path relationship. Among them, this information representation process can be specifically implemented by an encoding method, such as encoding the embedding information of each sub-path relationship through a Gate Recurrent Unit (GRU), so as to generate the path relationship information according to the encoded information of each sub-path relationship.
[0108] (103.2) Determine the matching score between the entity to be inferred and each result entity according to each path relationship information to obtain a set of matching scores.
[0109] Among them, the matching score is used to represent the matching degree between the entity to be inferred and the relevant result entity. Through the matching score, the similarity between the path relationship information of the entity to be inferred and the preset path relationship information can be reflected. Among them, the similarity between the path relationship information and the preset path relationship information can be specifically determined according to the situation between the sub-path relationship information included in the path relationship information and the preset sub-path relationship information included in the preset path relationship information.
[0110] In some embodiments, the step of "determining the matching score between the entity to be inferred and each result entity according to each path relationship information to obtain a set of matching scores" may include:
[0111] (103.2.1) Extract each preset path relationship of the preset entity pair from the preset entity relationship subgraph, and encode the preset path relationship information corresponding to each preset path relationship.
[0112] (103.2.2) Calculate the interaction coefficient between the path relationship information and the preset path relationship information, and determine the interaction coefficient as the matching score between the entity to be inferred and the corresponding result entity.
[0113] Specifically, in order to determine the matching score between the entity to be inferred and each result entity, the embodiment of the present application may first extract the preset path relationship related to the preset entity pair (i.e., the preset inference entity and the preset result entity) from the preset entity relationship subgraph, and represent the information of each preset path relationship in an encoded manner to obtain the preset path relationship information; furthermore, calculate the interaction coefficient between the path relationship information and the preset path relationship information, and the interaction coefficient can reflect the difference between the sub-path relationship information included in the path relationship information and the preset sub-path relationship information included in the preset path relationship information. Furthermore, the interaction coefficient can be determined as the matching score between the entity to be inferred and the corresponding result entity.
[0114] For example, taking "User 1" and "User 2" as a preset entity pair, after the preset path relationship of the preset entity pair is encoded, the preset path relationship information between "User 1" and "User 2" includes two sub-path relationship information, namely "son of" and "father in law". Suppose the entity to be inferred is User 3 and the result entity is User 4, and the path relationship information between User 3 and User 4 includes two sub-path relationship information, namely "son of" and "father in law". Then the interaction situation between this path relationship information and the preset path relationship information is the same, and the interaction coefficient is 1. Therefore, the matching score between the entity to be inferred and the result entity is 1. If the path relationship information between User 3 and User 4 includes "son of" and "friend", then the interaction coefficient between this path relationship information and the preset path relationship information is 0.5, and the matching score is 0.5. If the path relationship information between User 3 and User 4 includes "acquaintance" and "friend", then there is no interaction between this path relationship information and the preset path relationship information, so the interaction coefficient is 0 and the matching score is 0. The above path relationship information, interaction coefficient, etc. are only examples, and the embodiments of the present application do not limit this.
[0115] In some embodiments, the step of "calculating the interaction coefficient between the path relationship information and the preset path relationship information" may include:
[0116] Generate a path relationship information matrix according to the path relationship information and the preset path relationship information; process the path relationship information matrix through a preset attention mechanism to obtain a processed path similarity matrix; perform max pooling processing on the path similarity matrix to obtain a one-dimensional matrix; perform mapping processing on the one-dimensional matrix to obtain a mapped multi-dimensional matrix; sum the parameters in the multi-dimensional matrix column by column to obtain a multi-dimensional similarity matrix; perform multi-layer perception processing on the multi-dimensional similarity matrix to obtain the interaction coefficient between the path relationship information and the preset path relationship information.
[0117] It should be noted that since there may be multiple preset path relationships between the preset inference entity and the preset result entity of the preset entity pair, that is, the preset inference entity and the preset result entity can be inferred through multiple preset path relationships. It should be noted that each preset path relationship may represent a different entity relationship between the preset inference entity and the preset result entity; and, there may be multiple path relationships between the entity pair to be inferred and the result entity. Based on this, when calculating the interaction coefficient, the interaction coefficient between any one path relationship information and different preset path relationship information can be calculated.
[0118] Specifically, based on multiple preset path relationship information and multiple path relationship information, a path relationship information matrix is established; in order to strengthen the influence of path relationship pairs that match the target entity relationship (i.e., the preset entity relationship) on the matching result, an attention mechanism is introduced, and the path relationship information matrix is multiplied by attention parameters or attention coefficients to obtain a processed path similarity matrix; then, pooling processing is performed on the path similarity matrix, and the pooling processing method can specifically be maxpooling to obtain a one-dimensional matrix after pooling processing, and the one-dimensional matrix is mapped (kernel), such as mapping the one-dimensional matrix to 21 dimensions to obtain a multi-dimensional matrix; furthermore, the parameters of each row in the multi-dimensional matrix are summed up column by column to obtain a multi-dimensional similarity matrix; finally, the multi-dimensional similarity matrix is subjected to multi-layer perception processing, such as 2-layer perception (mlp), to obtain a 1*1 coefficient, and this coefficient is the interaction coefficient between the path relationship information of the current inference entity pair (the entity to be inferred and the corresponding result entity) and the preset path relationship information of the preset entity pair. Finally, the calculated interaction coefficient is determined as the matching score between the entity to be inferred and the corresponding result entity.
[0119] Through the above method, the matching score between the entity to be inferred and each result entity can be determined, obtaining a matching score set containing multiple matching scores, so as to subsequently determine the target result entity that best matches the entity to be inferred according to the matching score set.
[0120] 104. Determine the target matching score with the largest score from the matching score set, and determine the target result entity corresponding to the entity to be inferred according to the target matching score.
[0121] After the embodiment of the present application performs entity relationship reasoning and path relationship matching on the entity to be inferred, the matching score between the entity to be inferred and the corresponding result entity is obtained. It should be noted that when performing entity relationship reasoning, there may be multiple preset entity relationships of preset entity pairs applicable to the reasoning process of the entity to be inferred. After reasoning, there may be multiple result entities that can form corresponding entity relationships with the entity to be inferred; at this time, after performing path relationship matching, there may be multiple matching scores, and each matching score corresponds to an inference entity pair.
[0122] When there are multiple preset entity relationships applicable to the reasoning process of the entity to be reasoned, there will be multiple matching scores after the path relationship matching. It should be noted that the higher the matching score, the more likely the target result entity corresponding to the reasoning entity pair corresponding to the matching score is the correct entity answer. Therefore, in order to determine the target result entity that best matches the entity to be reasoned, the embodiments of the present application can select the maximum matching score from multiple matching scores as the target matching score. Thus, based on the target matching score, the target result entity corresponding to the entity to be reasoned is determined. In this way, the accuracy in determining the result entity corresponding to the entity to be reasoned is improved.
[0123] Through the above method, the target result entity that best matches the entity to be reasoned can be selected from multiple result entities, ensuring the accuracy in reasoning the result entity corresponding to the entity to be reasoned. Furthermore, the matching degree of the subsequently updated and complemented knowledge entities is improved, and the accuracy of the updated knowledge graph is improved, which is reliable.
[0124] 105. Establish a target entity relationship between the entity to be reasoned and the target result entity, and update the target entity relationship to the knowledge graph.
[0125] Specifically, after obtaining the target result entity, the embodiments of the present application can establish a target entity relationship according to the entity to be reasoned and the target result entity. The target entity relationship can be a preset entity relationship, and update the new knowledge entity pair composed of the entity to be reasoned and the target result entity with the target entity relationship to the knowledge graph. Or, a target entity relationship between the entity to be reasoned and the target result entity in the knowledge graph can also be established. In this way, there is a target entity relationship between the entity to be reasoned and the target result entity in the knowledge graph, realizing the update and complement of the knowledge graph.
[0126] As can be seen from the above, embodiments of the present application can obtain an entity to be inferred; perform entity relationship inference on the entity to be inferred through a first target sub-model to obtain an entity relationship sub-graph between the entity to be inferred and a result entity, where the first target sub-model is obtained by performing entity relationship training on a preset inference entity sample and a sample entity relationship sub-graph between the preset inference entity sample and a preset result entity sample; perform path relationship matching on the entity relationship sub-graph through a second target sub-model to obtain a set of matching scores, where the set of matching scores includes the matching scores of the path relationship between each result entity and the entity to be inferred, and the second target sub-model is obtained by performing sample path relationship training on the sample entity relationship sub-graph and the sample matching scores of the sample path relationship between the preset inference entity sample and the preset result entity sample; determine the target matching score with the largest score from the set of matching scores, and determine the target result entity corresponding to the entity to be inferred according to the target matching score; establish a target entity relationship between the entity to be inferred and the target result entity, and update the target entity relationship to the knowledge graph. Thus, the present solution can infer an entity relationship sub-graph between the entity to be inferred and the result entity, and perform matching on the path relationship between the entity to be inferred and the result entity based on the entity relationship sub-graph to obtain a set of matching scores. Furthermore, select the target result entity with the largest matching score with the entity to be inferred and update it to the knowledge graph; in this way, combine the path relationship and the entity relationship for knowledge entity mining, improve the matching degree when mining new knowledge entities, and do not require a large number of existing knowledge entity samples to ensure the accuracy of the knowledge graph.
[0127] According to the method described in the above embodiments, the following will give further detailed descriptions by way of examples.
[0128] Taking the update of the knowledge graph as an example, embodiments of the present application further describe the knowledge graph update method provided by the embodiments of the present application.
[0129] Figure 3 It is another step flow schematic diagram of the knowledge graph update method provided by embodiments of the present application. Figure 4 It is a schematic diagram of the path relationship of a preset entity pair provided by embodiments of the present application. Figure 5 It is a schematic diagram of the scenario of the knowledge graph update method provided by embodiments of the application. Figure 6 It is a schematic diagram of the interactive calculation scenario of the path relationship information of the knowledge graph update method provided by embodiments of the application. For ease of understanding, embodiments of the present application are described in conjunction with Figures 3 - 6 for description.
[0130] In embodiments of the present application, descriptions will be made from the perspective of a knowledge graph update device, which can be specifically integrated in a computer device such as a terminal and / or a server. When a processor on the computer device executes a program corresponding to the knowledge graph update method, the specific process of the knowledge graph update method is as follows:
[0131] 201. Jointly train a preset model based on a preset inference entity, a preset result entity, and a sample matching score between the preset inference entity and the preset result entity to obtain a trained target model.
[0132] It should be noted that the preset model can be a model jointly obtained by multiple sub-models, and the trained target model can include a first target sub-model and a second target sub-model.
[0133] Among them, the first target sub-model is used to perform entity relationship reasoning on the entity to be inferred to determine a result entity that can establish a relevant entity relationship with the entity to be inferred, and thus output an entity relationship sub-graph between the entity to be inferred and the result entity.
[0134] Among them, the second target sub-model belongs to a scoring model, which is used to extract the path relationship between the entity to be inferred and each result entity in the entity relationship sub-graph and represent the path information of each path relationship, so as to determine the interaction coefficient between each path and the preset path according to the path information, and thus calculate the matching score of the path relationship between each result entity and the entity to be inferred.
[0135] 202. Obtain the entity to be inferred.
[0136] Among them, a knowledge graph is a semantic network that describes entities and the relationships between entities. For example, entities can be information entities such as people, places, organizations, concepts, etc., and the types of relationships can be relationships between people, relationships between people and organizations, relationships between concepts and a certain object, etc. The above are only examples and are not limited here.
[0137] For example, the way to obtain the entity to be inferred can be: receive the text input by the user during information query, and the text can be text, fields, numbers, formulas, etc., and determine the entity to be inferred according to the received text. In addition, it is also possible to obtain the set of knowledge entities included in the knowledge graph, screen out the knowledge sub-entities to be established with entity relationships from the set of knowledge entities, and determine the knowledge sub-entities to be established with entity relationships as the entity to be inferred.
[0138] 203. Extract the preset entity relationship sub-graph of each preset entity pair in the knowledge graph through the first target sub-model, and obtain the preset path relationship of the preset entity pair in the preset entity relationship sub-graph.
[0139] Among them, the preset entity pair can be a knowledge entity pair that has established a corresponding preset entity relationship in the knowledge graph.
[0140] Among them, the preset entity relationship sub-graph can be a feature graph or information graph representing the entity relationship of the preset entity pair.
[0141] Among them, the preset path relationship is used to represent the path association relationship between the head entity (preset inference entity) and the tail entity (preset result entity) of the preset entity pair. For example, the entity relationship between user A and user B refers to the legal spouse relationship in reality, the entity relationship between user A and user C refers to the mother-child relationship in reality, and the entity relationship between user B and user C refers to the legal stepfather-stepson relationship in reality. Based on this, taking the path of the "mother-child relationship" between user A and user C as the first sub-path relationship, and the path of the "legal stepfather-stepson relationship" between user B and user C as the second sub-path relationship, the first sub-path relationship and the second sub-path relationship form the path relationship between user A and user B, which is used to represent the entity relationship between user A and user B.
[0142] 204. Expand the entity to be inferred according to the preset path relationship to obtain each result entity corresponding to the entity to be inferred.
[0143] To avoid the subsequent entity relationship subgraph from being too large, a preset expansion quantity threshold can be set during inference / expansion to limit the number of neighbor entities expanded per hop. Specifically, according to the preset expansion quantity threshold, the entity to be inferred can be expanded in multiple hops to obtain the expanded result entity, so as to facilitate the subsequent generation of the relevant entity relationship subgraph. In this way, the problem that the subsequently synthesized entity relationship subgraph is too large can be effectively avoided.
[0144] It should be noted that in order to make the inferred entity relationship subgraph more accurate, when expanding the entity to be inferred in each hop, the sub-path relationship between the entity to be inferred and each surrounding neighbor entity can be determined respectively, and the similarity between the sub-path relationship and the corresponding preset sub-path relationship in the preset entity relationship subgraph can be calculated. Thus, multiple target neighbor entities with higher similarity are selected from the multiple neighbor entities around the entity to be inferred; until the multi-hop expansion is completed, the result entity and the entity sequence from the entity to be inferred to the result entity are obtained, where the entity sequence can only include the entity to be inferred and the result entity; the entity sequence can also include the entity to be inferred, one or more intermediate entities, and the result entity, and the intermediate entity can be regarded as the result entity of the current hop expansion.
[0145] 205. Generate an entity relationship subgraph between the entity to be inferred and the result entity according to the entity to be inferred, the result entity, and the preset path relationship.
[0146] Among them, the entity relationship subgraph includes the entity relationship between the entity to be inferred and the corresponding result entity. In the embodiments of the present application, the entity relationship can be directly or indirectly represented by one or more path relationships between the entity to be inferred and the corresponding result entity; among them, the path relationship can be composed of one or more sub-path relationships.
[0147] 206. Extract the path relationship between the entity to be inferred and each result entity in the entity relationship subgraph through the second target submodel, and encode each path relationship to obtain the path relationship information corresponding to each path relationship.
[0148] Among them, the path relationship information can be composed of one or more sub-path relationship information from the entity to be inferred to the result entity. It should be noted that when there are multiple sub-path relationship information from the entity to be inferred to the result entity, the multiple sub-path relationship information is generated into path relationship sequence information according to the expansion time sequence of each expansion direction, and this relationship sequence information is used as the path relationship information.
[0149] In the embodiment of the present application, after obtaining the relevant entity relationship subgraph, the path relationship between the entity to be inferred and each result entity can be extracted to represent the information of each path relationship, so as to obtain the path relationship information corresponding to each path relationship. Among them, this information representation process can be specifically implemented by an encoding method, such as encoding the embedding information of each sub-path relationship through a Gate Recurrent Unit (GRU), so as to generate the path relationship information according to the encoding obtained for each sub-path relationship information.
[0150] It should be noted that if the preset path relationship corresponding to the preset entity relationship subgraph has not been information-represented, each preset path relationship of the preset entity pair can also be extracted from the preset entity relationship subgraph, and the preset path relationship information corresponding to each preset path relationship is encoded.
[0151] 207. Calculate the interaction coefficient between the path relationship information and the preset path relationship information, and determine the interaction coefficient as the matching score between the entity to be inferred and the corresponding result entity, obtaining multiple matching scores.
[0152] Specifically, the process of calculating the interaction coefficient between the path relationship information and the preset path relationship information is as follows: generate a path relationship information matrix according to the path relationship information and the preset path relationship information; process the path relationship information matrix through a preset attention mechanism to obtain a processed path similarity matrix; perform max pooling processing on the path similarity matrix to obtain a one-dimensional matrix; perform mapping processing on the one-dimensional matrix to obtain a mapped multi-dimensional matrix; perform a summation process on the parameters in the columns of the multi-dimensional matrix to obtain a multi-dimensional similarity matrix; perform a multi-layer perception process on the multi-dimensional similarity matrix to obtain the interaction coefficient between the path relationship information and the preset path relationship information.
[0153] Furthermore, the interaction coefficient can be determined as the matching score between the entity to be inferred and the corresponding result entity.
[0154] 208. Determine the target matching score with the highest score from the set of matching scores, and determine the target result entity corresponding to the entity to be inferred based on the target matching score.
[0155] It should be noted that the higher the matching score, the more likely the target result entity corresponding to the inference entity pair corresponding to the matching score is the correct entity answer. Therefore, in order to determine the target result entity that best matches the entity to be inferred, the embodiments of the present application can select the highest matching score from multiple matching scores as the target matching score. Thus, based on this target matching score, the target result entity corresponding to the entity to be inferred is determined. In this way, the accuracy in determining the result entity corresponding to the entity to be inferred is improved.
[0156] 209. Establish the target entity relationship between the entity to be inferred and the target result entity, and update the target entity relationship to the knowledge graph.
[0157] Specifically, establish the target entity relationship based on the entity to be inferred and the target result entity, and update the new knowledge entity pair composed of the entity to be inferred and the target result entity with the target entity relationship to the knowledge graph; or, the target entity relationship between the entity to be inferred and the target result entity in the knowledge graph can also be established. In this way, there is a target entity relationship between the entity to be inferred and the target result entity in the knowledge graph, realizing the update and completion of the knowledge graph.
[0158] To facilitate the understanding of the embodiments of the present application, the embodiments of the present application will be described with a specific application scenario example. Specifically, by performing the above steps 201-209, and in combination with Figure 4 , Figure 5 and Figure 6 , this application scenario example will be described, and this application scenario example can be implemented through a model. For example, as shown in Figure 4 , this application scenario example is that the target entity relationship of the support entity pair is "spouse" (support entity pair for the target relation “spouse”), and this support entity pair is the preset entity pair of the embodiments of the present application. The paths between "Elon Musk" and "Taluah" in the support entity pair include path 1 (path ) and path 2 (path ). Path 1 contains the path relationship information of "son of" and "fatherin law", and path 2 contains the path relationship information of "friend" and "acquaintance"; while the path between "Bill Gates" and "Melinda Gates" in the query entity pair (query entity pair) includes query path 1 (path ) and query path 2 (path ), where query path 1 contains relationship information of "son of", "father in law", and query path 2 contains relationship information of "friend", "acquaintance"; based on this, path interactions can be calculated according to the relationship information of the supporting entity pair and the query entity pair, then and the interaction coefficient between them is 1.0, and the interaction coefficient between them is 0.0, and the interaction coefficient between them is 0.0, and the interaction coefficient between them is 1.0. The specific example of this application scenario is as follows:
[0159] For example Figure 5 as shown, the application scenarios corresponding to this model are respectively the inference and matching scenarios, that is, this model can include two sub-models, and the joint implementation of these two sub-models can realize this application scenario example.
[0160] The inference scenario includes: (1) Extract the Support Subgraph, and this support subgraph is the preset entity relationship subgraph of the preset entity pair referred to in the embodiments of this application; specifically, for example Figure 5As shown, extract the path information between "ElonMusk" and "Taluah". The specific solution is to use two-side BFS, that is, first expand the paths in both directions, and then merge to obtain a support subgraph. (2) Reason about the query entity (the entity to be inferred referred to in the embodiments of the present application) based on the support subgraph; specifically, in order to ensure that the correct entity result can be in the inferred entity relationship subgraph (the entity relationship subgraph corresponding to the entity to be inferred), first, set the number of neighbor entities expanded per hop. Then, starting from the head entity (query entity), perform multi-hop expansion to obtain the expanded multi-hop query tail entity (result entity). Thus, generate the entity relationship subgraph corresponding to the query entity pair according to the support-relevant subgraph. Among them, after multi-hop expansion of the query entity, multiple intermediate entities obtained by expansion can be selected according to the similarity between the preset sub-path relationships of the sub-path relationships per hop. For example, encode each sub-path into sub-path relationship information (embedding) through the trained transE model, calculate the cosine similarity, and select a preset number of entities from the intermediate entities or result entities expanded per hop according to the calculated similarity. In this way, effectively prevent the entity relationship subgraph corresponding to the query entity pair from being too large.
[0161] The matching scenarios may include the following:
[0162] A. Represent the information of each path; specifically, use a gated recurrent unit (GRU) to encode the path relationship information of each path (such as embedding information).
[0163] B. Calculate the interaction coefficient of the path relationship information. Suppose 3 paths are extracted from the "support entity pair" and 4 paths are extracted from the "query entity pair". Then, calculate the similarity matrix for path to path, and a 3*4 path relationship information matrix can be obtained; for this matrix, the attention mechanism can be introduced to strengthen the influence of each path relationship on the matching degree, that is, process the 3*4 path relationship information matrix through the attention mechanism to obtain the similarity matrix; furthermore, first perform maxpooling on each row, then perform kernel (map from one dimension to 21 dimensions), sum on the columns, and finally pass through two layers of mlp to obtain a 1*1 score, which is the matching score for each query entity pair.
[0164] Among them, the similarity calculation formula can be as follows:
[0165]
[0166]
[0167] Among them, is the maximum value on the i-th row of the similarity matrix, and μ γ represents the mapping value, represents the similarity between the maximum value on the i-th row and the mapping value.
[0168] Among them, There are a total of 21 values, and each value is further log-transformed. Each row has 21 values. Finally, all rows are summed up column-wise (sum), and the resulting 21-dimensional value is the similarity φ between the query entity pair (hq, tq) and the support pair (hs, ts). This φ passes through two layers of MLP to obtain a 1*1 interaction coefficient. For the specific process of this part, please refer to Figure 6 .
[0169] C. If there are multiple support entity pairs, there will be multiple matching scores. At this time, the maximum score among all scores is taken as the interaction score between the "query entity pair" and the "support entity pair". Then, the tail entity in the "query entity pair" corresponding to this maximum interaction score is more likely to be the correct answer, that is, the target result entity.
[0170] In addition, before the application scenario of the model, the model needs to be trained. Specifically, the intersection of all nodes on the graph extended from the head entity and the candidate dataset is the candidate tail entity set of the head entity. The loss function of the model is margin loss. The positive example is the result (ground truth) that is originally correct in the supervised training, and the negative example is a certain node on the graph extended from the head node, and it is required that this node is not on the path from the head node to the correct tail node. If the correct tail node is at the 1st (2nd) hop, then the subsequent nodes extended from the correct tail node will not be used as negative examples either. For example, as Figure 5 shown, "William Henry Gates II" and "Paul Allen" in the reasoned query subgraph will not be used as negative examples because they are both on the path to the correct answer "Melinda Gates".
[0171] Specifically, the scoring formula of the model is as follows:
[0172]
[0173] Among them, g(h q , t q , S r ) represents the predicted score of querying the entity pair (hq, tq), and S r represents the set of supporting entity pairs corresponding to the entity relation relation(r). For example, if it is 5-shot, the number of supporting entity pairs included in the set is 5.
[0174] Specifically, the loss function formula during model training is as follows:
[0175]
[0176] Among them, (hq, tq) is a positive example, and (hq, tq-) is a negative example.
[0177] This application example, compared with the prior art, can be applied to many natural language processing tasks and products, such as question answering systems, knowledge reasoning, etc., and can improve the performance of these tasks and products in the case of small samples. At present, small-sample knowledge reasoning mainly relies on first-order neighbor information to represent relationship information, and such information is insufficient. This solution makes the prediction result more accurate by introducing path information and adopting more fine-grained knowledge.
[0178] As can be seen from the above, the embodiment of the present application can infer the entity relationship subgraph between the entity to be inferred and the result entity, and match the path relationship between the entity to be inferred and the result entity based on the entity relationship subgraph to obtain a set of matching scores. Furthermore, select the target result entity with the largest matching score with the entity to be inferred and update it to the knowledge graph; in this way, by combining knowledge entity mining of path relationships and entity relationships, the matching degree when mining new knowledge entities is improved, and a large number of existing knowledge entity samples are not required to ensure the accuracy of the knowledge graph.
[0179] To better implement the above method, the embodiment of the present application also provides a knowledge graph update device, and this knowledge graph update device can be integrated into a server.
[0180] For example, as Figure 7 shown, this knowledge graph update device may include an acquisition unit 401, an inference unit 402, a matching unit 403, a determination unit 404, and an update unit 405.
[0181] The acquisition unit 401 is used to acquire the entity to be inferred;
[0182] An inference unit 402 is configured to perform entity relationship inference on an entity to be inferred through a first target sub-model, so as to obtain an entity relationship sub-graph between the entity to be inferred and a result entity, where the first target sub-model is obtained by performing entity relationship training on a preset inference entity sample and a sample entity relationship sub-graph between the preset inference entity sample and a preset result entity sample;
[0183] A matching unit 403 is configured to perform path relationship matching on the entity relationship sub-graph through a second target sub-model to obtain a set of matching scores, where the set of matching scores includes the matching scores of the path relationship between each result entity and the entity to be inferred, and the second target sub-model is obtained by performing sample path relationship training on the sample entity relationship sub-graph and the sample matching scores of the sample path relationship between the preset inference entity sample and the preset result entity sample;
[0184] A determination unit 404 is configured to determine a target matching score with the largest score from the set of matching scores, and determine a target result entity corresponding to the entity to be inferred according to the target matching score;
[0185] An update unit 405 is configured to establish a target entity relationship between the entity to be inferred and the target result entity, and update the target entity relationship to the knowledge graph.
[0186] In some embodiments, the inference unit 402 is further configured to:
[0187] Extract a preset entity relationship sub-graph of each preset entity pair in the knowledge graph through the first target sub-model; based on the preset entity relationship sub-graph, perform inference on the entity to be inferred to obtain an entity relationship sub-graph between the entity to be inferred and the result entity.
[0188] In some embodiments, the inference unit 402 is further configured to:
[0189] Obtain the preset path relationship of the preset entity pair in the preset entity relationship sub-graph; expand the entity to be inferred according to the preset path relationship to obtain each result entity corresponding to the entity to be inferred; generate an entity relationship sub-graph between the entity to be inferred and the result entity according to the entity to be inferred, the result entity and the preset path relationship.
[0190] In some embodiments, the matching unit 403 is configured to:
[0191] Extract the path relationship between the entity to be inferred and each result entity in the entity relationship sub-graph through the second target sub-model, and encode each path relationship to obtain path relationship information corresponding to each path relationship; determine the matching score between the entity to be inferred and each result entity according to each path relationship information to obtain a set of matching scores.
[0192] In some embodiments, the matching unit 403 is further configured to:
[0193] Extract each preset path relationship of the preset entity pairs from the preset entity relationship sub-graph, and encode the preset path relationship information corresponding to each preset path relationship; calculate the interaction coefficient between the path relationship information and the preset path relationship information, and determine the interaction coefficient as the matching score between the entity to be inferred and the corresponding result entity.
[0194] In some embodiments, the matching unit 403 is further configured to:
[0195] Generate a path relationship information matrix according to the path relationship information and the preset path relationship information; process the path relationship information matrix through a preset attention mechanism to obtain a processed path similarity matrix; perform max pooling processing on the path similarity matrix to obtain a one-dimensional matrix; perform mapping processing on the one-dimensional matrix to obtain a mapped multi-dimensional matrix; perform summation processing on the parameters in the columns of the multi-dimensional matrix to obtain a multi-dimensional similarity matrix; perform multi-layer perception processing on the multi-dimensional similarity matrix to obtain the interaction coefficient between the path relationship information and the preset path relationship information.
[0196] In some embodiments, the knowledge graph updating device further includes a training unit, configured to:
[0197] Obtain the preset inference entities and preset result entities in the knowledge graph, and set the sample matching score between the preset inference entities and the preset result entities; input the preset inference entities into a preset model to obtain the predicted matching score between the preset inference entities and the preset result entities; obtain the matching difference value between the predicted matching score and the sample matching score; perform iterative training on the preset model based on the matching difference value until the matching difference value converges to obtain a trained target model, where the target model includes a first target sub-model and a second target sub-model.
[0198] As can be seen from the above, the embodiment of the present application can obtain the entity to be inferred through the obtaining unit 401; the inference unit 402 is used to perform entity relationship inference on the entity to be inferred through the first target sub-model to obtain an entity relationship sub-graph between the entity to be inferred and the result entity, where the first target sub-model is obtained by performing entity relationship training on the preset inference entity samples and the sample entity relationship sub-graph between the preset inference entity samples and the preset result entity samples; the matching unit 403 is used to perform path relationship matching on the entity relationship sub-graph through the second target sub-model to obtain a set of matching scores, and the set of matching scores includes the matching scores of the path relationship between each result entity and the entity to be inferred, where the second target sub-model is obtained by performing sample path relationship training on the sample entity relationship sub-graph and the sample matching scores of the sample path relationship between the preset inference entity samples and the preset result entity samples; the determining unit 404 is used to determine the target matching score with the largest score from the set of matching scores, and determine the target result entity corresponding to the entity to be inferred according to the target matching score; the updating unit 405 is used to establish the target entity relationship between the entity to be inferred and the target result entity, and update the target entity relationship to the knowledge graph. Thus, the present solution can infer the entity relationship sub-graph between the entity to be inferred and the result entity, and match the path relationship between the entity to be inferred and the result entity based on the entity relationship sub-graph to obtain a set of matching scores. Furthermore, the target result entity with the largest matching score with the entity to be inferred is selected and updated to the knowledge graph; in this way, by combining the path relationship and the entity relationship for knowledge entity mining, the matching degree when mining new knowledge entities is improved, without a large number of existing knowledge entity samples, ensuring the accuracy of the knowledge graph.
[0199] The embodiment of the present application also provides a computer device, as Figure 8 shown, which shows the structural schematic diagram of the computer device involved in the embodiment of the present application. Specifically:
[0200] The computer device may include a processor 501 with one or more processing cores, a memory 502 with one or more computer-readable storage media, a power supply 503, an input unit 504 and other components. Those skilled in the art can understand that Figure 8 the structural schematic diagram of the computer device shown in
[0201] The processor 501 is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 502, and by invoking the data stored in the memory 502, it executes various functions of the computer device and processes data. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 501 either.
[0202] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and knowledge graph updates by running the software programs and modules stored in the memory 502. The memory 502 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store data created according to the use of the computer device. In addition, the memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 502 may also include a memory controller to provide the processor 501 with access to the memory 502.
[0203] The computer device also includes a power supply 503 that powers each component. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 503 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0204] The computer device may also include an input unit 504, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0205] Although not shown, the computer device may also include a display unit, etc., which will not be elaborated here. Specifically, in the embodiment of this application, the processor 501 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 502 according to the following instructions, and the processor 501 will run the application programs stored in the memory 502 to achieve various functions as follows:
[0206] Obtain the entity to be inferred; perform entity relationship reasoning on the entity to be inferred through the first target sub-model to obtain an entity relationship sub-graph between the entity to be inferred and the result entity, where the first target sub-model is obtained by performing entity relationship training on the preset inference entity samples and the sample entity relationship sub-graph between the preset inference entity samples and the preset result entity samples; perform path relationship matching on the entity relationship sub-graph through the second target sub-model to obtain a set of matching scores, where the set of matching scores includes the matching scores of the path relationship between each result entity and the entity to be inferred, and the second target sub-model is obtained by performing sample path relationship training on the sample entity relationship sub-graph and the sample matching scores of the sample path relationship between the preset inference entity samples and the preset result entity samples; determine the target matching score with the largest score from the set of matching scores, and determine the target result entity corresponding to the entity to be inferred according to the target matching score; establish a target entity relationship between the entity to be inferred and the target result entity, and update the target entity relationship to the knowledge graph.
[0207] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated here.
[0208] Thus, it can be seen that this solution can infer the entity relationship sub-graph between the entity to be inferred and the result entity, and match the path relationship between the entity to be inferred and the result entity based on the entity relationship sub-graph to obtain a set of matching scores. Furthermore, select the target result entity with the largest matching score for the entity to be inferred and update it to the knowledge graph; in this way, combine the path relationship and the entity relationship for knowledge entity mining, improve the matching degree when mining new knowledge entities, and do not require a large number of existing knowledge entity samples to ensure the accuracy of the knowledge graph.
[0209] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling related hardware through instructions. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0210] Therefore, an embodiment of the present application provides a computer-readable storage medium, which stores multiple instructions that can be loaded by a processor to execute the steps in any of the knowledge graph update methods provided by the embodiments of the present application. For example, the instructions can execute the following steps:
[0211] Obtain the entity to be inferred; perform entity relationship inference on the entity to be inferred through the first target sub-model to obtain an entity relationship sub-graph between the entity to be inferred and the result entity, where the first target sub-model is obtained by performing entity relationship training on a preset inference entity sample and a sample entity relationship sub-graph between the preset inference entity sample and the preset result entity sample; perform path relationship matching on the entity relationship sub-graph through the second target sub-model to obtain a set of matching scores, where the set of matching scores includes the matching scores of the path relationship between each result entity and the entity to be inferred, and the second target sub-model is obtained by performing sample path relationship training on the sample entity relationship sub-graph and the sample matching scores of the sample path relationship between the preset inference entity sample and the preset result entity sample; determine the target matching score with the largest score from the set of matching scores, and determine the target result entity corresponding to the entity to be inferred according to the target matching score; establish a target entity relationship between the entity to be inferred and the target result entity, and update the target entity relationship to the knowledge graph.
[0212] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated herein.
[0213] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0214] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the knowledge graph update methods provided in the embodiments of the present application, the beneficial effects that can be achieved by any of the knowledge graph update methods provided in the embodiments of the present application can be realized. For details, refer to the previous embodiments and will not be elaborated herein.
[0215] The above has introduced in detail a knowledge graph update method, device, equipment, and computer-readable storage medium provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for updating a knowledge graph, characterized in that, Including: Obtain the entity to be inferred; Perform entity relationship reasoning on the entity to be inferred through the first target sub-model to obtain an entity relationship sub-graph between the entity to be inferred and the result entity, where the first target sub-model is obtained by entity relationship training using a preset inference entity sample and a sample entity relationship sub-graph between the preset inference entity sample and a preset result entity sample; Perform path relationship matching on the entity relationship sub-graph through the second target sub-model to obtain a set of matching scores, where the set of matching scores contains the matching scores of the path relationship between each result entity and the entity to be inferred, and the second target sub-model is obtained by sample path relationship training using the sample entity relationship sub-graph and the sample matching scores of the sample path relationship between the preset inference entity sample and the preset result entity sample; Determine the target matching score with the largest score from the set of matching scores, and determine the target result entity corresponding to the entity to be inferred according to the target matching score; Establish a target entity relationship between the entity to be inferred and the target result entity, and update the target entity relationship to the knowledge graph.
2. The method according to claim 1, characterized in that The step of performing entity relationship reasoning on the entity to be inferred through the first target sub-model to obtain an entity relationship sub-graph between the entity to be inferred and the result entity includes: Extract the preset entity relationship sub-graph of each preset entity pair in the knowledge graph through the first target sub-model; Based on the preset entity relationship sub-graph, perform reasoning on the entity to be inferred to obtain an entity relationship sub-graph between the entity to be inferred and the result entity.
3. The method according to claim 2, wherein The step of performing reasoning on the entity to be inferred based on the preset entity relationship sub-graph to obtain an entity relationship sub-graph between the entity to be inferred and the result entity includes: Obtain the preset path relationship of the preset entity pair in the preset entity relationship sub-graph; Expand the entity to be inferred according to the preset path relationship to obtain each result entity corresponding to the entity to be inferred; Generate an entity relationship sub-graph between the entity to be inferred and the result entity according to the entity to be inferred, the result entity and the preset path relationship.
4. The method according to claim 2, wherein The step of performing path relationship matching on the entity relationship sub-graph through the second target sub-model to obtain a set of matching scores includes: Extract the path relationship between the entity to be inferred and each result entity in the entity relationship sub-graph through the second target sub-model, and encode each path relationship to obtain path relationship information corresponding to each path relationship; Determine the matching score between the entity to be inferred and each result entity according to each path relationship information to obtain the set of matching scores.
5. The method according to claim 4, wherein The step of determining the matching score between the entity to be inferred and each result entity according to each path relationship information includes: Extract each preset path relationship of the preset entity pair from the preset entity relationship sub-graph and encode the preset path relationship information corresponding to each preset path relationship; Calculate the interaction coefficient between the path relationship information and the preset path relationship information, and determine the interaction coefficient as the matching score between the entity to be inferred and the corresponding result entity.
6. The method according to claim 5, characterized in that, The calculating the interaction coefficient between the path relationship information and the preset path relationship information includes: Generate a path relationship information matrix according to the path relationship information and the preset path relationship information; Process the path relationship information matrix through a preset attention mechanism to obtain a processed path similarity matrix; Perform max pooling processing on the path similarity matrix to obtain a one-dimensional matrix; Perform mapping processing on the one-dimensional matrix to obtain a mapped multi-dimensional matrix; Sum the parameters in the multi-dimensional matrix column by column to obtain a multi-dimensional similarity matrix; Perform multi-layer perception processing on the multi-dimensional similarity matrix to obtain the interaction coefficient between the path relationship information and the preset path relationship information.
7. The method according to claim 1, characterized in that, Before obtaining the entity relationship subgraph between the entity to be inferred and the result entity by performing entity relationship inference on the entity to be inferred through the first target sub-model, it further includes: Obtain the preset inference entity and the preset result entity in the knowledge graph, and set the sample matching score between the preset inference entity and the preset result entity; Input the preset inference entity into the preset model to obtain the predicted matching score between the preset inference entity and the preset result entity; Obtain the matching difference value between the predicted matching score and the sample matching score; Iteratively train the preset model based on the matching difference value until the matching difference value converges to obtain the trained target model, where the target model includes the first target sub-model and the second target sub-model.
8. A knowledge graph update device, characterized in that, It includes: An acquisition unit, configured to acquire the entity to be inferred; An inference unit, configured to perform entity relationship inference on the entity to be inferred through the first target sub-model to obtain the entity relationship subgraph between the entity to be inferred and the result entity, where the first target sub-model is obtained by performing entity relationship training on the preset inference entity sample and the sample entity relationship subgraph between the preset inference entity sample and the preset result entity sample; A matching unit, configured to perform path relationship matching on the entity relationship subgraph through the second target sub-model to obtain a set of matching scores, where the set of matching scores includes the matching scores of the path relationships between each result entity and the entity to be inferred, and the second target sub-model is obtained by performing sample path relationship training on the sample entity relationship subgraph and the sample matching scores of the sample path relationships between the preset inference entity sample and the preset result entity sample; A determination unit, configured to determine the target matching score with the largest score from the set of matching scores, and determine the target result entity corresponding to the entity to be inferred according to the target matching score; An update unit, configured to establish the target entity relationship between the entity to be inferred and the target result entity, and update the target entity relationship to the knowledge graph.
9. A computer device, characterized in that, It includes a processor and a memory. The memory stores a computer program, and the processor is configured to run the computer program in the memory to implement the steps in the knowledge graph update method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the knowledge graph update method according to any one of claims 1 to 7.
11. A computer program product, characterized in that, It includes computer instructions that execute the steps in the knowledge graph update method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Rule and path combined knowledge graph representation learning method
CN110069638A
Multi-source heterogeneous knowledge graph collaborative reasoning method and device based on entity alignment
CN112818137A