A method, device and medium for completing a knowledge graph

By extracting missing triplets and their associated images from the multimodal knowledge graph library, generating enhanced text and using multimodal large models for inference, the problem of multimodal information being ignored in the existing technology is solved, and the accurate completion and accuracy of the knowledge graph are achieved.

CN120123435BActive Publication Date: 2025-08-01ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510595172.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-01
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The existing knowledge graph completion method ignores multimodal information such as unstructured text and images associated with entity, resulting in incomplete semantic expression and low accuracy.

Method used

Missing triples and their associated images are extracted from the multimodal knowledge graph library, enhanced text is generated, and pre-constructed multimodal big model and inference prompt word engineering are used for inference, and the image and text information are completed.

Benefits of technology

It improves the accuracy and completeness of knowledge graph completion, provides a more reliable reflection of knowledge structure, and is suitable for practical applications in designated fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123435B_ABST
    Figure CN120123435B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, and medium for completing a knowledge graph. The method includes: extracting missing triples with missing fields and associated images of the missing triples from a multi-modal knowledge graph library in a specified domain; generating enhanced text for the missing triples, where the enhanced text is semantic enhanced text for the known fields in the missing triples; based on a pre-constructed inference prompt word engineering for the missing fields, through a pre-constructed multi-modal large model, inferring the missing triples, associated images, and enhanced text to obtain completed triples after completing the missing fields. Thus, making full use of the multi-modal information of the triples, fusing and complementing the image information and text information, accurately completing the missing triples, making the knowledge graph information in the specified domain more complete and accurate, better reflecting the knowledge structure and relationships in the specified domain, and providing a reliable basis for subsequent practical application scenarios of the knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of knowledge graph technology, and in particular to a method, device and medium for completing a knowledge graph. Background Art

[0002] Knowledge Graph Completion (KGC) is a technology that uses existing knowledge and structures in the knowledge graph to infer and generate new knowledge through a series of algorithms and models, in order to continuously improve and expand the knowledge graph. It can be widely used in search engines, recommendation systems and other fields.

[0003] Currently, knowledge graph completion primarily relies on structured triples within the knowledge graph, modeling entity relationships through rule-based reasoning (e.g., path sorting algorithms) or embedding models (e.g., TransE and RotatE models) to complete the knowledge graph. This approach relies solely on structured triples and ignores multimodal information such as unstructured text and images associated with entities. This results in incomplete semantic representation and low knowledge graph completion accuracy.

[0004] Therefore, how to improve the completion accuracy of knowledge graphs is an urgent problem to be solved by technical personnel in this field. Summary of the Invention

[0005] In view of this, one aspect of the present application provides a method for completing a knowledge graph, the method comprising:

[0006] Extract missing triples with missing fields and associated images of the missing triples from a multimodal knowledge graph library in a specified field;

[0007] Generating an enhanced text for the missing triple; the enhanced text is a semantically enhanced text for the known fields in the missing triple;

[0008] Obtaining a pre-built inference hint word project about the missing field and a pre-built multimodal large model;

[0009] Based on the inference prompt word project, the missing triples, the associated images and the enhanced text are inferred through the multimodal large model to obtain the completed triples after the missing fields are completed.

[0010] Optionally, the known fields include known entity fields and known relationship fields, and the known entity fields are head entity fields or tail entity fields;

[0011] The associated image is an image related to the known entity field; the enhanced text includes entity enhanced text and relationship enhanced text.

[0012] Optionally, generating the enhanced text for the missing triple includes:

[0013] Extracting the structured description text and unstructured description text of the known entity field from the domain knowledge base of the specified domain;

[0014] Fusing the structured description text and the unstructured description text to obtain the entity enhanced text;

[0015] Filtering out the complete triples containing the known relationship field from the multimodal knowledge graph library;

[0016] Based on the domain knowledge base, converting the complete triple into semantic description text to obtain the relationship enhanced text.

[0017] Optionally, based on the inference prompt engineering, through the multimodal large model, reasoning about the missing triple, the associated image, and the enhanced text to obtain the completed triple after filling in the missing field, including:

[0018] Extracting visual features from the associated image;

[0019] Based on the cross-modal attention mechanism, spatially semantically aligning the visual features with the enhanced text to generate a cross-modal joint embedding representation;

[0020] According to the enhanced text and the cross-modal joint embedding representation, based on the inference prompt engineering, reasoning about the missing field of the missing triple to obtain the completed triple.

[0021] Optionally, the missing field is the head entity field or the tail entity field; the output constraint conditions in the inference prompt engineering include:

[0022] Determining the type constraint conditions for constraining the type of the missing field according to the enhanced text and the cross-modal joint embedding representation;

[0023] According to the type constraint conditions, generating the completed field of the missing triple and the inference process explanation text of the completed field to obtain the completed triple.

[0024] Optionally, the method for completing the knowledge graph further includes:

[0025] Extracting a specified number of complete triples from the multimodal knowledge graph library;

[0026] Masking and hiding the specified field in the complete triple to obtain a masked triple; the specified field is the head entity field or the tail entity field;

[0027] Generate a completion field for the masked triple through the described knowledge graph completion method to obtain a masked completion triple;

[0028] Fine-tune and train the multi-modal large model according to the complete triple and the masked completion triple.

[0029] Optionally, the fine-tuning and training of the multi-modal large model according to the complete triple and the masked completion triple includes:

[0030] Obtain the actual type of the specified field and the inference type of the completion field;

[0031] Determine the entity cross-entropy loss between the specified field and the completion field, and the type cross-entropy loss between the actual type and the inference type;

[0032] Determine the joint cross-entropy loss according to the entity cross-entropy loss and the type cross-entropy loss;

[0033] According to the joint cross-entropy loss, use the gradient backpropagation algorithm to iteratively update and train the model parameters of the multi-modal large model until the preset iteration condition is reached.

[0034] Another aspect of the present application provides a knowledge graph completion device, and the device includes:

[0035] Missing triple extraction module, used to extract missing triples with missing fields and the associated images of the missing triples from the multi-modal knowledge graph library in a specified domain;

[0036] Enhanced text generation module, used to generate enhanced text for the missing triples; the enhanced text is semantic enhanced text for the known fields in the missing triples;

[0037] Model acquisition module, used to acquire a pre-constructed inference prompt word project for the missing field and a pre-constructed multi-modal large model;

[0038] Triple completion module, used to infer the missing triples, the associated images and the enhanced text through the multi-modal large model based on the inference prompt word project to obtain a completion triple after completing the missing field.

[0039] Another aspect of the present application provides a knowledge graph completion device, including a memory and a processor. A computer program that can run on the processor is stored on the memory, and when the processor executes the program, it implements the steps of the knowledge graph completion method.

[0040] Another aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method for completing the knowledge graph are implemented.

[0041] The beneficial effects of the method, device and medium for completing the knowledge graph provided by the present application are as follows: combining missing triples, associated images and enhanced text, and using a multimodal large model for reasoning, making full use of the multimodal information of the triples, integrating and complementing the image information and text information, improving the understanding and reasoning accuracy of the missing triples, so as to accurately complete the missing triples, making the knowledge graph information in the specified field more complete and accurate, better reflecting the knowledge structure and relationship in the specified field, and providing a reliable basis for the actual application scenarios of the subsequent knowledge graph. Description of the Drawings

[0042] Figure 1 It is a schematic flowchart of a method for completing a knowledge graph provided by an embodiment of the present application;

[0043] Figure 2 It is a schematic diagram of the principle of a method for completing a knowledge graph provided by an embodiment of the present application;

[0044] Figure 3 It is a schematic diagram of the principle of a method for completing a knowledge graph provided by another embodiment of the present application;

[0045] Figure 4 It is a schematic diagram of the principle of a method for completing a knowledge graph provided by still another embodiment of the present application;

[0046] Figure 5 It is a schematic structural diagram of a device for completing a knowledge graph provided by an embodiment of the present application;

[0047] Figure 6 It is a schematic structural diagram of a device for completing a knowledge graph provided by another embodiment of the present application.

[0048] The reference numerals are as follows: 50 is a missing triple extraction module, 51 is an enhanced text generation module, 52 is a model acquisition module, 53 is a triple completion module, 60 is a memory, 61 is a processor, 62 is a display screen, 63 is an input / output interface, 64 is a communication interface, 65 is a power supply, 66 is a communication bus, 601 is a computer program, 602 is an operating system, and 603 is data. Detailed Embodiments

[0049] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "said", and "the" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0050] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0051] Figure 1 The flowchart of a method for completing a knowledge graph provided by an embodiment of this application is shown as Figure 1 shown, and the method includes:

[0052] S10: Extract missing triples with missing fields and associated images of the missing triples from a multi-modal knowledge graph library in a specified domain;

[0053] Figure 2 The schematic diagram of the principle of a method for completing a knowledge graph provided by an embodiment of this application is shown as Figure 2 shown. In a specific embodiment, a multi-modal knowledge graph library in a specified domain is obtained. The specified domain includes, but is not limited to, the geology field, the medical and health field, and the sports field. And the obtained multi-modal knowledge graph library may include images, texts, audios, videos, etc. related to the specified domain, and this application does not make any limitations in this regard.

[0054] In order to improve the reliability and integrity of the data information in the multi-modal knowledge graph library, incomplete triples in the knowledge graph are completed. Specifically, missing triples with missing fields are extracted from the multi-modal knowledge graph library. Among them, it is worth noting that the triple structure in the knowledge graph is (head entity field, relationship field, tail entity field). Correspondingly, the missing fields in the missing triples can be any one or any two of the head entity field, the relationship field, and the tail entity field, and this application does not make any limitations in this regard.

[0055] However, it can be understood that in a specific embodiment, if there are two missing fields and only one known field remains, the amount of completed triple data after completing the missing fields is very large, and it is difficult to determine the triples that match the actual situation from them. Therefore, achieving a satisfactory effect requires consuming a large amount of computing resources. Thus, in an alternative embodiment, the missing field can be any one of the head entity field, the relationship field, and the tail entity field.

[0056] Furthermore, in order to accurately complete the missing fields, multi-dimensional data information can be obtained from a multi-modal knowledge graph library. Specifically, in an alternative embodiment, as Figure 2 shown, the associated image of the missing triple can be extracted.

[0057] S11: Generate enhanced text for the missing triple; the enhanced text is the semantic enhanced text for the known fields in the missing triple;

[0058] To more accurately complete the missing information in the missing triple and make full use of the known information of the missing triple, that is, to make full use of the known fields, so as to provide effective information for completing the missing fields. Specifically, generate enhanced text for the known fields in the missing triple. The enhanced text refers to text that contains more information, richer semantics, or a more explicit structure by processing, expanding, or supplementing the known fields. In fact, the enhanced text can also be understood as a more detailed and complete explanatory text for the known fields.

[0059] For ease of understanding, an example will be given below. For example, if the known field is the head entity field and the head entity field is "apple", the generated enhanced text can be "Apple is a fruit rich in cellulose, vitamin C, and various antioxidants."

[0060] S12: Obtain a pre-constructed inference prompt word project for the missing fields and a pre-constructed multi-modal large model;

[0061] S13: Based on the inference prompt word project, through the multi-modal large model, perform inference on the missing triple, the associated image, and the enhanced text to obtain the completed triple after completing the missing fields.

[0062] Furthermore, as Figure 2 shown, based on the pre-constructed inference Prompt project for the missing fields, input the extracted missing triple and associated image, as well as the generated enhanced text, into the pre-constructed multi-modal large model (Multimodal Large Models, MLLMs) for inference, so as to obtain the completed triple after completing the missing fields.

[0063] In a specific embodiment, the inference Prompt engineering can guide the multi-modal large model to more accurately infer the missing fields in the missing triples based on the associated images and associated image pairs. It can be understood that the multi-modal large model is a model that can understand and process multiple input forms. In the embodiments of the present application, multi-modal refers to the enhanced text and associated images as inputs. The multi-modal large model enhances the accuracy of filling in the missing fields by fusing the information of the enhanced text and the associated images.

[0064] Therefore, the method for completing the knowledge graph provided in the embodiments of the present application combines the missing triples, associated images, and enhanced text, and uses a multi-modal large model for inference, making full use of the multi-modal information of the triples, fusing and complementing the image information and text information, enhancing the understanding and inference accuracy of the missing triples, thereby accurately completing the missing triples, making the knowledge graph information in the specified domain more complete and accurate, better reflecting the knowledge structure and relationships in the specified domain, and providing a reliable basis for the actual application scenarios of the subsequent knowledge graph.

[0065] In an alternative embodiment, the known fields include known entity fields and known relationship fields, and the known entity field is the head entity field or the tail entity field;

[0066] The associated image is an image about the known entity field; the enhanced text includes entity-enhanced text and relationship-enhanced text.

[0067] It can be understood that in a triple including a head entity field, a relationship field, and a tail entity field, the relationship field is crucial for the accuracy of completing the triple. In addition, it can be understood that when the missing field in the missing triple is the relationship field, or any two fields, the computational complexity for completing the missing field is very large and the accuracy is not high, which may lead to the accuracy of the completed knowledge graph being difficult to meet the user's expectations.

[0068] Therefore, in an alternative embodiment, to ensure the accuracy of the finally completed knowledge graph and improve the completion efficiency, the known fields in the missing triples include known entity fields and known relationship fields, and the known entity field is one of the head entity field and the tail entity field, that is, the missing field is the head entity field or the tail entity field.

[0069] Accordingly, when obtaining the associated image of the missing triple, an image about the known entity field is obtained, and the enhanced text includes entity-enhanced text for enhancing the known entity and relationship-enhanced text for enhancing the known relationship field.

[0070] Therefore, in a specific embodiment, when completing a missing triple through a multimodal large model, enhanced text corresponding to the known entity field and the known relationship field is generated. Further, multimodal information fusion and reasoning are performed based on the associated image of the known entity and the generated enhanced text, so as to obtain an accurate completed triple.

[0071] Figure 3 The principle schematic diagram of a method for completing a knowledge graph provided by another embodiment of the present application. On the basis of the above embodiment, as an optional embodiment, generating enhanced text for the missing triple includes:

[0072] Extract the structured description text and unstructured description text of the known entity field from the domain knowledge base of the specified domain;

[0073] Fuse the structured description text and the unstructured description text to obtain entity enhanced text;

[0074] Screen out the complete triples containing the known relationship field from the multimodal knowledge graph library;

[0075] Based on the domain knowledge base, convert the complete triple into semantic description text to obtain relationship enhanced text.

[0076] As Figure 3 shown, in a specific embodiment, when the missing field of the missing triple is the head entity field or the tail entity field, generating enhanced text for the missing triple includes two parts, namely, the entity enhancement part and the relationship enhancement part.

[0077] In an optional embodiment, as Figure 3 shown, when enhancing the known entity field, based on the known domain knowledge base of the specified domain, extract the structured description text and unstructured description text of the known entity field. Among them, the known domain knowledge base includes but is not limited to open-source domain databases, domain literature libraries, and third-party knowledge sources, etc., and the present application does not limit the domain knowledge base.

[0078] The structured description text refers to text with clear information. For example, it may include but is not limited to text such as the type, function classification, and associated scenarios of the known entity. The unstructured description text refers to free text that requires further language understanding. For example, free text such as the function description and application examples of the known entity.

[0079] In an optional embodiment, the structured description text can be extracted from the domain knowledge base through an interface or data crawling technology. In another optional embodiment, the unstructured description text can be extracted from the domain knowledge base through a semantic parsing large model.

[0080] Further, the extracted structured description text and unstructured description text are fused to obtain an entity-enhanced text, which is used to complement the detailed information lost during the process of converting the knowledge graph structure into text.

[0081] In addition, as Figure 3 shown, in the embodiment of relationship enhancement, first, from the multimodal knowledge graph library, complete triples that are the same as the relationship field in the missing triple are screened out, and the mapping relationship between the entity field and the relationship field of the complete triple is constructed. That is, a complete triple refers to a triple in which the head entity field, the relationship field, and the tail entity field all exist, and the relationship field is the same as the known relationship field of the missing triple. For example, the missing triple is (Apple, belongs to, *), where "*" is used to represent the missing field, and the complete triple is (Durian, belongs to, Tropical fruit).

[0082] Further, as Figure 3 shown, each complete triple is converted into a natural language pattern text in a unified format, thereby obtaining a relationship-enhanced text. When converting a complete triple, it is converted based on the known domain knowledge base in a specified domain. In an optional embodiment, it can be converted based on a preset template description. For example, the template description is "Relationship R represents the [association pattern] between the head entity H and the tail entity T", where the association pattern is filled based on the domain knowledge base. For example, for the complete triple (Gene, regulates, Disease), the template description is "Gene H affects the progression of Disease T by regulating [biological pathway]". As Figure 3 shown, in a specific embodiment, after separately enhancing the known entity field and the known relationship field, an enhanced text including the entity-enhanced text and the relationship-enhanced text is obtained. This enhanced text can be used for subsequent reasoning about the missing fields in the missing triple.

[0083] Thus, the method for complementing a knowledge graph provided by the embodiments of the present application, based on the joint modeling of the entity-enhanced text and the relationship-enhanced text, uses the domain knowledge base to supplement structured and domain knowledge, improving the complementing efficiency and domain generalization.

[0084] Figure 4 This is a schematic diagram of the principle of a method for complementing a knowledge graph provided by another embodiment of the present application. In an optional embodiment, based on the inference prompt engineering, through a multimodal large model, inferences are made on the missing triple, the associated image, and the enhanced text to obtain a completed triple after complementing the missing fields, including:

[0085] Extract visual features from the associated image;

[0086] Based on the cross-modal attention mechanism, align the visual features and the enhanced text spatially and semantically to generate a cross-modal joint embedding representation;

[0087] Based on the enhanced text and cross-modal joint embedding representation, and based on the inference prompt engineering, infer the missing fields of the missing triples to obtain the completed triples.

[0088] As Figure 4 shown, in a specific embodiment where the multi-modal large model infers missing triples, associated images, and enhanced text, the multi-modal large model includes a multi-modal fusion unit. In an alternative embodiment, the multi-modal fusion unit performs visual feature encoding on the associated images input to the large model through an image encoder to obtain the visual features of the associated images. Further, based on the cross-modal attention mechanism of the multi-modal large model, the visual features are spatially semantically aligned with the enhanced text, thereby achieving cross-modal joint embedding representation.

[0089] Among them, it should be noted that the cross-modal joint embedding representation refers to a vector that represents information from both the associated images and the enhanced text in a multi-dimensional space, enabling data information from the two modalities to be compared or operated within the same semantic space. Thus, based on the end-to-end architecture of the multi-modal large model, the automatic fusion and inference of visual features and text features are realized.

[0090] Further, based on the enhanced text and cross-modal joint embedding representation, guided by the inference Prompt engineering, the multi-modal large model infers the missing fields in the missing triples to obtain the completed triples.

[0091] Therefore, the method for completing the knowledge graph provided by the embodiments of the present application aligns the image features and text descriptions spatially and semantically through the cross-modal attention mechanism, breaks through the limitations of a single modality, fully excavates the visual and semantic information of entity associations, and significantly improves the integrity and accuracy of knowledge graph completion.

[0092] In an alternative embodiment, the missing field is the head entity field or the tail entity field; the output constraint conditions in the inference prompt engineering include:

[0093] Based on the enhanced text and cross-modal joint embedding representation, determine the type constraint conditions for constraining the type of the missing field;

[0094] According to the type constraint conditions, generate the completed fields of the missing triples and the inference process explanation text for the completed fields to obtain the completed triples.

[0095] In a specific embodiment, in order to improve the ability of the multi-modal large model to capture implicit associations in triples, such as Figure 4As shown, in the pre-constructed inference Prompt engineering, the output constraints are used for inference output in a chain of thought. Specifically, first, based on the enhanced text and cross-modal joint embedding representation, the type constraints for the missing fields are generated. These type constraints are used to constrain the types of the missing fields, that is, to constrain the types of the completed fields in the generated completed triples.

[0096] That is to say, in a specific embodiment, according to the enhanced text and cross-modal joint embedding representation, it can be determined what type the finally generated completed field belongs to, and thus the corresponding type constraints are dynamically generated according to different enhanced texts and cross-modal joint embedding representations.

[0097] Furthermore, based on the type constraints, the completed field with the highest inference generation probability, that is, the most likely or most conforming to the enhanced text and cross-modal joint embedding representation, and the explanatory text of the inference process of the completed field are generated to obtain the completed triple. Among them, the explanatory text of the inference process refers to the explanatory text of the entire derivation process of generating the completed field, so that users can view and verify the rationality and accuracy of the derivation results and processes.

[0098] Thus, the method for completing the knowledge graph provided by the embodiments of the present application introduces a chain of thought reasoning mechanism, decomposes the completion task into two steps: type constraint generation and entity reasoning, and combines logical explanation generation to improve the model's ability to capture implicit associations.

[0099] To further improve the accuracy of the knowledge graph completion of the present application, in an optional embodiment, the method for completing the knowledge graph further includes:

[0100] Extract a specified number of complete triples from the multi-modal knowledge graph library;

[0101] Mask and hide the specified fields in the complete triples to obtain masked triples; the specified fields are the head entity fields or the tail entity fields;

[0102] Generate the completed fields of the masked triples through the method for completing the knowledge graph provided by any of the above embodiments to obtain masked completed triples;

[0103] Fine-tune and train the multi-modal large model according to the complete triples and the masked completed triples.

[0104] In a specific embodiment, the model parameters of the multi-modal large model are continuously optimized until a high-precision model is obtained, so that the subsequent multi-modal large model can quickly complete the knowledge graph in a specified domain. Specifically, such as Figure 4As shown, a specified number of complete triples are extracted from the multi-modal knowledge graph library in a specified domain, that is, from the multi-modal knowledge graph library, triples with known head entity fields, relationship fields, and tail entity fields in a certain number are extracted. In a specific embodiment, the extracted complete triples are used as a training data set to fine-tune and train the multi-modal large model.

[0105] Specifically, the specified fields in the complete triples are masked and hidden by the model optimization unit, where the specified fields are head entity fields or tail entity fields. Thus, the complete triples are converted into a missing triple (i.e., a masked triple) identical to any of the above embodiments.

[0106] To verify whether the accuracy of the current multi-modal large model meets the requirements, the masked triples are completed by the knowledge graph completion method provided by any of the above embodiments to obtain masked completed triples. Thus, the complete triples and the corresponding masked completed triples form a training data set to fine-tune and train the multi-modal large model through the training data set.

[0107] It can be understood that the generated masked completed triples are inferred triples, while the complete triples are real actual triples. By analyzing the differences between a large number of inferred triples and actual triples, the multi-modal large model can be fine-tuned and trained.

[0108] Based on the above embodiments, as an optional embodiment, fine-tuning and training the multi-modal large model according to the complete triples and the masked completed triples includes:

[0109] Obtain the actual type of the specified field and the inferred type of the completed field;

[0110] Determine the entity cross-entropy loss between the specified field and the completed field, and the type cross-entropy loss between the actual type and the inferred type;

[0111] Determine the joint cross-entropy loss according to the entity cross-entropy loss and the type cross-entropy loss;

[0112] According to the joint cross-entropy loss, through the gradient backpropagation algorithm, the model parameters of the multi-modal large model are iteratively updated and trained until the preset iteration condition is reached.

[0113] It can be understood that in a specific embodiment, when inferring and completing the missing triples, the missing fields can be inferred, that is, the types of the specified fields masked in the above embodiments. Therefore, in an optional embodiment, to improve the model optimization accuracy, the differences between the complete triples and the masked completed triples can be compared from multiple dimensions.

[0114] Specifically, obtain the actual type of the specified field (i.e., the entity field before masking), and obtain the inferred type of the completed field in the inferred masked completed triple. Further, calculate the type cross-entropy loss between the actual type and the inferred type.

[0115] In addition, starting from the field itself, calculate the entity cross-entropy loss between the specified field (i.e., the entity field before masking) and the inferred completed field. Further, based on the type cross-entropy loss and the entity cross-entropy loss, the Figure 4 joint cross-entropy loss shown can be calculated.

[0116] In an optional embodiment, using the complete triple as the training data set, continuously perform inference through the knowledge graph completion method provided by any of the above embodiments to obtain the masked completed triple. Further, continuously calculate the joint cross-entropy loss between the complete triple and the masked completed triple, and based on the joint cross-entropy loss, use the gradient backpropagation algorithm to iteratively update the model parameters of the multi-modal large model until the preset iteration condition is reached.

[0117] Among them, the preset iteration condition can be the preset number of iterations, and the preset number of iterations can be set according to actual business requirements, which is not limited in this application. The preset iteration condition can also be that the joint cross-entropy loss converges to a preset value.

[0118] That is to say, under the complete triple training data set, with the goal of minimizing the joint cross-entropy loss, continuously iteratively update the model parameters of the multi-modal large model to obtain a high-performance multi-modal large model for subsequent use in actual missing triple completion services.

[0119] It should be noted that for the process of updating the model parameters of the multi-modal large model, the model parameters can be continuously updated in actual business to adapt to knowledge graph completion tasks in different specified fields. It can also be that before the knowledge graph completion task, the model is first trained to reach the expected level and then the missing triples are completed, which is not limited in this application.

[0120] To make the technical solutions provided in this application clearer to those skilled in the art, the following takes the FB15k-237 multi-modal knowledge graph library as the knowledge graph to be completed, and the missing triple is (Zhang San, nationality, *) as an example, where "*" represents the missing field (i.e., the tail entity field), and the knowledge graph completion method provided in this application is described in detail.

[0121] In the FB15k-237 multimodal knowledge graph library, there are 1,454 entities, 237 relations, and 310,116 training triples. In this multimodal knowledge graph library, the structure of each triple is (head entity H, relation R, tail entity T).

[0122] In a specific embodiment, the known entity is the head entity Zhang San (this head entity is a person's name). From the known domain knowledge base, the structured description text and unstructured description text of the known entity Zhang San are extracted. At the same time, the associated image of the known entity Zhang San is extracted from the multimodal knowledge graph library.

[0123] Among them, for example, the structured description text is: "Place of birth: Beijing", "Occupation: Lawyer", "Spouse: Li Si". The unstructured description text is: "Zhang San, is a lawyer and writer, born in Beijing on January 17, 1964, graduated from University A, and his spouse is Li Si. Li Si is a teacher and also graduated from University A...".

[0124] Furthermore, in an alternative embodiment, the key information fragments in the above-mentioned structured description text and unstructured description text are extracted through the tokenizer of the large language model, so as to fuse the above-mentioned structured description text and unstructured description text to generate an entity-enhanced text with a certain length. This entity-enhanced text supplements the missing details in the knowledge graph and provides semantic constraints for subsequent reasoning.

[0125] In addition, in a specific embodiment, the complete triples with the same relation field "nationality" are extracted from the multimodal knowledge graph library. For example, the extracted complete triple is (Wang Wu, nationality, China). Further, each complete triple is transformed into a semantic description text, so as to obtain a relation-enhanced text. For example, the transformation is based on a preset template description, and the template description is: "The relation field 'nationality' represents the relationship between the head entity [Wang Wu] and the tail entity [China]". Thus, a batch of relation-enhanced texts are generated for input into the multimodal large model to explicitly strengthen the learning of the relation pattern.

[0126] Furthermore, the associated image of the known entity Zhang San is feature-encoded through a pre-trained multimodal large model to obtain the visual features of the associated image. And the visual features and the above-mentioned obtained enhanced texts (including entity-enhanced texts and relation-enhanced texts) are spatially semantically aligned through a cross-modal attention mechanism, so as to generate a cross-modal joint embedding representation.

[0127] Thus, under the reasoning Prompt engineering of chain-of-thought, according to the enhanced text and cross-modal joint embedding representation, the type constraint conditions for the missing fields are inferred and generated. For example, in the above example, the type constraint conditions for the missing field (i.e., the tail entity field) are: "country, and related to the birthplace, citizenship, or long-term residence of the head entity". Further, according to the type constraint conditions, the field with the highest probability is inferred and generated, that is, the field that is most likely or most conforms to the enhanced text and cross-modal joint embedding representation, as well as the reasoning process explanation text for completing the field. In the above example, it can be inferred that the completed field for the tail entity field is "China", and thus, the completed triple is obtained.

[0128] In an alternative embodiment, to enable the multi-modal large model to better understand the dry features of the associated images in the domain knowledge base and be familiar with the knowledge graph completion task, and further improve the performance of the multi-modal large model in knowledge graph completion, it is necessary to perform supervised fine-tuning optimization on the model.

[0129] Specifically, using a specified number of complete triples in the FB15k-237 multi-modal knowledge graph library as the supervision signal, the head entity field or the tail entity field is masked and hidden, and the masked triple is processed in the manner of the above embodiment. The unmasked entity field and the unmasked relationship field in the masked triple are semantically enhanced to obtain the enhanced text, which is then input into the multi-modal model to obtain the masked completed triple.

[0130] Further, a multi-task loss function is designed. First, calculate the type cross-entropy loss between the actual type "country" of the masked field and the type of the completed field inferred by the multi-modal large model. At the same time, calculate the entity cross-entropy loss between the masked field "China" and the completed field.

[0131] Thus, according to the entity cross-entropy loss and the type cross-entropy loss, the joint cross-entropy loss is determined, and through the gradient backpropagation algorithm, the model parameters of the multi-modal large model are iteratively updated and trained until the preset iteration condition is reached.

[0132] In a specific embodiment, after fine-tuning, the performance of the multi-modal large model on the FB15k-237 test set reaches 59.1%, which is 13% higher than the baseline model (e.g., the TransE model), and the consistency between the reasoning process explanation text of the completed field and the manual annotation reaches 91%.

[0133] Thus, by integrating multi-modal data with a step-by-step reasoning mechanism, problems such as incomplete semantic expression, difficulty in reasoning complex relationships, and poor cross-domain generalization when completing a knowledge graph based on structured triples are avoided, enabling accurate completion of missing triples and enhancing the interpretability of multi-modal large models. In addition, through the collaborative design of a dual-driven enhancement framework based on entity-enhanced text and relationship-enhanced text and a chain reasoning engine for multi-modal large models, full-process optimization from knowledge enhancement to reasoning completion is achieved.

[0134] In the above embodiments, the method for completing a knowledge graph is described in detail. The present application also provides an embodiment corresponding to a device for completing a knowledge graph.

[0135] Figure 5 The following is a schematic structural diagram of a device for completing a knowledge graph provided by an embodiment of the present application, as Figure 5 shown. The device includes:

[0136] A missing triple extraction module 50, configured to extract missing triples with missing fields and associated images of the missing triples from a multi-modal knowledge graph library in a specified domain;

[0137] An enhanced text generation module 51, configured to generate enhanced text for the missing triples; the enhanced text is semantic enhanced text for the known fields in the missing triples;

[0138] A model acquisition module 52, configured to acquire a pre-constructed inference prompt word project for the missing fields and a pre-constructed multi-modal large model;

[0139] A triple completion module 53, configured to perform reasoning on the missing triples, associated images, and enhanced text through the multi-modal large model based on the inference prompt word project to obtain completed triples with the missing fields completed.

[0140] In addition, the device for completing a knowledge graph provided by an embodiment of the present application further includes:

[0141] A description text extraction module, configured to extract structured description text and unstructured description text of known entity fields from a domain knowledge base in a specified domain;

[0142] A fusion module, configured to fuse the structured description text and the unstructured description text to obtain entity-enhanced text;

[0143] A screening module, configured to screen out complete triples containing known relationship fields from the multi-modal knowledge graph library;

[0144] A conversion module, configured to convert the complete triples into semantic description text based on the domain knowledge base to obtain relationship-enhanced text.

[0145] A visual feature extraction module for extracting visual features from associated images;

[0146] A semantic alignment module for spatially semantically aligning visual features with enhanced text based on a cross-modal attention mechanism to generate a cross-modal joint embedding representation;

[0147] An inference module for inferring the missing fields of a missing triple based on inference prompt engineering according to the enhanced text and the cross-modal joint embedding representation to obtain a completed triple.

[0148] A constraint condition determination module for determining type constraint conditions for constraining the types of missing fields according to the enhanced text and the cross-modal joint embedding representation;

[0149] A completed field generation module for generating the completed fields of a missing triple and the inference process explanation text of the completed fields according to the type constraint conditions to obtain a completed triple.

[0150] A complete triple extraction module for extracting a specified number of complete triples from a multi-modal knowledge graph library;

[0151] A masking module for masking and hiding specified fields in a complete triple to obtain a masked triple; the specified fields are the head entity field or the tail entity field;

[0152] A masked completed triple generation module for generating the completed fields of a masked triple through a knowledge graph completion method to obtain a masked completed triple;

[0153] A fine-tuning training module for fine-tuning and training a multi-modal large model according to the complete triple and the masked completed triple.

[0154] A field type acquisition module for acquiring the actual type of a specified field and the inference type of the completed field;

[0155] A cross-entropy loss determination module for determining the entity cross-entropy loss between the specified field and the completed field, and the type cross-entropy loss between the actual type and the inference type;

[0156] A joint cross-entropy loss determination module for determining the joint cross-entropy loss according to the entity cross-entropy loss and the type cross-entropy loss;

[0157] An iteration module for iteratively updating and training the model parameters of a multi-modal large model according to the joint cross-entropy loss through the gradient backpropagation algorithm until a preset iteration condition is reached.

[0158] Figure 6 The structural schematic diagram of a knowledge graph completion device provided by another embodiment of the present application, as Figure 6As shown in the figure, the knowledge graph completion device includes: a memory 60 for storing computer programs;

[0159] a processor 61 for implementing the steps of the knowledge graph completion method as mentioned in the above embodiments when executing the computer program.

[0160] The knowledge graph completion device provided in this embodiment may include, but is not limited to, a tablet computer, a notebook computer, a desktop computer, etc.

[0161] Among them, the processor 61 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 61 may be implemented in at least one hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 61 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 61 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 61 may further include an artificial intelligence (AI) processor for processing computational operations related to machine learning.

[0162] The memory 60 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 60 may further include a high-speed random access memory and a non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 60 is at least used to store the following computer program 601. After the computer program is loaded and executed by the processor 61, it can implement the relevant steps of the knowledge graph completion method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 60 may further include an operating system 602 and data 603, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 602 may include Windows, Unix, Linux, etc. The data 603 may include, but is not limited to, the relevant data involved in the knowledge graph completion method.

[0163] In some embodiments, the knowledge graph completion device may further include a display screen 62, an input / output interface 63, a communication interface 64, a power supply 65, and a communication bus 66.

[0164] Those skilled in the art can understand that Figure 6 the structure shown in does not constitute a limitation on the knowledge graph completion device, and it may include more or fewer components than shown in the figure.

[0165] The knowledge graph completion device provided by the embodiments of the present application includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the knowledge graph completion method in the above embodiments.

[0166] It should be noted that although the operations are depicted in a specific order in the drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all of the illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of the various system modules and components in the above embodiments should not be construed as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

Claims

1. A method for completing a knowledge graph, characterized in that, The method includes: Extracting missing triples with missing fields and the associated images of the missing triples from a multi-modal knowledge graph library in a specified domain; Generating enhanced text for the missing triples; the enhanced text is semantic enhanced text for the known fields in the missing triples; Obtaining a pre-constructed inference prompt word project for the missing fields and a pre-constructed multi-modal large model; Based on the inference prompt word project, through the multi-modal large model, inferring the missing triples, the associated images and the enhanced text to obtain a completed triple after completing the missing fields; The inferring the missing triples, the associated images and the enhanced text through the multi-modal large model based on the inference prompt word project to obtain a completed triple after completing the missing fields includes: Extracting visual features from the associated images; Based on the cross-modal attention mechanism, spatially semantically aligning the visual features with the enhanced text to generate a cross-modal joint embedding representation; According to the enhanced text and the cross-modal joint embedding representation, based on the inference prompt word project, inferring the missing fields of the missing triples to obtain the completed triples; The missing fields are head entity fields or tail entity fields; the output constraint conditions in the inference prompt word project include: Determining type constraint conditions for constraining the types of the missing fields according to the enhanced text and the cross-modal joint embedding representation; According to the type constraint conditions, generating the completed fields of the missing triples and the inference process explanation text of the completed fields to obtain the completed triples.

2. The method for completing a knowledge graph according to claim 1, wherein The known fields include known entity fields and known relationship fields; The associated images are images about the known entity fields; the enhanced text includes entity enhanced text and relationship enhanced text.

3. The method for completing the knowledge graph according to claim 2, wherein, The generating the enhanced text for the missing triples includes: Extracting the structured description text and unstructured description text of the known entity fields from the domain knowledge library in the specified domain; Fusing the structured description text and the unstructured description text to obtain the entity enhanced text; Filtering out complete triples containing the known relationship fields from the multi-modal knowledge graph library; Based on the domain knowledge library, transforming the complete triples into semantic description text to obtain the relationship enhanced text.

4. The method for completing a knowledge graph according to claim 1, characterized in that The method further includes: Extracting a specified number of complete triples from the multi-modal knowledge graph library; Masking and hiding the specified fields in the complete triples to obtain masked triples; the specified fields are head entity fields or tail entity fields; Generating the completed fields of the masked triples through the knowledge graph completion method according to any one of claims 1 to 3 to obtain masked completed triples; Fine-tuning and training the multi-modal large model according to the complete triples and the masked completed triples.

5. The method for completing a knowledge graph according to claim 4, wherein The fine-tuning and training the multi-modal large model according to the complete triples and the masked completed triples includes: Obtain the actual type of the specified field and the inference type of the complemented field; Determine the entity cross-entropy loss between the specified field and the complemented field, and the type cross-entropy loss between the actual type and the inference type; Determine the joint cross-entropy loss based on the entity cross-entropy loss and the type cross-entropy loss; According to the joint cross-entropy loss, use the gradient backpropagation algorithm to iteratively update and train the model parameters of the multimodal large model until the preset iteration condition is reached.

6. An apparatus for completing a knowledge graph, characterized in that The device includes: A missing triple extraction module, configured to extract missing triples with missing fields and the associated images of the missing triples from a multimodal knowledge graph library in a specified domain; An enhanced text generation module, configured to generate enhanced text for the missing triples; the enhanced text is semantic enhanced text for the known fields in the missing triples; A model acquisition module, configured to obtain a pre-constructed inference prompt word project for the missing fields and a pre-constructed multimodal large model; A triple complementation module, configured to, based on the inference prompt word project, use the multimodal large model to reason about the missing triples, the associated images, and the enhanced text, and obtain complemented triples after complementing the missing fields; A visual feature extraction module, configured to extract visual features from the associated images; A semantic alignment module, configured to, based on a cross-modal attention mechanism, perform spatial semantic alignment between the visual features and the enhanced text to generate a cross-modal joint embedding representation; An inference module, configured to, based on the enhanced text and the cross-modal joint embedding representation, and based on the inference prompt word project, infer the missing fields of the missing triples to obtain the complemented triples; A constraint condition determination module, configured to determine type constraint conditions for constraining the types of the missing fields according to the enhanced text and the cross-modal joint embedding representation; A complemented field generation module, configured to generate the complemented fields of the missing triples and the inference process explanation text of the complemented fields according to the type constraint conditions to obtain the complemented triples.

7. A knowledge graph completion device, comprising a memory and a processor, where a computer program that can run on the processor is stored on the memory, and is characterized in that When the processor executes the program, it implements the steps of the knowledge graph complementation method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the knowledge graph complementation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Knowledge graph completion method and system based on modal decoupling and integrated reasoning

    CN117573890A

  • Large language model knowledge graph completion method based on subgraph structure information enhancement

    CN118674026A