Knowledge graph construction method, storage medium and electronic equipment
By extracting the content information of the pictures in the document and converting it into a plain text document, the problem that the existing technology cannot extract picture information is solved, the integrity and accuracy of the knowledge graph are achieved, and the accurate mapping of entity relationships and pictures is supported.
Patent Information
- Application Number
- CN202510343490.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-20
AI Technical Summary
Existing knowledge graph construction methods cannot extract image information in documents, resulting in the completeness and accuracy of knowledge graphs being affected.
By determining the pictures and their location information in the document, the content information of the picture is extracted and converted into a plain text document, and then the entity relationship extraction of the plain text document is carried out to build a knowledge graph.
Effectively retain key information in the picture, ensure the integrity and accuracy of the knowledge graph, and realize the accurate mapping of entity relationships and pictures, supporting the correct reference of entity relationships to the picture.
Smart Images

Figure CN120181202A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of knowledge graphs, and particularly to a method for constructing a knowledge graph, a storage medium, and an electronic device. Background Art
[0002] A knowledge graph is a structured semantic network used to describe entities, concepts, and their relationships in the real world. It organizes information in the form of a graph, where nodes represent entities or concepts, and edges represent the relationships between them.
[0003] In related technologies, usually, a document is input into a large language model for automatic extraction of a knowledge graph. However, the knowledge graph construction framework based on a large language model can only process pure text data and cannot perform entity relationship extraction on the pictures in the original document, resulting in the loss of information contained in the pictures in the extracted knowledge graph, thereby affecting the integrity and accuracy of the knowledge graph. Summary of the Invention
[0004] This application provides a method for constructing a knowledge graph, a storage medium, and an electronic device to solve the problem that the current knowledge graph construction method cannot extract picture information in a document.
[0005] To solve the above problems, this application adopts the following technical solutions:
[0006] In a first aspect, an embodiment of this application provides a method for constructing a knowledge graph, and the method includes:
[0007] Determine the pictures in the initial document and the position information of the pictures;
[0008] Extract the content information of the pictures, and based on the content information and the position information, convert the initial document into a pure text document;
[0009] Perform entity relationship extraction on the pure text document to obtain the entity relationship extraction result of the pure text document;
[0010] Construct a knowledge graph based on the entity relationship extraction result.
[0011] In an embodiment of this application, extracting the content information of the pictures includes:
[0012] Input the pictures into a multimodal large language model to obtain the content information corresponding to the pictures output by the multimodal large language model.
[0013] In an embodiment of this application, based on the content information and the position information, converting the initial document into a pure text document includes:
[0014] Replace the picture with the corresponding content information and position information of the picture to obtain the pure text document.
[0015] In one embodiment of the present application, entity relationship extraction is performed on the pure text document to obtain the entity relationship extraction result of the pure text document, including:
[0016] Divide the pure text document into multiple text blocks;
[0017] Input the multiple text blocks into a large language model to obtain the entity relationship extraction results corresponding to each of the multiple text blocks output by the large language model;
[0018] Based on the entity relationship extraction results corresponding to each of the multiple text blocks, obtain the entity relationship extraction result of the pure text document.
[0019] In one embodiment of the present application, dividing the pure text document into multiple text blocks includes:
[0020] For any picture, determine the associated paragraph of the picture; the associated paragraph includes the previous paragraph and / or the subsequent paragraph;
[0021] Divide the associated paragraph and the content information corresponding to the picture into the text blocks.
[0022] In one embodiment of the present application, inputting the multiple text blocks into a large language model to obtain the entity relationship extraction results corresponding to each of the multiple text blocks output by the large language model includes:
[0023] Input the source recognition prompt word and the multiple text blocks into the large language model, so that the large language model determines whether the entity relationship extraction results corresponding to each of the multiple text blocks originate from a picture in response to the source recognition prompt word, and stores the position information of the picture as the position index of the entity relationship extraction result when it is determined that the entity relationship extraction result of any text block originates from a picture.
[0024] In one embodiment of the present application, the entity relationship extraction result includes entity names and / or relationships between entities; based on the entity relationship extraction results corresponding to each of the multiple text blocks, obtaining the entity relationship extraction result of the pure text document includes:
[0025] Perform an alignment operation on the entity names in the multiple text blocks to obtain the aligned entity names; and / or, perform a merging operation on the relationships between entities in the multiple text blocks to obtain the merged relationships between entities;
[0026] Based on the aligned entity names and / or the merged relationships between entities, obtain the entity relationship extraction result of the pure text document.
[0027] In one embodiment of the present application, an alignment operation is performed on the entity names in the multiple text blocks to obtain the aligned entity names, including:
[0028] Determine the string similarity, and / or semantic similarity, and / or context similarity between the multiple entity names;
[0029] Based on the string similarity, and / or the semantic similarity, and / or the context similarity, determine the comprehensive similarity;
[0030] Based on the comprehensive similarity, perform an alignment operation on the multiple entity names to obtain the aligned entity names.
[0031] In a second aspect, based on the same inventive concept, an embodiment of the present application provides a knowledge graph construction device, which includes:
[0032] A picture determination module, configured to determine the pictures in the initial document and the position information of the pictures;
[0033] A document conversion module, configured to extract the content information of the pictures, and based on the content information and the position information, convert the initial document into a plain text document;
[0034] An information extraction module, configured to perform entity relationship extraction on the plain text document to obtain the entity relationship extraction result of the plain text document;
[0035] A graph construction module, configured to construct a knowledge graph based on the entity relationship extraction result.
[0036] In one embodiment of the present application, the document conversion module includes:
[0037] A picture content recognition sub-module, configured to input the picture into a multi-modal large language model to obtain the content information corresponding to the picture output by the multi-modal large language model.
[0038] In one embodiment of the present application, the document conversion module includes:
[0039] An information replacement sub-module, configured to replace the picture with the content information and position information corresponding to the picture to obtain the plain text document.
[0040] In one embodiment of the present application, the information extraction module includes:
[0041] A document chunking sub-module, configured to divide the plain text document into multiple text blocks;
[0042] An information extraction sub-module, configured to input the multiple text blocks into a large language model to obtain the entity relationship extraction results corresponding to the multiple text blocks output by the large language model;
[0043] An information fusion sub-module, configured to obtain the entity relationship extraction result of the pure text document based on the entity relationship extraction results corresponding to the multiple text blocks.
[0044] In an embodiment of the present application, the document chunking sub-module includes:
[0045] An associated paragraph determination unit, configured to determine the associated paragraphs of the picture; the associated paragraphs include the previous paragraphs and / or the subsequent paragraphs;
[0046] A text block division unit, configured to divide the associated paragraphs and the content information corresponding to the picture into the text blocks.
[0047] In an embodiment of the present application, the information extraction sub-module includes:
[0048] A prompt unit, configured to input a source identification prompt word and the multiple text blocks into the large language model, so that the large language model determines whether the entity relationship extraction results corresponding to the multiple text blocks are from pictures in response to the source identification prompt word, and stores the position information of the picture as the position index of the entity relationship extraction result when it is determined that the entity relationship extraction result of any text block is from a picture.
[0049] In an embodiment of the present application, the entity relationship extraction result includes entity names and / or relationships between entities; the information fusion sub-module includes:
[0050] An alignment unit, configured to perform an alignment operation on the entity names in the multiple text blocks to obtain the aligned entity names; and / or, perform a merging operation on the relationships between entities in the multiple text blocks to obtain the merged relationships between entities;
[0051] A fusion unit, configured to obtain the entity relationship extraction result of the pure text document based on the aligned entity names and / or the merged relationships between entities.
[0052] In an embodiment of the present application, the alignment unit includes:
[0053] A first similarity determination sub-unit, configured to determine the string similarity, and / or semantic similarity, and / or context similarity between the multiple entity names;
[0054] A second similarity determination sub-unit, configured to determine the comprehensive similarity based on the string similarity, and / or the semantic similarity, and / or the context similarity;
[0055] An entity name alignment subunit, configured to perform an alignment operation on the multiple entity names based on the comprehensive similarity to obtain aligned entity names.
[0056] In a third aspect, based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, on which an executable program is stored, and when the executable program is executed by a processor, the knowledge graph construction method proposed in the first aspect of the present application is implemented.
[0057] In a fourth aspect, based on the same inventive concept, an embodiment of the present application provides an electronic device, including:
[0058] A memory, configured to store an executable program;
[0059] A processor;
[0060] When the executable program is executed by the processor, the knowledge graph construction method proposed in the first aspect of the present application is implemented.
[0061] Compared with the prior art, the present application has the following advantages:
[0062] A knowledge graph construction method provided by an embodiment of the present application first determines pictures in an initial document and the position information of the pictures; then extracts the content information of the pictures, and based on the content information and the position information, converts the initial document into a pure text document; finally, performs entity relationship extraction on the pure text document to obtain an entity relationship extraction result of the pure text document, and constructs a knowledge graph based on the entity relationship extraction result. By extracting the content information of the pictures in the initial document, the embodiment of the present application can convert the initial document with mixed pictures and texts into a pure text document, thereby retaining the key information in the pictures and ensuring the integrity and accuracy of the knowledge graph; at the same time, by attaching the position information of the pictures to the entity relationships, an accurate mapping between the entity relationships and the pictures can be established, thereby realizing the correct reference of the entity relationships to the pictures. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0064] Figure 1 is a flowchart of the steps of a knowledge graph construction method in an embodiment of the present application;
[0065] Figure 2It is a schematic diagram of the functional modules of a knowledge graph construction device in an embodiment of the present application;
[0066] Figure 3 It is a schematic diagram of the structure of an electronic device in an embodiment of the present application. Detailed implementation manners
[0067] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings in the embodiments of the present application. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully conveyed to those skilled in the art.
[0068] It should be noted that the pictures in the document usually contain key entities (such as people, objects, scenes) and the relationships between them (such as spatial relationships, action relationships). For example, in a medical literature, the text may describe the symptoms of a certain disease, while the picture may show the morphological characteristics of the diseased tissue. If the picture information is ignored, the knowledge graph will not be able to comprehensively reflect the content of the document, thus affecting subsequent applications; in the field of education, the pictures and texts in textbooks jointly convey knowledge, and ignoring the pictures will cause the knowledge graph to be unable to support comprehensive teaching assistance functions.
[0069] In the related art, the knowledge graph construction framework based on large language models can only process pure text data and cannot process the pictures in the original document, that is, it cannot extract entities and relationships from the pictures. At the same time, in some application scenarios, such as using the knowledge graph for question answering, there is a need to reference pictures. However, since the traditional knowledge graph construction method cannot analyze the pictures in the original document, it is difficult to establish the mapping relationship between the entity relationship and the original document, and thus it is difficult to correctly reference the pictures.
[0070] Aiming at the problem that the current knowledge graph construction method cannot extract the picture information in the document, the present application aims to provide a knowledge graph construction method. By extracting the content information of the pictures in the initial document, the initial document mixed with pictures and texts can be converted into a pure text document, thereby retaining the key information in the pictures and ensuring the integrity and accuracy of the knowledge graph; at the same time, by attaching the position information of the pictures to the entity relationship, the accurate mapping between the entity relationship and the pictures can be established, and thus the correct reference of the entity relationship to the pictures can be realized.
[0071] Referring to Figure 1 , a knowledge graph construction method of the present application is shown, and the method may include the following steps:
[0072] S101: Determine the pictures in the initial document and the position information of the pictures.
[0073] It should be noted that the initial document is a document containing at least one picture, and this embodiment does not limit the format of the initial document and the number of pictures.
[0074] In this embodiment, for initial documents in different formats (such as PDF, web pages, scanned documents, etc.), corresponding document recognition tools (such as PDFMiner, Apache Tika) can be used to extract the text structure information of the initial document, including the coordinate information of elements such as text paragraphs and image blocks. Furthermore, based on the extracted text structure information, each picture in the initial document and the position information of each picture can be determined.
[0075] In a specific implementation, an open-source document recognition tool can be used to design a document traverser. By traversing the initial document, the document traverser can sequentially identify the pictures in the initial document and record the positions where the pictures are located. Specifically, the document traverser will identify the current traversed content to determine whether the current traversed content is a picture; if it is determined that the current traversed content is a picture, the picture will be extracted and the position identifier will be incremented by 1; if it is determined that the current traversed content is not a picture, the current position will be skipped and the position identifier will be incremented by 1. In this way, after the traversal is completed, each extracted picture has a corresponding position index, which is used to reflect the position information of the picture in the initial document.
[0076] In this embodiment, the layout structure of the initial document can also be recognized through computer vision technology, and then the pictures and the position information of the pictures can be extracted. Specifically, the initial document can be input into a preset layout segmentation model to obtain the pictures in the initial document and the position information of the pictures output by the layout segmentation model.
[0077] S102: Extract the content information of the pictures, and based on the content information and the position information, convert the initial document into a plain text document.
[0078] In this embodiment, the content information of the pictures can specifically include information such as the entity type, relationship between entities, and entity attributes of the pictures.
[0079] In a specific implementation, an object detection model can be used to identify entity types such as objects and people in each picture, and output the bounding boxes and class labels. For example, if a vehicle is included in the picture, the vehicle and the position coordinates representing the bounding box of the vehicle will be output; a vision-language model can also be used to generate picture descriptions of each picture for parsing the relationship between entities. For example, it is output that the vehicle is located in the leftmost lane of the road; OCR (Optical Character Recognition) technology can also be used to extract the text information in each picture, such as vehicle model, license plate, etc. information, and then based on the text information, the entity attributes can be extracted.
[0080] In this embodiment, after extracting the content information of each picture, for any picture, the picture can be replaced with the corresponding content information and position information of the picture to obtain a pure text document.
[0081] In this embodiment, by replacing the picture with the corresponding content information and position information, the integrity of the context information can be effectively retained, avoiding fragmentation or information loss in the document content; at the same time, by writing the position information of the picture into the pure text document, it is convenient to establish an accurate mapping between the subsequent entity relationship and the picture.
[0082] S103: Perform entity relationship extraction on the pure text document to obtain the entity relationship extraction result of the pure text document.
[0083] In this embodiment, before performing entity relationship extraction on the pure text document, preprocessing such as text cleaning, sentence splitting and word segmentation, and part-of-speech tagging can be performed on the pure text document. Among them, text cleaning is used to preprocess the input pure text document to remove noise information in the document, such as extra spaces, special symbols, headers and footers, footnotes, etc., to improve the efficiency and accuracy of subsequent processing; sentence splitting and word segmentation is used to split the cleaned text according to the sentence structure to form independent sentence units, and then perform word segmentation on each sentence to split the sentence into individual word units; part-of-speech tagging is used to perform part-of-speech tagging on the segmented words to clarify the grammatical attributes of each word, such as nouns, verbs, adjectives, etc. Part-of-speech tagging helps to understand the role of words in sentences and provides a grammatical basis for subsequent entity recognition and relationship extraction.
[0084] In this embodiment, the entity relationship extraction process specifically includes an entity recognition stage and a relationship recognition stage; among them, the entity recognition stage is used to recognize entities in the pure text document to obtain the entity extraction result, and the relationship recognition stage is used to recognize the relationships between entities in the pure text document to obtain the relationship extraction result. That is to say, the entity relationship extraction result includes the entity extraction result and the relationship extraction result.
[0085] In this embodiment, during the process of performing entity relationship extraction on the pure text document, the position information of the picture will also be stored as an attribute in the entity relationship, thereby establishing a mapping between the picture and the entity relationship. Among them, the entity relationship specifically includes entities and / or relationships between entities.
[0086] S104: Construct a knowledge graph based on the entity relationship extraction result.
[0087] In this embodiment, after obtaining the entity relationship extraction result, operations such as entity alignment and relationship merging can be performed on the entity relationship extraction result based on rules. The extraction results of scattered text blocks can be fused into a globally consistent knowledge representation. Finally, a knowledge graph is constructed using graph construction software such as NetworkX (a Python-based library for graph theory and network analysis, used to create, manipulate, and study complex network structures), that is, the extracted entities and relationships are represented as a graphical structure, and data is stored in the form of nodes (entities) and edges (relationships) to obtain the final knowledge graph.
[0088] In this embodiment, by converting the initial document with pictures into a plain text document, the construction of a knowledge graph for documents with pictures can be realized, effectively avoiding the loss of picture information, and thus ensuring the integrity and accuracy of the knowledge graph; at the same time, in the entity relationship extraction task, by establishing a mapping between pictures and entity relationships, the correct reference of pictures and entity relationships can be realized.
[0089] In a feasible embodiment, the step of extracting the content information of the picture in S102 may specifically include the following sub-steps:
[0090] S102-1: Input the picture into a multimodal large language model to obtain the content information corresponding to the picture output by the multimodal large language model.
[0091] It should be noted that multimodal large language models (MLLMs) are a class of models that combine the natural language processing capabilities of large language models (LLMs) with the ability to understand and generate various modal data (such as vision, audio, video, etc.).
[0092] It should be noted that the core mechanism for multimodal large language models to understand image content stems from the semantic association between modalities achieved through cross-modal representation alignment and multimodal fusion mechanisms.
[0093] In this embodiment, the multimodal large language model can use Vision Transformer (abbreviated as ViT) or Convolutional Neural Network (abbreviated as CNN) to perform block processing on images, generating low-dimensional dense vectors (Patch Embeddings). Through the self-attention mechanism, local image features (such as edges and textures) and global context (such as the spatial relationship of objects) are extracted to form hierarchical visual features. This process maps the pixel space to a latent space aligned with text semantics.
[0094] In this embodiment, during the model training stage, a contrastive learning framework (such as CLIP, Contrastive Language-Image Pre-training) can be adopted to pre-train the model on a large-scale image-text pair dataset. By maximizing the cosine similarity of paired image-text pairs (Contrastive Loss), while minimizing the similarity of unpaired samples, the construction of an image-text joint embedding space is achieved. This stage forces the model to learn the geometric alignment of visual concepts (such as "dog") and language symbols (the word "dog") in the vector space.
[0095] In this embodiment, the multimodal large language model uses the cross-attention mechanism to dynamically establish the association between visual features and language tokens (the basic units after the decomposition of language text) during the decoding stage. For example, when generating an image description, each autoregressive step of the language decoder will selectively focus on the regions in the image feature map that are relevant to the current semantics (such as focusing on the sky region when generating "airplane") through Query-Key-Value operations. This process realizes fine-grained semantic matching between modalities.
[0096] In this embodiment, the multimodal large language model can also learn the symbol grounding ability through massive data, that is, establishing a statistical association between abstract semantic symbols (such as "joy" in the text) and visual features (such as the texture pattern of a smiling face). With the generalization ability of Transformer, the model can combine existing concepts to understand new scenarios (such as an animal breed never recognized before, by decomposing and recognizing it as a combination of features such as "fur + feline + stripes").
[0097] In this embodiment, the multimodal large language model has multiple layers of abstract representations internally. The bottom layer processes perceptual features such as color / shape, the middle layer identifies object parts and spatial relationships, and the top layer infers scene intentions and social common sense (for example, judging a "wedding scene" requires integrating elements such as white veils, crowds, and decorations). Through the visualization of attention weights, it can be observed that the model dynamically adjusts the feature attention granularity in different tasks.
[0098] In this embodiment, through large-scale multimodal pre-training data (such as LAION-5B), the parameter sharing mechanism of the unified Transformer architecture, and the explicit constraint of the self-supervised learning objective on cross-modal associations, the multi-modal data source processing ability of the multimodal large language model can be achieved. In addition, the interpretability and reasoning ability of the model can be further improved through differentiable symbolic logic injection and causal intervention.
[0099] In this embodiment, compared with using multiple models to separately identify information such as entity types, entity descriptions, and entity attributes in pictures, leveraging the processing ability of the multimodal large language model for cross-modal data such as pictures can extract the content information of pictures more efficiently and accurately.
[0100] In a feasible embodiment, S103 may specifically include the following sub-steps:
[0101] S103-1: Divide the pure text document into multiple text blocks.
[0102] In this embodiment, considering that most language models (such as the Transformer architecture) have limitations on the input length, by dividing the pure text document into multiple text blocks, it can be ensured that the size of each block meets the input requirements of the model, avoiding errors caused by exceeding the limit; at the same time, complex tasks can be decomposed into multiple small tasks, reducing the computational complexity of single processing.
[0103] In this embodiment, to ensure semantic integrity, for any picture, the associated paragraphs of the picture can be determined; the associated paragraphs and the content information corresponding to the picture are divided into text blocks.
[0104] It should be noted that the associated paragraphs include the preceding paragraphs and / or the following paragraphs. Among them, the preceding paragraphs refer to the associated paragraphs located above the picture, such as the text quoting above the picture; the following paragraphs refer to the associated paragraphs located below the picture, such as the explanatory paragraphs of the picture.
[0105] In specific implementation, the associated paragraphs of the picture can be identified through rules (such as keywords like "picture description:", "schematic diagram:") or models (such as an NLP (Natural Language Processing) classifier).
[0106] In this embodiment, by merging the content information of the picture and the associated paragraphs into a semantic block, it is possible to effectively avoid splitting the picture-text association, thereby retaining the complete logic in the initial document and improving the integrity and accuracy of knowledge.
[0107] In this embodiment, for the pure text content that has nothing to do with the content information of the picture, chunking operations can be performed according to a preset chunking strategy. Among them, the chunking strategy can include a sentence-based chunking strategy, a paragraph-based chunking strategy, or a semantic unit-based chunking strategy. Among them, the sentence-based chunking strategy is used to split the text by sentences, and each sentence or a group of sentences is used as a text chunk; the paragraph-based chunking strategy is used to chunk by paragraph, and each paragraph is used as an independent text chunk; the semantic unit-based chunking strategy is used to chunk according to the semantic structure of the text, such as dividing by theme, event, or logical unit.
[0108] S103-2: Input multiple text chunks into the large language model to obtain the entity relationship extraction results corresponding to each of the multiple text chunks output by the large language model.
[0109] In this embodiment, by inputting each text chunk into the large language model respectively, the powerful language understanding ability of the large language model can be utilized to analyze each text chunk, identify the entities therein (such as person names, organizations, locations, etc.) and the relationships between the entities (such as "belong to", "be located in", "be associated with", etc.).
[0110] Exemplarily, if there is a sentence in the text chunk: Company A released the latest version of Model B. The entities extracted by the large language model from this sentence are the two entities "Company A" and "the latest version of Model B", and the extracted relationship is "released", and the direction of the relationship is from "Company A" to "the latest version of Model B".
[0111] In this embodiment, based on some existing entities and relationships, the large language model can also predict the potential relationships between entities that are not explicitly defined by learning the known knowledge graph structure, thereby discovering new knowledge connections and enriching and improving the content of the knowledge graph.
[0112] Exemplarily, if the known knowledge graph structure indicates that the parent company of Company A is B and the parent company of B is C, then the large language model can infer that the parent company of A is C; another example is that if the known knowledge graph structure indicates that the headquarters of Company X is located in Shanghai and person Y works in Company X, then the large language model can infer that the place of residence of person Y is Shanghai.
[0113] In this embodiment, the entity relationship extraction results output by the large language model for each text chunk can be presented in a structured form (such as JSON format), including the name, type of the entity, and the relationship between the entities.
[0114] In this embodiment, during the process of using a large language model to extract entities and relationships, in order to enable the large language model to implement the function of entity relationship referring to pictures, a task of "source identification" will also be added when writing prompts, that is, a source identification prompt will be added. This source identification prompt is used to instruct the large language model to determine whether the entity relationship extraction result of a text block comes from a picture.
[0115] In this embodiment, the source identification prompt and multiple text blocks are input into the large language model, so that the large language model, in response to the source identification prompt, determines whether the entity relationship extraction results corresponding to the multiple text blocks come from pictures, and when it is determined that the entity relationship extraction result of any text block comes from a picture, the position information of the picture is stored as the position index of the entity relationship extraction result.
[0116] In a specific implementation, if the large language model determines that the entity relationship extraction result of a text block contains the position information of a picture, it determines that the entity relationship extraction result of the text block comes from a picture.
[0117] In this embodiment, if the large language model determines that the entity relationship extracted from any text block comes from a picture, the picture is stored as an attribute of the entity relationship, that is, the position index representing the picture source is stored in the attributes of the entity relationship. In this way, a mapping between the picture and the entity relationship is established.
[0118] In this embodiment, by establishing a mapping relationship between the entity relationship and the picture, the large language model can achieve the dual output of "answer + image evidence". For example, in a question-answering system built based on a large language model, when a user queries the large language model about "the appearance of a certain brand of car model", the large language model can directly return a text description and relevant pictures, and mark the picture source according to the position index (such as "quoted from page 2 of the document"), and the user can locate the position of the picture in the document according to the position index.
[0119] S103-3: Based on the entity relationship extraction results corresponding to multiple text blocks, obtain the entity relationship extraction result of the pure text document.
[0120] In this embodiment, considering that the same entity in different text blocks may have different expressions, such as "J.K. Rowling" and "the author of Harry Potter"; at the same time, the relationships between entities in different text blocks may be redundant or conflicting. Therefore, the entity relationship extraction results of each text block will be integrated to obtain a globally consistent entity relationship extraction result. Specifically, the entity relationship extraction result includes the entity name and / or the relationship between entities; S103-3 may include the following sub-steps:
[0121] S103-3-1: Align the entity names in multiple text blocks to obtain the aligned entity names; and / or, merge the relationships between entities in multiple text blocks to obtain the merged relationships between entities.
[0122] It should be noted that the alignment operation is used to merge entities in different text blocks that refer to the same real object into a unique identifier. Exemplarily, the entity lists extracted from multiple text blocks may have duplicates or variants, such as Variant 1, Variant 2, and Variant 3. Then, the set of unique entities after alignment is {entity ID: [Variant 1, Variant 2, Variant 3]}, where the entity ID is the unique entity name after alignment.
[0123] In this embodiment, to accurately determine whether entities in different text blocks refer to the same object, entity alignment will be achieved by calculating the similarity of entity names.
[0124] In a specific implementation, determine the string similarity, and / or semantic similarity, and / or context similarity between multiple entity names; based on the string similarity, and / or semantic similarity, and / or context similarity, determine the comprehensive similarity; based on the comprehensive similarity, perform an alignment operation on multiple entity names to obtain the aligned entity names.
[0125] In this embodiment, the string similarity is used to represent the similarity at the character level and can effectively handle surface differences such as spelling mistakes, abbreviations, and case variants. Specifically, algorithms such as the Levenshtein distance, Jaccard similarity, and cosine similarity can be used to calculate the string similarity between entity names.
[0126] In this embodiment, the semantic similarity is used to represent the similarity at the semantic level and can solve complex problems such as synonyms, aliases, and cross-language differences. Specifically, methods such as word vectors and semantic recognition models can be used to calculate the semantic similarity between entity names.
[0127] In this embodiment, the context similarity is used to represent the similarity of the context features of entities. Specifically, the context similarity between entity names can be determined based on the attribute consistency and co-occurring entities between entities. Among them, the attribute consistency represents the attribute similarity between different entity names. For example, the attributes of entity A are: {occupation: writer, nationality: UK}, and the attributes of entity B are {occupation: novelist, nationality: UK}, then it is determined that the attributes of entity A and entity B are similar. The co-occurring entity refers to other entities that co-occur with the entity. Entity A often co-occurs with Harry Potter and Hogwarts, and entity B often co-occurs with Harry Potter and Hogwarts, then it is determined that the contexts of entity A and entity B are similar.
[0128] In this embodiment, the comprehensive similarity can be determined by integrating the string similarity, semantic similarity, context similarity, as well as the first weight, second weight, and third weight; wherein, the sum of the first weight, second weight, and third weight is 1. Specifically, the comprehensive similarity can be calculated according to the following formula:
[0129] S = w1·a + w2·b + w3·c;
[0130] wherein, S represents the comprehensive similarity; a represents the string similarity; b represents the semantic string similarity; c represents the context similarity; w1 represents the first weight corresponding to the string similarity; w2 represents the second weight corresponding to the semantic string similarity; w3 represents the third weight corresponding to the context similarity.
[0131] In this embodiment, the first weight, second weight, and third weight can be dynamically adjusted according to requirements. For example, when dealing with abbreviations, the first weight corresponding to the string similarity can be increased.
[0132] In this embodiment, by comprehensively considering the string similarity, semantic similarity, and context similarity during entity alignment, the accuracy and robustness of alignment can be significantly improved, and thus a more complete and accurate entity relationship network can be generated.
[0133] In this embodiment, the merging operation is used to integrate the relationships between entities extracted from different blocks, eliminate redundancy, and resolve conflicts.
[0134] In this embodiment, the merging operation can specifically include standardization, matching and deduplication, and conflict handling. Among them, standardization is used to unify the format of relationship predicates. For example, "founded" and "established" are unified as "founded"; matching and deduplication are used to delete duplicate relationships between entities and merge relationships with the same semantics but different expressions. For example, "born in" and "place of birth" are unified as "place of birth"; conflict handling is used to handle contradictory descriptions of the same relationship between entities in different text blocks. Specifically, the time priority strategy can be adopted for conflict handling, that is, the data with the latest timestamp is selected as the final relationship between entities.
[0135] In this embodiment, by merging the relationships between entities in multiple text blocks, the consistency and accuracy of the relationships between entities can be ensured.
[0136] S103 - 3 - 2: Based on the aligned entity names and / or the merged relationships between entities, obtain the entity relationship extraction result of the plain text document.
[0137] In this embodiment, by performing entity alignment and relationship merging on multiple text blocks, the extraction results of scattered text blocks can be fused into a globally consistent knowledge representation, and thus a more complete and accurate knowledge graph can be constructed.
[0138] In a second aspect, based on the same inventive concept, referring to Figure 2 , an embodiment of the present application provides a knowledge graph construction device 200, and the knowledge graph construction device 200 includes:
[0139] A picture determination module 201, configured to determine pictures in an initial document and the position information of the pictures;
[0140] A document conversion module 202, configured to extract the content information of the pictures, and based on the content information and the position information, convert the initial document into a plain text document;
[0141] An information extraction module 203, configured to perform entity relationship extraction on the plain text document to obtain an entity relationship extraction result of the plain text document;
[0142] A graph construction module 204, configured to construct a knowledge graph based on the entity relationship extraction result.
[0143] In an embodiment of the present application, the document conversion module 202 includes:
[0144] A picture content recognition sub-module, configured to input the pictures into a multi-modal large language model to obtain the content information corresponding to the pictures output by the multi-modal large language model.
[0145] In an embodiment of the present application, the document conversion module 202 includes:
[0146] An information replacement sub-module, configured to replace the pictures with the content information and position information corresponding to the pictures to obtain a plain text document.
[0147] In an embodiment of the present application, the information extraction module 203 includes:
[0148] A document chunking sub-module, configured to divide the plain text document into multiple text chunks;
[0149] An information extraction sub-module, configured to input the multiple text chunks into a large language model to obtain the entity relationship extraction results corresponding to the multiple text chunks output by the large language model;
[0150] An information fusion sub-module, configured to obtain an entity relationship extraction result of the plain text document based on the entity relationship extraction results corresponding to the multiple text chunks.
[0151] In an embodiment of the present application, the document chunking sub-module includes:
[0152] An associated paragraph determination unit, configured to determine the associated paragraphs of the pictures; the associated paragraphs include the previous paragraphs and / or the subsequent paragraphs;
[0153] A text block division unit for dividing the content information corresponding to the associated paragraphs and pictures into a text block.
[0154] In an embodiment of the present application, the information extraction sub-module includes:
[0155] A prompt unit for inputting the source recognition prompt word and multiple text blocks into the large language model, so that the large language model determines whether the entity relationship extraction results corresponding to the multiple text blocks respectively come from pictures in response to the source recognition prompt word, and stores the position information of the picture as the position index of the entity relationship extraction result when it is determined that the entity relationship extraction result of any text block comes from a picture.
[0156] In an embodiment of the present application, the entity relationship extraction result includes entity names and / or relationships between entities; the information fusion sub-module includes:
[0157] An alignment unit for performing an alignment operation on the entity names in multiple text blocks to obtain the aligned entity names; and / or, performing a merging operation on the relationships between entities in multiple text blocks to obtain the merged relationships between entities;
[0158] A fusion unit for obtaining the entity relationship extraction result of the pure text document based on the aligned entity names and / or the merged relationships between entities.
[0159] In an embodiment of the present application, the alignment unit includes:
[0160] A first similarity determination sub-unit for determining the string similarity, and / or semantic similarity, and / or context similarity between multiple entity names;
[0161] A second similarity determination sub-unit for determining the comprehensive similarity based on the string similarity, and / or semantic similarity, and / or context similarity;
[0162] An entity name alignment sub-unit for performing an alignment operation on multiple entity names based on the comprehensive similarity to obtain the aligned entity names.
[0163] It should be noted that the specific implementation manner of the knowledge graph construction device 200 in the embodiment of the present application refers to the specific implementation manner of the knowledge graph construction method proposed in the first aspect of the embodiment of the present application, which will not be elaborated here.
[0164] In a third aspect, based on the same inventive concept, the embodiment of the present application provides a computer-readable storage medium, on which an executable program is stored, and when the executable program is executed by a processor, it implements the knowledge graph construction method proposed in the first aspect of the present application.
[0165] It should be noted that the specific implementation of the computer-readable storage medium in the embodiments of the present application refers to the specific implementation of the knowledge graph construction method proposed in the first aspect of the embodiments of the present application, which will not be elaborated here.
[0166] Fourthly, referring to Figure 3 , based on the same inventive concept, the embodiments of the present application provide an electronic device 300, including:
[0167] A memory 301 for storing an executable program;
[0168] A processor 302;
[0169] When the executable program is executed by the processor 302, the knowledge graph construction method proposed in the first aspect of the present application is implemented.
[0170] It should be noted that the specific implementation of the electronic device 300 in the embodiments of the present application refers to the specific implementation of the knowledge graph construction method proposed in the first aspect of the embodiments of the present application, which will not be elaborated here.
[0171] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0172] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0173] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions in the process Figure 1One or more processes and / or boxes Figure 1 The functions specified in one box or more boxes.
[0174] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process Figure 1 One or more processes and / or boxes Figure 1 The steps of the functions specified in one box or more boxes.
[0175] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.
[0176] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.
[0177] The above provides a detailed introduction to a method for constructing a knowledge graph, a storage medium and an electronic device. Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A knowledge graph construction method, characterized in that: The method comprises: Determine the image in the initial document and the location information of the image; Extracting content information of the image, and converting the initial document into a plain text document based on the content information and the location information; Performing entity relationship extraction on the plain text document to obtain an entity relationship extraction result of the plain text document; Based on the entity relationship extraction results, a knowledge graph is constructed.
2. The knowledge graph construction method according to claim 1, characterized in that: Extracting content information of the picture, including: The image is input into a multimodal large language model to obtain content information corresponding to the image output by the multimodal large language model.
3. The knowledge graph construction method according to claim 1, characterized in that: Based on the content information and the location information, converting the initial document into a plain text document includes: The image is replaced with the content information and location information corresponding to the image to obtain the plain text document.
4. The knowledge graph construction method according to claim 1, characterized in that: Performing entity relationship extraction on the plain text document to obtain an entity relationship extraction result of the plain text document includes: Dividing the plain text document into a plurality of text blocks; Inputting the multiple text blocks into a large language model, and obtaining entity relationship extraction results corresponding to the multiple text blocks output by the large language model; Based on the entity relationship extraction results corresponding to each of the multiple text blocks, the entity relationship extraction result of the plain text document is obtained.
5. The knowledge graph construction method according to claim 4, characterized in that: Divide the plain text document into multiple text blocks, including: Determine the associated paragraphs of the picture; the associated paragraphs include the preceding paragraph and / or the succeeding paragraph; The content information corresponding to the associated paragraphs and the pictures is divided into the text blocks.
6. The knowledge graph construction method according to claim 4, characterized in that: Inputting the multiple text blocks into a large language model to obtain entity relationship extraction results corresponding to the multiple text blocks output by the large language model, including: The source identification prompt word and the multiple text blocks are input into the large language model, so that the large language model determines whether the entity relationship extraction results corresponding to each of the multiple text blocks are derived from the picture in response to the source identification prompt word, and when it is determined that the entity relationship extraction result of any of the text blocks is derived from the picture, the position information of the picture is stored as the position index of the entity relationship extraction result.
7. The knowledge graph construction method according to claim 4, characterized in that: The entity relationship extraction result includes entity names and / or relationships between entities; based on the entity relationship extraction results corresponding to each of the multiple text blocks, the entity relationship extraction result of the plain text document is obtained, including: Performing an alignment operation on the entity names in the multiple text blocks to obtain aligned entity names; and / or, performing a merging operation on the entity relationships in the multiple text blocks to obtain merged entity relationships; Based on the aligned entity names and / or the merged entity relationships, an entity relationship extraction result of the plain text document is obtained.
8. The knowledge graph construction method according to claim 7, characterized in that: Performing an alignment operation on the entity names in the multiple text blocks to obtain aligned entity names includes: Determining string similarity, and / or semantic similarity, and / or contextual similarity between the plurality of entity names; Determining a comprehensive similarity based on the string similarity, and / or the semantic similarity, and / or the context similarity; Based on the comprehensive similarity, an alignment operation is performed on the multiple entity names to obtain aligned entity names.
9. A computer-readable storage medium having an executable program stored thereon, characterized in that: When the executable program is executed by the processor, it implements the knowledge graph construction method as described in any one of claims 1 to 8.
10. An electronic device, characterized in that: include: A memory for storing an executable program; processor; When the executable program is executed by the processor, the knowledge graph construction method as described in any one of claims 1 to 8 is implemented.