Multi-modal knowledge graph construction and retrieval method and system based on large language model
Through the multimodal knowledge graph construction method based on large language models, the applicability and efficiency of the knowledge graph generation method in the existing technology is solved, and high-precision multimodal data processing and complex semantic query are realized, and multi-level inheritance relationships and cross-domain knowledge fusion are supported.
Patent Information
- Application Number
- CN202510591095.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-15
AI Technical Summary
The existing knowledge graph generation methods have problems such as limited applicability, high maintenance cost, low efficiency, poor generalization ability, low accuracy of extraction results, insufficient semantic depth, and inability to integrate multimodal data.
A multimodal knowledge graph construction method based on a large language model is adopted. By receiving multimodal documents, a guided problem set is generated, a JSON model containing entity levels, attributes and relationships is built, and a domain knowledge graph is created in Neo4j, supporting deep collaborative processing and high-precision query of multimodal data.
It realizes high-precision semantic query of knowledge graphs, significantly improves query targeted and complex semantic reasoning capabilities, can express multi-level inheritance relationships, integrates text, images, and tabular data, and supports cross-domain knowledge fusion and multi-hop query.
Smart Images

Figure CN120492682A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and knowledge engineering technology, and in particular to a multimodal knowledge graph construction and retrieval method and system based on a large language model. Background Art
[0002] A knowledge graph is a series of various graphics that show the development process and structural relationship of knowledge. It uses visualization technology to describe knowledge resources and their carriers, and to mine, analyze, construct, draw and display knowledge and their interrelationships.
[0003] The main generation methods and shortcomings of existing knowledge graphs: (1) Rule-based method: This method relies on manually defined templates (such as regular expressions) to extract entities and relationships, which has limited applicability and high maintenance costs.
[0004] (2) Supervised learning method: This method requires a large amount of labeled data to train the model, which consumes a lot of manpower and material resources, is inefficient, has poor generalization ability and is difficult to migrate to new fields.
[0005] (3) Unsupervised / semi-supervised methods: This method reduces labeling dependency through clustering or remote supervision, reduces manpower and material resources, and improves efficiency, but the extraction results are less accurate and lack semantic depth.
[0006] (4) Neural network method: directly using large language models (LLM) to generate graphs, such as GPT-4, but the output lacks structured layers and there is a risk of hallucination.
[0007] (5) Knowledge graph generation tools: such as Neo4j's Graph Builder. This type of method is currently a hot topic of research. However, the relationships extracted are mostly explicit descriptions, such as "A belongs to B," lacking implicit semantics, such as "A's design flaw caused B's failure." The generated graphs are usually flat and cannot express multi-level inheritance relationships. In addition, most current tools only support text data and cannot integrate device structure diagrams in images or performance parameters in tables. Summary of the Invention
[0008] The purpose of the present invention is to provide a multimodal knowledge graph construction and retrieval method and system based on a large language model, which can automatically construct a hierarchical knowledge graph from multimodal data and realize high-precision semantic query.
[0009] To achieve this object, the present invention adopts the following technical solutions: A multimodal knowledge graph construction and retrieval method based on a large language model is provided, comprising the following steps: Receive multimodal documents uploaded by users; Read multimodal document information and generate a set of guiding questions based on the multimodal document information; Build a JSON model containing entity hierarchies, attributes, and relationships based on question sets and multimodal documents; Extract data from multimodal document data to fill and improve the JSON model and verify consistency; Map the populated JSON model to Neo4j nodes and edges, create a domain knowledge graph, and set corresponding properties; Generate corresponding Cypher query statements based on the questions or query requirements entered by the user, execute query operations in the Neo4j domain knowledge graph, and retrieve relevant knowledge and information; Generate user responses in natural language based on domain knowledge graph data information.
[0010] As a preferred solution for constructing a retrieval method based on a multimodal knowledge graph with a large language model, The step of reading the multimodal document information and generating a set of guiding questions based on the multimodal document information includes: Read and store multimodal documents; Segment multimodal documents into text blocks and image blocks; Loading a set of questions related to the current document domain from a stored historical question library; Combine the questions in the historical question library with the text block and image block content of the document to generate a series of initial question sets related to the document.
[0011] As a preferred solution for constructing a retrieval method based on a multimodal knowledge graph with a large language model, In the step of segmenting the multimodal document into text blocks and image blocks, the text content is segmented into a plurality of text blocks, and the image content is segmented into a plurality of image blocks.
[0012] The present invention also provides a multimodal knowledge graph construction retrieval system based on a large language model, characterized by comprising: The query generation agent is used to receive the multimodal documents uploaded by the user; read the multimodal document information and generate a set of guiding questions based on the multimodal document information; Domain model generation agent, which is used to construct a JSON model containing entity hierarchy, attributes and relationships based on the question set and multimodal documents; Domain model filling agent, used to extract data from multimodal document data to fill and improve the JSON model and verify consistency; The knowledge graph generation agent is used to map the populated JSON model into Neo4j nodes and edges, create a domain knowledge graph, and set corresponding properties; The knowledge graph query agent is used to generate corresponding Cypher query statements based on the questions or query requirements entered by the user, perform query operations in the Neo4j domain knowledge graph, retrieve relevant knowledge and information; and generate user responses in the form of natural language based on the domain knowledge graph data information.
[0013] Beneficial effects of the present invention: 1. Significantly improve the query relevance of the knowledge graph; this invention proactively guides the knowledge graph construction process by generating a predefined set of guiding questions, ensuring that the generated entity hierarchy, attributes, and relationships are highly aligned with user needs. Compared to the passive retrieval model of traditional methods, this design can reduce the generation of a large number of irrelevant entities and accurately identify key relationships. This targeted extraction mechanism significantly improves the accuracy of subsequent query responses, effectively addressing the fragmented answers and poor practicality of traditional RAG systems.
[0014] 2. Implement complex semantic reasoning and deep relationship mining. By designing a domain model in a multi-level nested JSON format, this system can express semantic relationships at least three levels deep, such as "construction process → pile expansion → mechanical pile expansion → mechanical equipment → reaming drilling tools." This supports advanced application scenarios such as causal chain tracing, cross-domain knowledge integration, and multi-hop querying.
[0015] 3. Multimodal data integration; the present invention realizes deep collaborative processing of text, image and table data. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.
[0017] Figure 1 This is a flow chart of a retrieval method for constructing a multimodal knowledge graph based on a large language model according to an embodiment of the present invention; Figure 2 This is a schematic block diagram of a structure of a retrieval system based on a multimodal knowledge graph of a large language model according to an embodiment of the present invention; Figure 3 A multimodal document according to an embodiment of the present invention; Figure 4 A set of questions according to an embodiment of the present invention; Figure 5This is a JSON model according to an embodiment of the present invention; Figure 6 Creating a Neo4j node according to an embodiment of the present invention; Figure 7 The creation of a Neo4j relationship in one embodiment of the present invention; Figure 8 A knowledge graph for generating bored pile construction specifications according to an embodiment of the present invention DETAILED DESCRIPTION The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0018] The following describes the embodiments of the present disclosure through specific examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.
[0019] Reference Figure 1 An embodiment of the present invention provides a multimodal knowledge graph construction and retrieval method based on a large language model, comprising the following steps: S01: User uploads multimodal documents; the present invention supports users to upload multimodal documents in various formats, such as operation manuals, maintenance reports, standard operating procedures (SOPs) in PDF format, etc. These documents contain rich text and image information and are the basic data source for knowledge graph construction.
[0020] S02: Read document information, query and generate an intelligent agent to simulate the thinking of domain experts to generate a set of guiding questions; S03: The domain model generation agent constructs a JSON model containing entity hierarchy, attributes and relationships based on the question set and user documents; The Domain Model Generation Agent uses question sets and user documents as its foundation. Using natural language processing and information extraction techniques, it identifies entities, attributes, and relationships within the documents and organizes them into a hierarchical JSON model. This model defines the hierarchical structure of entities, their attributes, and the relationships between them, providing a detailed framework for building a knowledge graph.
[0021] S04: The domain model filling agent extracts data from the document data to fill in and improve the JSON model and verify the consistency; The domain model population agent extracts specific data from the document's text and image blocks and populates it into the corresponding entities and attributes in the JSON model, further improving the model's content. Simultaneously, the populated model undergoes consistency verification to ensure that the information is accurate, logically coherent, and in line with the specifications and requirements of the domain knowledge.
[0022] S05: The knowledge graph generation agent maps the filled JSON model to Neo4j nodes and edges, creates a domain knowledge graph, and sets corresponding properties, thereby creating a structured domain knowledge graph in Neo4j to achieve knowledge visualization and associative storage.
[0023] S06: The knowledge graph query agent generates Cypher statements to query the Neo4j domain knowledge graph; The knowledge graph query agent generates corresponding Cypher query statements based on the user's questions or query requirements, executes query operations in the Neo4j domain knowledge graph, retrieves knowledge and information related to the questions, and provides data support for generating user responses.
[0024] S07: Generate user responses based on domain knowledge graph data. Based on the data information obtained by the knowledge graph query agent from the Neo4j domain knowledge graph, the system combines natural language generation technology to generate accurate, clear, and complete user responses. The query results are presented to the user in natural language to meet the user's question-and-answer needs.
[0025] In some embodiments, the above step S02 includes the following steps: S021: User uploads multimodal document; the user uploads the multimodal document to the server through the system interface, and the system reads and stores the document in preparation for subsequent processing.
[0026] S022: The system segments the uploaded multimodal document into text blocks and image blocks; the system segments the uploaded multimodal document into multiple text blocks and multiple image blocks, so as to facilitate subsequent separate processing and information extraction.
[0027] S023: The query generation agent loads the historical question library; the query generation agent loads a set of questions related to the current document domain from the stored historical question library. These historical questions provide the agent with templates and references for domain questions, which helps to generate more targeted guiding questions.
[0028] S024: Generate an initial set of questions. The query generation agent combines questions from the historical question library with the text and image content of the document, simulating the thinking of domain experts to generate a series of initial questions related to the document. These questions are intended to guide the subsequent information extraction and knowledge graph construction process, covering key information points and potential knowledge associations in the document.
[0029] Take the construction of a knowledge graph of bored pile construction specifications as a specific implementation case: S01: User input document The public construction specification "Construction Standard for Bored and Cast-in-Place Piles (DG / TJ 08-202-2020J11042-2020)" contains multimodal data such as text, illustrations, and tables. Figure 3 shown.
[0030] S02: Generate a set of guiding questions, such as Figure 4 shown.
[0031] S03-04: Build a JSON model containing entity hierarchy, attributes and relationships, extract data from document data to fill in and improve the JSON model, and verify consistency. Figure 5 shown.
[0032] S05: Map the populated JSON model to Neo4j nodes and edges, create a domain knowledge graph, and set the corresponding properties, such as Figure 6 and Figure 7 shown.
[0033] S06 generates a knowledge graph of bored pile construction specifications, such as Figure 8 shown.
[0034] Reference Figure 2 An embodiment of the present invention further provides a multimodal knowledge graph construction and retrieval system based on a large language model, characterized by including: The query generation agent is used to receive the multimodal documents uploaded by the user; read the multimodal document information and generate a set of guiding questions based on the multimodal document information; Domain model generation agent, which is used to construct a JSON model containing entity hierarchy, attributes and relationships based on the question set and multimodal documents; Domain model filling agent, used to extract data from multimodal document data to fill and improve the JSON model and verify consistency; The knowledge graph generation agent is used to map the populated JSON model into Neo4j nodes and edges, create a domain knowledge graph, and set corresponding properties; The knowledge graph query agent is used to generate corresponding Cypher query statements based on the questions or query requirements entered by the user, perform query operations in the Neo4j domain knowledge graph, retrieve relevant knowledge and information; and generate user responses in the form of natural language based on the domain knowledge graph data information.
[0035] In the description of the present invention, it should be understood that the terms "middle", "length", "upper", "lower", "front", "back", "vertical", "horizontal", "inner", "outer", "radial", "circumferential", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0036] In the present invention, unless otherwise expressly specified or limited, a first feature "on" a second feature may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediary. "Multiple" means at least two, such as two or three, unless otherwise expressly specified or limited.
[0037] In the present invention, unless otherwise specified or limited, the terms "installed," "connected," "connect," "fixed," etc. should be understood in a broad sense. For example, they can refer to fixed connection, detachable connection, or integration; mechanical connection, electrical connection, or communication; direct connection or indirect connection through an intermediate medium; internal communication between two elements or interaction between two elements, unless otherwise specified. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0038] The above is only for explaining the embodiments of the present invention and is not intended to limit the present invention. For those skilled in the art, any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present invention without creative work should be included in the scope of protection of the present invention.
Claims
1. A multimodal knowledge graph construction and retrieval method based on a large language model, characterized in that: The following steps are involved: Receive multimodal documents uploaded by users; Read multimodal document information and generate a set of guiding questions based on the multimodal document information; Build a JSON model containing entity hierarchies, attributes, and relationships based on question sets and multimodal documents; Extract data from multimodal document data to fill and improve the JSON model and verify consistency; Map the populated JSON model to Neo4j nodes and edges, create a domain knowledge graph, and set corresponding properties; Generate corresponding Cypher query statements based on the questions or query requirements entered by the user, execute query operations in the Neo4j domain knowledge graph, and retrieve relevant knowledge and information; Generate user responses in natural language based on domain knowledge graph data information.
2. The multimodal knowledge graph construction and retrieval method based on a large language model according to claim 1 is characterized in that: The step of reading the multimodal document information and generating a set of guiding questions based on the multimodal document information includes: Read and store multimodal documents; Segment multimodal documents into text blocks and image blocks; Loading a set of questions related to the current document domain from a stored historical question library; Combine the questions in the historical question library with the text block and image block content of the document to generate a series of initial question sets related to the document.
3. The multimodal knowledge graph construction and retrieval method based on a large language model according to claim 2 is characterized in that: In the step of segmenting the multimodal document into text blocks and image blocks, the text content is segmented into a plurality of text blocks, and the image content is segmented into a plurality of image blocks.
4. A multimodal knowledge graph construction retrieval system based on a large language model, characterized by: include: Query generation agent, which is used to receive multimodal documents uploaded by users; and reading multimodal document information and generating a set of guiding questions based on the multimodal document information; Domain model generation agent, which is used to construct a JSON model containing entity hierarchy, attributes and relationships based on the question set and multimodal documents; Domain model filling agent, used to extract data from multimodal document data to fill and improve the JSON model and verify consistency; The knowledge graph generation agent is used to map the populated JSON model into Neo4j nodes and edges, create a domain knowledge graph, and set corresponding properties; The knowledge graph query agent is used to generate corresponding Cypher query statements based on the questions or query requirements entered by the user, execute query operations in the Neo4j domain knowledge graph, and retrieve relevant knowledge and information; And based on the domain knowledge graph data information, user responses are generated in the form of natural language.