Complex problem decomposition and multi-modal knowledge retrieval method combined with large model

By combining the complex problem decomposition and multimodal knowledge retrieval methods of large models, multimodal knowledge data are extracted and transformed, complex problems are decomposed and sub-problems are generated for knowledge retrieval, and finally a knowledge graph is constructed to generate answers, which solves the shortcomings of existing systems in dealing with complex problems and multimodal knowledge retrieval, and achieves more efficient and accurate answer generation.

CN119938832AActive Publication Date: 2025-05-06NORTHEASTERN UNIV CHINA

Patent Information

Application Number
CN202411982091.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Existing machine reading comprehension and question-answer systems have shortcomings in dealing with complex problems and retrieving multimodal knowledge, making it difficult to accurately understand and analyze complex problems, and lack the ability to efficiently integrate and process multimodal information.

Method used

Using a method of combining complex problem decomposition and multimodal knowledge retrieval of large models, the semantic information of image modal knowledge data in the multimodal knowledge base is extracted, and complex problems are decomposed based on entity recognition and dependency syntax analysis model, sub-problems are generated for knowledge retrieval, and finally a knowledge graph is constructed to generate answers.

Benefits of technology

It improves the accuracy and comprehensiveness of large models when dealing with complex problems, enhances the retrieval ability of multimodal knowledge base, can more accurately understand and analyze complex problems, and generate detailed answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938832A_ABST
    Figure CN119938832A_ABST
Patent Text Reader

Abstract

The invention provides a complex problem decomposition and multi-modal knowledge retrieval method combined with a large model, and relates to the technical field of artificial intelligence and information retrieval. According to the method, for a multi-modal knowledge base, semantic information contained in image modal knowledge data is extracted, a knowledge retrieval method with the image understanding ability is provided, and semantic and structural relations among numerous sub-problems possibly contained in complex problems are constructed by decomposing the complex problems; the method is used for complex question decomposition and multi-modal knowledge retrieval to generate accurate and comprehensive answers, so that a large model is helped to answer complex questions more accurately and more comprehensively. In addition, by integrating related background knowledge of complex problems, the retrieval capability of a retrieval tool on the multi-modal knowledge base is expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and information retrieval technology, and in particular to a complex problem decomposition and multimodal knowledge retrieval method combined with a large model. Background Art

[0002] With the widespread dissemination of Internet multimedia information and the rapid development of artificial intelligence technology, machine reading comprehension and question-answering systems have become a research hotspot. The core of this type of task is to understand the questions raised by users, retrieve the relevant information required for the questions and generate accurate answers based on the knowledge base provided by users or built into the system. This research direction can not only significantly reduce the time users spend searching and reading materials, but also improve the flexibility of the question-answering interaction process, and has important application value in scenarios such as human-computer dialogue, intelligent teaching assistants, and medical consultations.

[0003] In recent years, large language models (referred to as large models) represented by ChatGPT, GPT-4, and LLaMA systems have demonstrated excellent inductive summarization and diverse instruction-following capabilities with their large model parameter volume and pre-trained corpus size. In order to further improve the accuracy of large models in answering questions (i.e., reduce the large model hallucination problem), the knowledge base is retrieved according to the user's question to provide context for the large model. This method of combining knowledge retrieval to enhance capabilities has become the mainstream method for current machine reading comprehension and question-answering tasks. However, this method has certain limitations: first, it lacks the ability to handle complex problems. The questions raised by users often contain multiple sub-questions, and even further reasoning about the questions is required to understand the user's intentions. In such scenarios, large models may find it difficult to accurately understand and analyze the problems, and retrieval tools cannot directly match the content required to answer complex questions from the knowledge base. Secondly, in many application scenarios (such as education and medical care), users' questions may involve information combined with images and text. At this time, key information may need to be answered in combination with images, tables, or other non-textual data. However, most current methods still focus on text modality retrieval, which makes it difficult to efficiently integrate and process multimodal information.

[0004] In summary, although the current machine reading comprehension and question-answering systems have made great progress driven by large models, they still face challenges such as the lack of ability to decompose complex problems, retrieve and integrate multimodal knowledge, etc. Therefore, in order to solve the problem of what knowledge to retrieve and how to retrieve knowledge, developing a method with a large model as the core that can handle complex problems and retrieve multimodal knowledge is the key to improving the accuracy and comprehensiveness of the answers of the question-answering system.

[0005] As the general capabilities and instruction-following abilities of large models continue to increase, existing question-answering systems have shifted to knowledge retrieval research oriented toward large models. That is, different retrieval algorithms are used to retrieve corresponding knowledge for related questions, thereby answering questions based on the summarization and reasoning abilities of large models.

[0006] Simply put, because people ask flexible and diverse questions, and the background knowledge required to answer the corresponding questions is also different, traditional machine reading comprehension and question answering systems can no longer meet people's needs. This is because as the complexity and flexibility of the questions increase, it is difficult for the understanding model to correctly answer the user's questions, which is why traditional tasks are to answer questions in the form of multiple-choice questions.

[0007] As a generative model, the big model has not only accumulated a lot of basic knowledge in the pre-training stage, but also can follow the various instructions given by the user to a certain extent. The most important thing is that the big model can generate corresponding answers according to the user's different questioning styles and requirements, so it is more practical.

[0008] The Chinese patent "CN202410213818 A large language model question and answer method and device based on knowledge retrieval enhancement" enhances the question and answer capability of the large model by retrieving the knowledge that may be involved in the question from the knowledge base, and then combining historical questions and answers in multiple rounds of conversations.

[0009] The Chinese patent "CN202410343664 A large language model reasoning method and system based on multi-level knowledge retrieval enhancement" optimizes the accuracy and relevance of the answers generated by the large model through layer-by-layer retrieval of the operation knowledge base and the industry knowledge base. The specific process is as follows: First, the question statement is vectorized and compared with the questions and answers in the operation knowledge base. If successful, the corresponding answer is output; if it fails, it turns to the industry knowledge base, obtains highly relevant text fragments through segmented vectorization matching, and generates answers in combination with the large model. The core methods include BGE algorithm vectorization, cosine similarity calculation, multi-way recall and precise sorting. Aims to alleviate the "hallucination" of large models and achieve efficient and accurate industry knowledge question and answer

[0010] The current related technologies either focus on how to better search the knowledge base or how to sort the retrieved knowledge, but ignore the question of "what knowledge should be retrieved". In other words, the existing technology has an assumption that the given question (also known as the query) itself is a text that can be directly used to search the knowledge base. This text itself is not complicated and its intent and semantics can be directly understood by the search tool. However, in actual applications, the questions asked by users may be very complex. It is necessary to carefully analyze and understand the sub-questions and reasoning paths behind the complex questions, and decompose them before they can be used directly by the search tool to search the knowledge base, or use the search method optimized by the existing technology to retrieve the corresponding knowledge.

[0011] In addition, the knowledge bases used in the prior art are basically text-based. However, in many cases, when the knowledge base is provided in a multimodal form (especially in a document form), it contains a lot of image content, such as PDF, word documents, etc. Some content is not directly contained in the text, but requires reference to the accompanying images to understand. Therefore, the prior art still lacks a simple and efficient method for searching a multimodal knowledge base. Summary of the invention

[0012] The technical problem to be solved by the present invention is to address the deficiencies of the above-mentioned prior art and to provide a complex problem decomposition and multimodal knowledge retrieval method combining a large model, which is used for complex problem decomposition and multimodal knowledge retrieval to generate accurate and comprehensive answers.

[0013] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0014] The present invention provides a complex problem decomposition and multimodal knowledge retrieval method combined with a large model, which specifically includes the following steps:

[0015] Step 1: Select the large model to be used as the basic model;

[0016] Step 2: According to the specific business scenario requirements, build a named entity recognition dataset for training the selected basic model to obtain the trained basic model f LLM ;

[0017] Constructing named entity recognition dataset D NER ={(S p ; {e p1 ,e p2 ,…,e pd})|p=1,2,…,D} is used to train the basic model, where S p is the pth text in a data set with a sample size of D, {e p1 ,e p2 ,…,e pd} is the named entity recognition dataset D NER The d entities contained in the p-th text;

[0018] Step 3: Establish a multimodal knowledge base, preprocess it, and convert the multimodal knowledge base into a plain text knowledge base, specifically including: extracting the semantic information contained in the image modal knowledge data in the multimodal knowledge base, converting the image modal knowledge data into text modal knowledge data, and then converting the text modal knowledge data into a text modal knowledge vector;

[0019] The multimodal knowledge base refers to a knowledge base containing text-image modality knowledge data, and the data in the knowledge base is not limited to be multimodal data. The modality of the data depends on the actual business and usage scenarios. If a given knowledge base contains image modality knowledge data, the image modality data needs to be preprocessed;

[0020] The multimodal knowledge base includes the system's built-in knowledge base and the user-provided knowledge base. The system's built-in knowledge base refers to the existing professional knowledge base, including but not limited to knowledge graphs, graph databases, and vector query databases; the user-provided knowledge base refers to the multimodal context that users provide to the big model when using it to help answer questions, that is, general text or documents provided by the user;

[0021] When preprocessing the multimodal data contained in the multimodal knowledge base, the system's built-in knowledge base only needs to be processed offline once, and the data contained in the system's built-in knowledge base can be directly retrieved and used after preprocessing operations; for the user-given knowledge base, real-time processing is required; if the user does not provide a knowledge base when asking a question, the real-time processing of the user-given knowledge base is skipped, and only the system's built-in knowledge base is preprocessed; conversely, if the user provides a knowledge base when asking a question, the preprocessing of the system's built-in knowledge base is skipped, and only the user-provided knowledge base is processed in real time;

[0022] Step 3.1: Use the optical character recognition tool OCR Extract the optical character information of each image contained in all image modal knowledge data in the multimodal knowledge base, and use a large model f with multimodal processing capabilities MLLM Identify semantic information contained in the optical character information of each image;

[0023] Step 3.1.1: Use the optical character recognition tool OCR Extracting optical character information of each image contained in all image modality knowledge data in a multimodal knowledge base;

[0024] Using optical character recognition tools OCRExtract the optical character information in the image modality knowledge data. For the image I located at position P in the multimodal knowledge base, the optical character information of the image is represented by OCR=f OCR (I), if the image does not contain optical character information, the OCR is empty;

[0025] Step 3.1.2: Based on the optical character recognition information, position information and context information of the image, a large model with multimodal processing capabilities is used to identify the semantic information contained in the optical character information of each image;

[0026] For the position P of image I, the content above the position δ1 row away is defined as C 1 (δ1), the next content δ2 rows away is C 2 (δ2), the optical character information of the image is finally recognized to contain the semantic information F:

[0027] F=f MLLM (I; C 1 (δ1);OCR;C 2 (δ2)) (1)

[0028] Step 3.2: Replace the image corresponding to each position in the multimodal knowledge base with the optical character information of the image containing semantic information, and convert the multimodal knowledge base into a plain text knowledge base;

[0029] Step 4: Receive a complex question given by the user, decompose the complex question based on the entity recognition model, and use the dependency syntax analysis model to obtain the entity information in the complex question;

[0030] The complex problem is a problem composed of two or more entities, or a problem composed of two or more sub-problems;

[0031] Step 4.1: Based on the entity recognition model f NER Obtaining the set of potential entities in complex problems As shown in the following formula:

[0032]

[0033] Among them, Q is a complex problem, {e1,e2,…,e E} are the E potential entities identified from the complex question Q;

[0034] Step 4.2: Parse the model f through syntax DSP Obtaining key entity sets in complex problems And update the potential entity set to obtain the updated potential entity set

[0035] Key entities are defined as all types of nouns obtained after dependency syntactic analysis;

[0036] Key Entity Collection And the updated potential entity set is shown in the following formula:

[0037]

[0038] Among them, {s1,s2,…,s N} is the syntax analysis model f DSP The N key entities in the complex problem obtained, {c1,c2,…,c m} are the m key potential entities in the updated potential entity set;

[0039] Step 4.3: Based on dependency syntactic structure and key entity set A set of related entities is formed by direct, indirect or clause-based associations between nouns. Among them (s i ,s j ) is the associated entity pair consisting of the i-th associated entity and the j-th associated entity, (s i ,s j ) z is the zth associated entity pair in the associated entity set, and n is the number of associated entity pairs contained in the complex question Q;

[0040] Step 5: According to the entity information in the complex problem, the complex problem is converted into multiple sub-problems, and the direct knowledge and related knowledge of each sub-problem are retrieved from the plain text knowledge base;

[0041] Step 5.1: Update the potential entity set The key potential entities in the direct question template T sub Generate a set of direct subproblems Q sub ;

[0042] Based on direct question template T sub , generate a set of direct sub-problems Q sub As shown below:

[0043]

[0044] Among them, q y The direct sub-question set Q sub The direct sub-problem in c y is a key potential entity; for any key potential entity c y , after filling in the direct question template, generate its corresponding direct sub-question q y , until the updated potential entity set All the key potential entities in are filled;

[0045] Step 5.2: Associating entity collections Based on the associated question template T rel Generate a set of associated sub-problems Q rel ;

[0046] Based on the associated question template T rel , generate a set of associated sub-problems Q rel As shown below:

[0047]

[0048] Among them, q i,j is the set of associated sub-problems Q rel The associated subproblem in (q i,j ) w is the set of associated sub-problems Q rel The w-th associated subproblem in ;

[0049] Step 5.3: Retrieve the set of direct sub-questions Q in the plain text knowledge base sub Each direct subproblem q y , obtain the information about each direct sub-problem q y Direct knowledge of K sub ;

[0050] For the direct sub-problem q y , using a search tool that matches a plain text knowledge base retrieval Retrieve relevant knowledge from the plain text knowledge base. If direct knowledge is not retrieved, the basic model will answer the question separately as reference knowledge and generate direct knowledge K sub If the retrieved direct knowledge is one or more than one, all the retrieved direct knowledge are sorted according to the relevance scores of each direct knowledge according to the basic model, and each direct knowledge is spliced ​​from high to low in terms of relevance scores to generate direct knowledge K sub :

[0051] K sub =f retrieval (Q sub )={k u |u=1,2,…,U} (6)

[0052] Among them, k u For the direct subproblem q y The u-th direct knowledge, U is the number of direct knowledge, f retrieval To use search tools to search in plain text knowledge base and return the retrieved knowledge;

[0053] Step 5.4: Retrieve the collection of related sub-questions Q in the plain text knowledge base rel Each associated subproblem q i,j , get the associated sub-problem q i,j The associated knowledge K rel ;

[0054] For the associated sub-problem q i,j , retrieve its related knowledge from the plain text knowledge base. If no related knowledge is retrieved, concatenate the “maybe there is no relationship” and the answer directly answered by the basic model as reference knowledge and generate related knowledge K rel If the retrieved associated knowledge is one or more than one, all the retrieved associated knowledge are sorted according to the relevance scores of each associated knowledge according to the basic model, and each associated knowledge is spliced ​​from high to low in terms of relevance scores to generate associated knowledge K rel :

[0055] K rel =f retrieval (Q rel )={(k i,j ) v |v=1,2,…,V} (7)

[0056] Among them, (k i,j ) v For the associated subproblem q i,j The vth associated knowledge, V is the number of associated knowledge;

[0057] Step 6: Embedding model f based on knowledge graph KGE Construct a knowledge graph G for complex problems, sub-problem sets, direct knowledge related to sub-problems, and associated knowledge;

[0058] Step 6.1: Define the set R of binary logical relations between all key entities in the complex problem;

[0059] The binary logical relationship set R between the key entities in the complex problem includes three types: symmetric relationship, antisymmetric relationship and no relationship;

[0060] For the potential entity set Any two entities e l and e h , the symmetric relationship is shown as follows:

[0061]

[0062] Among them, r is the relationship between two entities. l and e hUnder the premise that there is a relationship r, e can be directly derived h and e l If there is also a relationship r between them, then there is a symmetric relationship between the two entities;

[0063] The antisymmetric relationship is shown below:

[0064]

[0065] In contrast to the symmetric relationship, knowing the entity e l and e h Under the premise that there is a relationship r, it can be deduced that e h and e l If there is no relationship r between them, then there is an antisymmetric relationship between the two entities;

[0066] No relationship, that is, there is no semantic relationship between the two entity pairs;

[0067] Step 6.2: Use the knowledge graph to embed the model f KGE The encoder f emb , the complex problem, the direct knowledge corresponding to the direct sub-problems obtained by decomposing the complex problem, and the associated knowledge corresponding to the associated sub-problems obtained by decomposing the complex problem are all converted into the knowledge graph embedding model f KGE The required embedding vector;

[0068] Using the knowledge graph embedding model f KGE The encoder f emb , get the embedding vector E of the complex question Q Q =f emb (Q);

[0069] Each direct subproblem q y and several pieces of direct knowledge, so as to integrate the entity and its background knowledge, and use the knowledge graph embedding model f KGE The encoder f emb , get the embedding vector E of direct knowledge y =f emb (q y ;k u );

[0070] Each associated subproblem q i,j and several pieces of associated knowledge are spliced ​​together to fuse the background knowledge of associated entity pairs and the possible relationships between them, and use the knowledge graph embedding model f KGE The encoder f emb , get the embedding vector E of the associated knowledge i,j =f emb (q i,j ;(ki,j ) v );

[0071] Step 6.3: Input the embedding vector of complex question Q, the embedding vector of direct knowledge, and the embedding vector of associated knowledge into the knowledge graph embedding model f KGE , get the knowledge graph G corresponding to the decomposition of the complex problem;

[0072] Using the knowledge graph embedding model f KGE The resulting set of potential entities Any entity e in l and e h The logical relationship between them is:

[0073] r l.h =f KGE (E Q ,E l ,E h ,E l,h ),r l,h ∈R (10)

[0074] Further obtain the potential entity set The logical relationship in lh Any entity e l and e h The knowledge graph G is:

[0075] G={{(e l ,k l ),r l.h ,(e h ,k h )} g |g=1,2,…,G} (11)

[0076] Step 7: Integrate the elements contained in the knowledge graph G into a new question Q * , and the trained basic model f LLM Generate new question Q * Complete and generate answer A;

[0077] Step 7.1: Determine the entity-direct knowledge pairs (e l ,k l ) exceeds the maximum context length specified by the basic model, and summarizes and abbreviates the direct knowledge that exceeds the maximum context length specified by the basic model so that the overall length of the direct knowledge does not exceed the maximum context length Θ of the basic model;

[0078] Define the maximum context length Θ that the base model can handle, and define each entity-direct knowledge pair (e l ,kl )'s maximum length θ l , if an entity-direct knowledge pair (e l ,k l ) exceeds its specified maximum length θ l , then the basic model f LLM Direct knowledge of l Summarize and abbreviate until the length limit is met, so the maximum context length θ that the base model can handle is:

[0079]

[0080] Step 7.2: Fill the elements contained in the original complex question Q and the knowledge graph G into the new question template T in sequence new Generate new question Q * , and the trained basic model f LLM According to the new question Q * Generate response A:

[0081]

[0082] Define a new question template T new , fill each element contained in the knowledge graph G into the new question template T in turn new After that, a new question Q is formed * , based on the trained basic model f LLM The generated response A is the new question Q * The answer to , which is also the answer to the original question Q.

[0083] The beneficial effect of adopting the above technical solution is that: the present invention provides a complex problem decomposition and multimodal knowledge retrieval method combined with a large model. For a multimodal knowledge base, by extracting semantic information contained in image modal knowledge data, a knowledge retrieval method with image understanding ability is provided. By decomposing complex problems, semantic and structural relationships between numerous sub-problems that may be contained in complex problems are constructed, thereby helping the large model to answer complex problems more accurately and comprehensively. In addition, by integrating relevant background knowledge of complex problems, the retrieval ability of the retrieval tool for the multimodal knowledge base is expanded. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 A flowchart of a complex problem decomposition and multimodal knowledge retrieval method combining a large model provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0085] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0086] This implementation is a complex problem decomposition and multimodal knowledge retrieval method that combines large models, such as Figure 1 As shown, the following steps are included:

[0087] Step 1: Select the large model to be used as the basic model;

[0088] This embodiment does not limit the large language model (LLM) selected, and a large model with multimodal processing capabilities can be selected, including but not limited to GPT-4o, Gemini, LLaVA and LLaMA; in some embodiments, if a large model without multimodal processing capabilities is selected, it is necessary to additionally use a model with image content description capabilities, such as a BLIP model with image content understanding capabilities. In this embodiment, a large model with multimodal processing capabilities is selected as the basic model f LLM ;

[0089] Step 2: According to the specific business scenario requirements, build a named entity recognition dataset for training the selected basic model to obtain the trained basic model f LLM ;

[0090] In this embodiment, business scenarios include but are not limited to real application scenarios in specific fields such as education, medical, and legal fields;

[0091] Named Entity Recognition (NER) refers to the technology of identifying all potential named entities (hereinafter referred to as entities) in a given text. Potential entities include names of people, places, names of organizations, and professional terms in specific business scenarios.

[0092] Constructing named entity recognition dataset D NER ={(S p ; {e p1 ,e p2 ,…e pD})|p=1,2,…,D} is used to train the basic model, where S i is the pth text in a dataset with D samples, which is usually a natural sentence. p1 ,e p2 ,…e pd} is the named entity recognition dataset D NERThe d entities contained in the pth text; for example, the text "A obtained B" is constructed into two entities "A" and "B", forming a named entity recognition dataset D NER A sample of

[0093] In this embodiment, it is not limited to build a named entity recognition dataset. Whether to build a named entity recognition dataset can be determined based on whether there is an optimization requirement for one or more business scenarios. If there is an optimization requirement for one or more business scenarios, a named entity recognition dataset is built to optimize the basic model f LLM Otherwise, there is no need to build a named entity recognition dataset. Use prompt engineering and other technologies to build a named entity recognition dataset for the required business scenarios, and fine-tune the basic model f LLM training;

[0094] Step 3: Establish a multimodal knowledge base, preprocess it, and convert the multimodal knowledge base into a plain text knowledge base, specifically including: extracting the semantic information contained in the image modal knowledge data in the multimodal knowledge base, converting the image modal knowledge data into text modal knowledge data, and then converting the text modal knowledge data into a text modal knowledge vector;

[0095] In this embodiment, a multimodal knowledge base refers to a knowledge base that contains text-image modal knowledge data, and it is not limited to that the data in the knowledge base must be multimodal data. The modality of the data depends on the actual business and usage scenarios. If a given knowledge base only contains knowledge data in pure text modality, it does not affect the subsequent steps of this embodiment. If a given knowledge base contains image modality knowledge data, the image modality data needs to be preprocessed.

[0096] The multimodal knowledge base includes the system's built-in knowledge base and the user-provided knowledge base. The system's built-in knowledge base refers to the existing professional knowledge base, including but not limited to knowledge graphs, graph databases, and vector query databases. The user-provided knowledge base refers to the multimodal context that users provide to the big model when using it, which helps answer questions, that is, general text or documents provided by the user.

[0097] When preprocessing the multimodal data contained in the multimodal knowledge base, the system's built-in knowledge base only needs to be processed offline once. The data contained in the system's built-in knowledge base can be directly retrieved and used after preprocessing operations are performed, without repeating the preprocessing operations. For the user-given knowledge base, real-time processing is required to meet the user's special needs. If the user does not provide a knowledge base when asking a question, the real-time processing of the user-given knowledge base is skipped, and only the system's built-in knowledge base is preprocessed. On the contrary, if the user provides a knowledge base when asking a question, the preprocessing of the system's built-in knowledge base is skipped, and only the user-provided knowledge base is processed in real time.

[0098] Step 3.1: Use the optical character recognition tool OCR Extract the optical character information of each image contained in all image modal knowledge data in the multimodal knowledge base, and use a large model f with multimodal processing capabilities MLLM Identify semantic information contained in the optical character information of each image;

[0099] Since existing knowledge retrieval tools are often based on text retrieval, they lack the ability to retrieve image data, or directly provide image content to large models through visual encoders, which greatly affects the context processing ability of large models. In order to enhance the retrieval ability of image modal knowledge data and reduce the amount of image modal knowledge data, this embodiment uses the method of modal conversion to use optical character recognition tools f OCR and image content description techniques to convert image modality information into text modality;

[0100] Step 3.1.1: Use the optical character recognition tool OCR Extracting optical character information of each image contained in all image modality knowledge data in a multimodal knowledge base;

[0101] For forms, receipts and other formatted text or digital knowledge information provided in the form of images, it is difficult for traditional visual encoders and image content description tools to accurately recognize the text or digital information therein. Therefore, this embodiment uses an optical character recognition tool f OCR Extract the optical character information in the image modality knowledge data. For an image I located at position P in the knowledge base, the optical character information of the image is represented by OCR=f OCR (I), if the image does not contain optical character information, the OCR is empty;

[0102] Step 3.1.2: Based on the optical character recognition information, position information and context information of the image, a large model with multimodal processing capabilities is used to identify the semantic information contained in the optical character information of each image;

[0103] In order to identify the semantic information in the image, such as the general content of the image, a large model with multimodal processing capabilities is used, and the optical character recognition information, position information and context information of the image are provided to identify the semantic information in the image modal knowledge data; for the position P where the image I is located, the content above it at a distance of δ1 lines is defined as C 1 (δ1), the next content δ2 rows away is C 2 (δ2), the optical character information of the image is finally recognized to contain semantic information F x for:

[0104] F=f LLM (I; C1 (δ1);OCR;C 2 (δ2)) (1)

[0105] Step 3.2: Replace the image corresponding to each position in the multimodal knowledge base with the optical character information of the image containing semantic information, and convert the multimodal knowledge base into a plain text knowledge base;

[0106] Step 4: Receive a complex question given by the user, decompose the complex question based on the entity recognition model, and use the dependency syntax analysis model to obtain the entity information in the complex question;

[0107] A complex question is defined as a question consisting of two or more entities, or a question composed of two or more sub-questions. Complex questions cannot be answered directly, but need to be answered based on the reasoning path formed by the entities involved in the complex question and the relationship between multiple entities. If the reasoning path formed is unclear, or the basic model used does not have the relevant background knowledge of the entity, it may not be possible to answer the question correctly or comprehensively. For example, for questions consisting of a single entity such as "Who is A", it can be regarded as a simple question; for questions containing multiple entities such as "The date on which A obtains B", it is necessary not only to know who A is, but also what B is, and the meaning of the entity of date in this sentence. Therefore, this type of question containing multiple entities is a complex question;

[0108] In this embodiment, the specific named entity recognition model used is not limited. NER (hereinafter referred to as entity recognition model), any existing model can be selected, such as the entity recognition model based on BERT; in addition, according to the named entity recognition dataset D NER You can choose whether to further train the selected entity recognition model based on the specific business scenario. In this embodiment, the training method used for the entity recognition model is not limited;

[0109] Dependency parsing modelf DSP (hereinafter referred to as the syntactic analysis model) decomposes the complex question text input by the user into a syntactic tree with a dependency structure including noun subject, direct object, indirect object, adjective clause, and copula according to the dependency syntax.

[0110] Step 4.1: Based on the entity recognition model f NER Obtaining the set of potential entities in complex problems As shown in the following formula:

[0111]

[0112] Among them, Q is a complex problem, {e1,e2,…,e E} are the E potential entities identified from the complex question Q;

[0113] Step 4.2: Parse the model f through syntax DSP Obtaining key entity sets in complex problems And update the potential entity set to obtain the updated potential entity set

[0114] Key entities are defined as all types of nouns obtained after dependency syntactic analysis. These nouns act as noun subjects, direct objects, indirect objects, and various types of clauses in sentences, playing a key role in syntax. Since these nouns are also part of entities, in order to avoid differences between potential entities and key entities, after obtaining the potential entity set, the potential entity set is updated as shown in the following formula:

[0115]

[0116] Among them, {s1,s2,…,s N} is the syntax analysis model f DSP The N key entities in the complex problem obtained, {c1,c2,…,c m} are the m key potential entities in the updated potential entity set;

[0117] Step 4.3: Based on dependency syntactic structure and key entity set A set of related entities is formed by direct, indirect or clause-based associations between nouns. Among them (s i ,s j ) is the associated entity pair consisting of the i-th associated entity and the j-th associated entity, (s i ,s j ) z is the zth associated entity pair in the associated entity set, and n is the number of associated entity pairs contained in the complex question Q;

[0118] The associated entity refers to an entity pair with a formal association between subject, object and clause in the dependency syntactic tree. Specifically, the noun subject is associated with its direct object, the direct object is associated with its indirect object, and any noun is associated with its subsequent clause. If a noun is not associated with any other noun, it is treated as an exception and associated with any other noun alone.

[0119] Step 5: According to the entity information in the complex problem, the complex problem is converted into multiple sub-problems, and the direct knowledge and related knowledge of each sub-problem are retrieved from the plain text knowledge base;

[0120] In this embodiment, the search tools and search methods used are not limited, and the search tools that match the plain text knowledge base can be used. retrieval For example;

[0121] Step 5.1: Update the potential entity set The key potential entities in the direct question template T sub Generate a set of direct subproblems Q sub ;

[0122] Direct question templates refer to simple questions that can be directly retrieved and answered after each entity is filled in according to the template; in this embodiment, the question "What / who is c y ?” as a direct question template to generate direct sub-questions corresponding to key potential entities. For example, for the entity “A”, a direct sub-question of the form “What / who is A?” is generated;

[0123] Based on direct question template T sub , generate a set of direct sub-problems Q sub As shown below:

[0124]

[0125] Among them, q y The direct sub-question set Q sub The direct sub-problem in c y is a key potential entity; for any key potential entity c y , after filling in the direct question template, generate its corresponding direct sub-question q y , until the updated potential entity set All the key potential entities in are filled;

[0126] Step 5.2: Associating entity collections Based on the associated question template T rel Generate a set of associated sub-problems Q rel ;

[0127] The associated question template refers to a question that asks what kind of relationship exists between two entities in each element contained in the associated entity set; in this embodiment, "s i and j "What is the relationship between A and B?" is the association question template. For the association entity pair (A, B), an association sub-question of the form "What is the relationship between A and B?" is generated;

[0128] Based on the associated question template T rel , associated sub-problem set Q rel As shown below:

[0129]

[0130] Among them, q i,j is the set of associated sub-problems Q rel The associated subproblem in (q i,j ) w is the set of associated sub-problems Q rel The w-th associated subproblem in ;

[0131] Step 5.3: Retrieve the set of direct sub-questions Q in the plain text knowledge base sub Each direct subproblem q a , obtain the information about each direct sub-problem q y Direct knowledge of K sub ;

[0132] For the direct sub-problem q y , using a search tool that matches a plain text knowledge base retrieval Retrieve relevant knowledge from the plain text knowledge base. If direct knowledge is not retrieved, the basic model will answer the question separately as reference knowledge and generate direct knowledge K sub If the retrieved direct knowledge is one or more than one, all the retrieved direct knowledge are sorted according to the relevance scores of each direct knowledge according to the basic model, and each direct knowledge is spliced ​​from high to low in terms of relevance scores to generate direct knowledge K sub :

[0133] K sub =f retrieval (Q sub )={k u |u=1,2,…,U} (6)

[0134] Among them, k u For the direct subproblem q y The u-th direct knowledge, U is the number of direct knowledge, f retrieval To use search tools to search in plain text knowledge base and return the retrieved knowledge;

[0135] Step 5.4: Retrieve the collection of related sub-questions Q in the plain text knowledge base rel Each associated subproblem q i,j , get the associated sub-problem q i,j The associated knowledge K rel ;

[0136] For the associated sub-problem q i,j , retrieve its related knowledge from the plain text knowledge base. If no related knowledge is retrieved, concatenate the “maybe there is no relationship” and the answer directly answered by the basic model as reference knowledge and generate related knowledge Krel If the retrieved related knowledge is one or more, all the retrieved related knowledge are sorted according to the relevance scores of each related knowledge according to the basic model, and each related knowledge is spliced ​​from high to low in terms of relevance scores to generate related knowledge K rel :

[0137] K rel =f retrieval (Q rel )={(k i,j ) v |v=1,2,…,V} (7)

[0138] Among them, (k i,j ) v For the associated subproblem q i,j The vth associated knowledge, V is the number of associated knowledge;

[0139] Step 6: Embedding model f based on knowledge graph KGE Construct a knowledge graph G for complex problems, sub-problem sets, direct knowledge related to sub-problems, and associated knowledge;

[0140] Knowledge Graph Embedding (KGE) refers to a type of model that captures the semantic relationship between entities by learning the vector representation between entities and relationships, thereby helping knowledge graph modeling and reasoning, including but not limited to TransE, RotateE, and ConvE. In this embodiment, the knowledge graph embedding model used is not limited. By establishing a knowledge graph with semantic relationships, it helps large models better complete reasoning on complex problems;

[0141] Step 6.1: Define the set R of binary logical relations between all key entities in the complex problem;

[0142] The relationship type set mainly refers to the types of semantic relationships that may exist between entities in the complex questions raised by users, and even the relationships between relationships. They are generally represented by binary logical relationships in discrete mathematics, including symmetric relationships, antisymmetric relationships, inverse relationships, transitive relationships, and combination relationships, etc. By establishing a knowledge graph with entity semantics as nodes and logical relationships as nodes, with the help of discrete mathematics theories, it is possible to help large models establish reasoning paths for complex problems, thereby answering complex problems more accurately and comprehensively. In this embodiment, the binary logical relationship set R between key entities in complex problems includes three types: symmetric relationships, antisymmetric relationships, and no relationships;

[0143] For the potential entity set Any two entities e l and eh , the symmetric relationship is shown as follows:

[0144]

[0145] Among them, r is the relationship between two entities. l and e h Under the premise that there is a relationship r, e can be directly derived h and e l If there is also a relationship r between them, then there is a symmetrical relationship between the two entities, such as classmates, friends, lovers, etc.

[0146] The antisymmetric relationship is shown below:

[0147]

[0148] In contrast to the symmetric relationship, knowing the entity e l and e h Under the premise that there is a relationship r, it can be deduced that e h and e l If there is no relationship r between them, then the two entities have an antisymmetric relationship, such as a parent-child relationship;

[0149] No relationship, that is, there is no semantic relationship between the two entity pairs;

[0150] Step 6.2: Use the knowledge graph to embed the model f KGE The encoder f emb , the complex problem, the direct knowledge corresponding to the direct sub-problems obtained by decomposing the complex problem, and the associated knowledge corresponding to the associated sub-problems obtained by decomposing the complex problem are all converted into the knowledge graph embedding model f KGE The required embedding vector;

[0151] Using the knowledge graph embedding model f KGE The encoder f emb , get the embedding vector E of the complex question Q Q =f emb (Q);

[0152] Each direct subproblem q y and several pieces of direct knowledge, so as to integrate the entity and its background knowledge, and use the knowledge graph embedding model f KGE The encoder f emb , get the embedding vector E of direct knowledge y =f emb (q y ;k u );

[0153] Each associated subproblem qi,j and several pieces of associated knowledge are spliced ​​together to fuse the background knowledge of associated entity pairs and the possible relationships between them, and use the knowledge graph embedding model f KGE The encoder f emb , get the embedding vector E of the associated knowledge i,j =f emb (q i,j ;(k i,j ) v );

[0154] Step 6.3: Input the embedding vector of complex question Q, the embedding vector of direct knowledge, and the embedding vector of associated knowledge into the knowledge graph embedding model f KGE , get the knowledge graph G corresponding to the decomposition of the complex problem;

[0155] Using the knowledge graph embedding model f KGE The resulting set of potential entities Any entity e in l and e h The logical relationship between them is:

[0156] r l.h =f KGE (E Q ,E l ,E h ,E l,h ),r l,h ∈R (10)

[0157] Further obtain the potential entity set The logical relationship in l,h Any entity e l and e h The knowledge graph G is:

[0158] G={{(e l ,k l ),r l.h ,(e h ,k h )} g |g=1,2,…,G} (11)

[0159] Step 7: Integrate the elements contained in the knowledge graph G into a new question Q * , and the trained basic model f LLM Generate new question Q * Complete and generate answer A;

[0160] Step 7.1: Determine the entity-direct knowledge pairs (e l ,k l) exceeds the maximum context length specified by the basic model, and summarizes and abbreviates the direct knowledge that exceeds the maximum context length specified by the basic model so that the overall length of the direct knowledge does not exceed the maximum context length Θ of the basic model;

[0161] Define the maximum context length Θ that the base model can handle, and define each entity-direct knowledge pair (e l ,k l )'s maximum length θ l , if an entity-direct knowledge pair (e l ,k l ) exceeds its specified maximum length θ l , then the basic model f LLM Direct knowledge of l Summarize and abbreviate until the length limit is met, so the maximum context length θ that the base model can handle is:

[0162]

[0163] The maximum context length that the basic model can handle generally depends on the selected basic model itself. In this embodiment, for simple calculation, the maximum length of each entity-knowledge pair is set to be equal, and the maximum context length that the basic model can handle is equally divided into the number of elements in the knowledge graph;

[0164] Step 7.2: Fill the elements contained in the original complex question Q and the knowledge graph G into the new question template T in sequence new Generate new question Q * , and the trained basic model f LLM According to the new question Q * Generate response A:

[0165]

[0166] In this embodiment, a new question template T is defined new Please answer the question based on the given background knowledge: Q; Related background knowledge: (e l ,k l ), (e h ,k h ), e l and e h The logical relationship between them is r l,h "; Fill each element contained in the knowledge graph G into the new question template T in turn new After that, a new question Q is formed * , based on the trained basic model f LLM The generated response A is the new question Q *The answer to , which is also the answer to the original question Q.

[0167] In this way, not only the relevant background knowledge is integrated and the retrieval ability of the retrieval tool for multimodal knowledge base is expanded, but also by decomposing complex problems, the semantic and structural relationships between the many sub-problems that may be contained in the complex problems are constructed, thereby helping the large model to answer complex problems more accurately and comprehensively.

[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A complex problem decomposition and multimodal knowledge retrieval method combining a large model, characterized by: The following steps are involved: Step 1: Select the large model to be used as the basic model; Step 2: According to the specific business scenario requirements, build a named entity recognition dataset for training the selected basic model to obtain a trained basic model; Step 3: Establish a multimodal knowledge base, preprocess it, and convert the multimodal knowledge base into a plain text knowledge base; Step 4: Receive a complex question given by the user, decompose the complex question based on the entity recognition model, and use the dependency syntax analysis model to obtain the entity information in the complex question; Step 5: According to the entity information in the complex problem, the complex problem is converted into multiple sub-problems, and the direct knowledge and related knowledge of each sub-problem are retrieved from the plain text knowledge base; Step 6: Based on the knowledge graph embedding model, construct a knowledge graph for complex problems, sub-problem sets, direct knowledge related to sub-problems, and associated knowledge; Step 7: Integrate the elements contained in the knowledge graph into new questions, and use the trained basic model to generate new questions and generate answers.

2. The complex problem decomposition and multimodal knowledge retrieval method combining a large model according to claim 1, characterized in that: Step 2 describes the construction of a named entity recognition dataset: Constructing named entity recognition dataset D NER ={(S p ; {e p1 ,e p2 ,…,e pd })|p=1,2,…,D} is used to train the basic model, where S p is the pth text in a data set with a sample size of D, {e p1 ,e p2 ,…,e pd } is the named entity recognition dataset D NER The d entities contained in the p-th text.

3. The complex problem decomposition and multimodal knowledge retrieval method combined with a large model according to claim 2, characterized in that: The step 3 comprises: Extracting semantic information contained in the image modality knowledge data in the multimodal knowledge base, converting the image modality knowledge data into text modality knowledge data, and then converting the text modality knowledge data into a text modality knowledge vector; The multimodal knowledge base refers to a knowledge base containing text-image modality knowledge data, and the data in the knowledge base is not limited to be multimodal data. The modality of the data depends on the actual business and usage scenarios. If a given knowledge base contains image modality knowledge data, the image modality data needs to be preprocessed; The multimodal knowledge base includes a system built-in knowledge base and a user-provided knowledge base. The system built-in knowledge base refers to an existing professional knowledge base, including but not limited to a knowledge graph, a graph database, and a vector query database; the user-provided knowledge base refers to a multimodal context provided by a user to the big model when the user uses the big model, which is helpful for answering questions, that is, general text or documents provided by the user; When preprocessing the multimodal data contained in the multimodal knowledge base, the system's built-in knowledge base only needs to be processed offline once, and the data contained in the system's built-in knowledge base can be directly retrieved and used after preprocessing operations; for the knowledge base given by the user, real-time processing is required; if the user does not provide a knowledge base when asking a question, the real-time processing of the user-given knowledge base is skipped, and only the system's built-in knowledge base is preprocessed; conversely, if the user provides a knowledge base when asking a question, the preprocessing of the system's built-in knowledge base is skipped, and only the user-provided knowledge base is processed in real time.

4. The complex problem decomposition and multimodal knowledge retrieval method combined with a large model according to claim 3, characterized in that: The specific method of step 3 is: Step 3.1: Use the optical character recognition tool OCR Extract the optical character information of each image contained in all image modal knowledge data in the multimodal knowledge base, and use a large model f with multimodal processing capabilities MLLM Identify semantic information contained in the optical character information of each image; Step 3.1.1: Use optical character recognition tools OCR Extracting optical character information of each image contained in all image modality knowledge data in a multimodal knowledge base; Using optical character recognition tools OCR Extract the optical character information in the image modality knowledge data. For the image I located at position P in the multimodal knowledge base, the optical character information of the image is represented by OCR=f OCR (I), if the image does not contain optical character information, the OCR is empty; Step 3.1.2: Based on the image’s optical character recognition information, location information, and context information, use a large model with multimodal processing capabilities f MLLM Identify semantic information contained in the optical character information of each image; For the position P of image I, the content above the position δ1 row away is defined as C 1 (δ1), the next content δ2 rows away is C 2 (δ2), the optical character information of the image is finally recognized to contain the semantic information F: F=f MLLM (I;C 1 (δ1);OCR;C 2 (δ2)) (1) Step 3.2: Replace the image corresponding to each position in the multimodal knowledge base with the optical character information of the image containing semantic information, and convert the multimodal knowledge base into a plain text knowledge base.

5. The complex problem decomposition and multimodal knowledge retrieval method combining a large model according to claim 4, characterized in that: The complex problem described in step 4 is a problem composed of two or more entities, or a problem composed of two or more sub-problems; The specific method of step 4 is: Step 4.1: Based on the entity recognition model f NER Obtaining the set of potential entities in complex problems As shown in the following formula: Among them, Q is a complex problem, {e1,e2,…,e E } are the E potential entities identified from the complex question Q; Step 4.2: Parse the model f through syntax DSP Obtaining key entity sets in complex problems And update the potential entity set to obtain the updated potential entity set Key entities are defined as all types of nouns obtained after dependency syntactic analysis; Key Entity Collection And the updated potential entity set is shown in the following formula: Among them, {s1,s2,…,s N } is the syntax analysis model f DSP The N key entities in the complex problem obtained, {c1,c2,…,c m } are the m key potential entities in the updated potential entity set; Step 4.3: Based on dependency syntactic structure and key entity set A set of related entities is formed by direct, indirect or clause-based associations between nouns. Among them (s i ,s j ) is the associated entity pair consisting of the i-th associated entity and the j-th associated entity, (s i ,s j ) z is the zth associated entity pair in the associated entity set, and n is the number of associated entity pairs contained in the complex question Q.

6. The complex problem decomposition and multimodal knowledge retrieval method combined with a large model according to claim 5, characterized in that: The step 5 comprises the following steps: Step 5.1: Update the potential entity set The key potential entities in the direct question template T sub Generate a set of direct sub-problems Q sub ; Based on direct question template T sub , generate a set of direct sub-problems Q sub As shown below: Among them, q y The direct sub-problem set Q sub The direct sub-problem in c y is a key potential entity; for any key potential entity c y , after filling in the direct question template, generate its corresponding direct sub-question q y , until the updated potential entity set All the key potential entities in are filled; Step 5.2: Associating entity collections Based on the associated question template T rel Generate a set of associated sub-problems Q rel ; Based on the associated question template T rel , generate a set of associated sub-problems Q rel As shown below: Among them, q i,j is the set of associated sub-problems Q rel The associated subproblem in (q i,j ) w is the set of associated sub-problems Q rel The w-th associated subproblem in ; Step 5.3: Retrieve the set of direct sub-questions Q in the plain text knowledge base sub Each direct subproblem q y , obtain the information about each direct sub-problem q y Direct knowledge of K sub ; For the direct sub-problem q y , using a search tool that matches a plain text knowledge base retrieval Retrieve relevant knowledge from the plain text knowledge base. If direct knowledge is not retrieved, the basic model will answer the question separately as reference knowledge and generate direct knowledge K sub If the retrieved direct knowledge is one or more than one, all the retrieved direct knowledge are sorted according to the relevance scores of each direct knowledge according to the basic model, and each direct knowledge is spliced ​​from high to low in terms of relevance scores to generate direct knowledge K sub : K sub =f retrieval (Q sub )={k u |u=1,2,…,U} (6) Among them, k u For the direct subproblem q y The u-th direct knowledge, U is the number of direct knowledge, f retrieval To use search tools to search in plain text knowledge base and return the retrieved knowledge; Step 5.4: Retrieve the collection of related sub-questions Q in the plain text knowledge base rel Each associated subproblem q i,j , get the associated sub-problem q i,j The associated knowledge K rel ; For the associated sub-problem q i,j , retrieve its related knowledge from the plain text knowledge base. If no related knowledge is retrieved, concatenate the "maybe there is no relationship" and the answer directly answered by the basic model as reference knowledge and generate related knowledge K rel If the retrieved related knowledge is one or more, all the retrieved related knowledge are sorted according to the relevance scores of each related knowledge according to the basic model, and each related knowledge is spliced ​​from high to low in terms of relevance scores to generate related knowledge K rel : K rel =f retrieval (Q rel )={(k i,j ) v |v=1,2,…,V} (7) Among them, (k i,j ) v For the associated subproblem q i,j The vth associated knowledge, V is the number of associated knowledge.

7. The complex problem decomposition and multimodal knowledge retrieval method combined with a large model according to claim 6, characterized in that: The specific method of step 6 is: Step 6.1: Define the set R of binary logical relations between all key entities in the complex problem; The binary logical relationship set R between the key entities in the complex problem includes three types: symmetric relationship, antisymmetric relationship and no relationship; For the potential entity set Any two entities e l and e h , the symmetric relationship is shown as follows: Among them, r is the relationship between two entities. l and e h Under the premise that there is a relationship r, e can be directly derived h and e l If there is also a relationship r between them, then there is a symmetric relationship between the two entities; The antisymmetric relationship is shown as follows: In contrast to the symmetric relationship, knowing the entity e l and e h Under the premise that there is a relationship r, it can be deduced that e h and e l If there is no relationship r between them, then there is an antisymmetric relationship between the two entities; No relationship, that is, there is no semantic relationship between the two entity pairs; Step 6.2: Use the knowledge graph to embed the model f KGE The encoder f emb , the complex problem, the direct knowledge corresponding to the direct sub-problems obtained by decomposing the complex problem, and the associated knowledge corresponding to the associated sub-problems obtained by decomposing the complex problem are all converted into the knowledge graph embedding model f KGE The required embedding vector; Using the knowledge graph embedding model f KGE The encoder f emb , get the embedding vector E of the complex question Q Q =f emb (Q); Each direct subproblem q y and several pieces of direct knowledge, so as to integrate the entity and its background knowledge, and use the knowledge graph embedding model f KGE The encoder f emb , get the embedding vector E of direct knowledge y =f emb (q y ;k u ); Each associated subproblem q i,j and several pieces of associated knowledge are spliced ​​together to fuse the background knowledge of associated entity pairs and the possible relationships between them, and use the knowledge graph embedding model f KGE The encoder f emb , get the embedding vector E of the associated knowledge i,j =f emb (q i,j ;(k i,j ) v ); Step 6.3: Input the embedding vector of complex question Q, the embedding vector of direct knowledge, and the embedding vector of associated knowledge into the knowledge graph embedding model f KGE , get the knowledge graph G corresponding to the decomposition of the complex problem; Using the knowledge graph embedding model f KGE The resulting set of potential entities Any entity e in l and e h The logical relationship between them is: r l.h =f KGE (E Q ,E l ,E h ,E l,h ),r l,h ∈R (10) Further obtain the potential entity set The logical relationship in l,h Any entity e l and e h The knowledge graph G is: G={{(e l ,k l ),r l.h ,(e h ,k h )} g |g=1,2,…,G} (11)。 8. The complex problem decomposition and multimodal knowledge retrieval method combining a large model according to claim 7, characterized in that: The specific method of step 7 is: Step 7.1: Determine the entity-direct knowledge pairs (e l ,k l ) exceeds the maximum context length specified by the basic model, and summarizes and abbreviates the direct knowledge that exceeds the maximum context length specified by the basic model so that the overall length of the direct knowledge does not exceed the maximum context length Θ of the basic model; Define the maximum context length Θ that the base model can handle, and define each entity-direct knowledge pair (e l ,k l )'s maximum length θ l , if an entity-direct knowledge pair (e l ,k l ) exceeds its specified maximum length θ l , then the basic model f LLM Direct knowledge of l Summarize and abbreviate until the length limit is met, so the maximum context length θ that the base model can handle is: Step 7.2: Fill the elements contained in the original complex question Q and the knowledge graph G into the new question template T in sequence new Generate new question Q * , and the trained basic model f LLM According to the new question Q * Generate response A: Define a new question template T new , fill each element contained in the knowledge graph G into the new question template T in turn new After that, a new question Q is formed * , based on the trained basic model f LLM The generated response A is the new question Q * The answer to , which is also the answer to the original question Q.

Citation Information

Patent Citations

  • Large language model reasoning method and system based on multi-level knowledge retrieval enhancement

    CN118036753A

  • Knowledge retrieval enhancement-based large language model question and answer method and device

    CN118113836A

  • Intelligent question answering method and system based on domain knowledge graph

    CN117648984A

  • Visual question and answer method, system and device based on thinking chain and storage medium

    CN117891965A

  • Question and answer type retrieval method and system based on large language model

    CN118051590A

Cited By

  • Knowledge reasoning method and device for integrated circuit wafer process and medium

    CN120744140A

  • A knowledge reasoning method, device and medium for an integrated circuit wafer process

    CN120744140B