Construction method of low-illusion multi-Agent picture question-answering system based on structured understanding and medium
By constructing a low-idiot multi-agent picture question and answer system based on structured understanding, the inefficiency and hallucination problems of multimodal large language model in the judgment of image elements relationships is solved, high-quality picture content understanding and question and answer are achieved, and application reliability in professional fields is enhanced.
Patent Information
- Application Number
- CN202510292748.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-04
AI Technical Summary
The multimodal large language model is inefficient in judging the relationship between picture elements and numerical elements, resulting in serious hallucinations in professional applications and may cause accidents.
A low-illusion multi-Agent image question and answer system based on structured understanding is built. By constructing a joint training of training data sets and four agents, including image classification, knowledge graph generation, data table generation and Text2SQL Agent, it uses Lora's fine-tuned multimodal large language model to enhance the ability to understand image content.
It alleviates the illusion problem of multimodal large language model, realizes high-quality picture content understanding and question-and-answer, and improves application reliability in professional fields.
Smart Images

Figure CN120256561A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large language models, and specifically, it relates to a construction method and medium of a low-hallucination multi-Agent picture question-answering system based on structured understanding. Background Art
[0002] With the rise of multi-modal large language models, it has become possible to transfer the capabilities of large language models to the visual field. More and more multi-modal large language models, such as GPT-4o, cogVLM2, mPLUG-Owl, etc., have demonstrated powerful picture question-answering capabilities.
[0003] However, due to the inefficiency of multi-modal large language models in judging the relationships between picture elements and numerical elements, serious hallucinations occur in data-type picture question-answering, which may lead to serious accidents in applications in professional fields such as finance and medicine; the inventors of the present application found that the reason for the serious hallucinations is that the large language model lacks the ability to understand picture content. On the premise of understanding low-quality picture content and conducting picture content question-answering, the hallucination problem of the large model will inevitably occur. Summary of the Invention
[0004] The purpose of the present invention is to provide a construction method of a low-hallucination multi-Agent picture question-answering system based on structured understanding to solve the technical problems existing in the prior art.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows: A construction method of a low-hallucination multi-Agent picture question-answering system based on structured understanding, including: Step S1: Construction of a training data set for picture structured understanding: For text form pictures, data tables, picture classification, and text2sql data, training data sets are constructed respectively; Step S2: Joint training of four Agents: According to the four Agents, corresponding prompt data is constructed, and the prompt data is mixed together for training the original multi-modal large language model to obtain the trained multi-modal large language model; Step S3: Based on the trained multi-modal large language model, a multi-Agent system is constructed, and the multi-Agent system includes: Picture classification Agent: When the user inputs a multi-modal question, the picture is classified into types: form pictures, table pictures, and other pictures, where the multi-modal question includes pictures and text questions; Knowledge graph generation Agent: Identifies form pictures, generates a knowledge graph of the picture content, and stores it in a semi-structured database; Data table generation Agent: Recognize tabular images, generate table content of the image content, and store it in a relational database; General large language model: Process multi-modal questions and answers for other images and directly return the answer results; at the same time, be responsible for converting SQL data into natural language; Text2SQL Agent: For text questions, generate query statements, and query the above databases based on the semi-structured database and relational database to obtain query data; Output module: Convert the query data into natural language and output it to the user.
[0006] In one implementation, the specific method of step S1 is as follows: For text form image data, use the OCR model to extract text data, and use the semantic relationship judgment model in NLP to generate a knowledge graph of the text data; For data table type data, use the API interface of office Excel for reverse construction. Based on a known data table, generate a data graph through the drawing API of office Excel; For image classification and text2sql data, use a multi-modal large model to generate the image category and the corresponding SQL data for the question.
[0007] In one implementation, in step S2, the construction method of the prompt data is as follows; (2.1) For the image classification Agent, construct the following prompt data: Question template: Please classify the following image; Image information: base64 string of the image; Agent return result: image category, form image, data graph or other images; (2.2) For the knowledge graph generation Agent, construct the following prompt data: Question template: Please convert the following image into a knowledge graph; Image information: base64 string of the image; Agent return result: knowledge graph of the image text content, represented in json structure; (2.3) For the data table generation Agent, construct the following prompt data: Question template: Please convert the following image into a data table; Image information: base64 string of the image; Agent return result: table of the image text content, represented in markdown structure; (2.4) For the SQL generation Agent, use the following prompt data: Question template: Given the data table structure: (insert knowledge graph or data table structure), the user's input question is: (insert user question). Please generate a MySQL query statement? Agent returns the result: query statement.
[0008] In one embodiment, in step S2, during the training process, LoRA fine-tuning is used to fine-tune the original multi-modal large language model to obtain the trained multi-modal large language model.
[0009] In one embodiment, the specific method of the fine-tuning is as follows: During the training process, based on the prompt data, low-rank adaptation training is performed on the original multi-modal large language model. Side branch matrices A and B are added to both ends of the original multi-modal large language model for slight adjustment of the original weights of the model. The calculation method is:
[0010] Where, is the weight of a certain layer of the original model, is the adjusted weight, The shape of is n, the dimension of the side branch matrix B is The dimension of the side branch matrix A is m < n.
[0011] To achieve the above object, the present invention also provides a computer-readable storage medium, on which a computer program is stored. The computer program is executed by a processor to implement the construction method of the low-hallucination multi-Agent picture question-answering system based on structured understanding as described above.
[0012] Compared with the prior art, the present invention has the following beneficial effects: (1) According to the present invention, through different data construction schemes, the training data construction problems of the four Agents of picture classification, knowledge graph generation, data table generation, and text2sql are solved; (2) According to the present invention, through the LoRA fine-tuning scheme trained by specific prompts, the joint training problem of the four Agents is solved; (3) According to the present invention, a highly structured picture understanding chain is built, which can realize low-hallucination picture question-answering. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a schematic diagram of the principle of Embodiment 1 of the present invention. DETAILED DESCRIPTION
[0014] To enable those skilled in the art to have a clearer understanding and knowledge of the present invention, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described below are only used to explain the present invention for easy understanding, and the technical solutions provided by the present invention are not limited to the technical solutions provided by the following embodiments, nor should the technical solutions provided by the embodiments limit the protection scope of the present invention.
[0015] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The form, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the layout form of its components may also be more complex.
[0016] Embodiment 1 As Figure 1 shown, this embodiment provides a construction method for a low-hallucination multi-Agent picture Q&A system based on structured understanding, aiming to enhance the picture content understanding ability of the multi-modal large language model and alleviate the hallucination problem of the multi-modal large language model. This technical solution mainly includes three aspects: I. Construction of the training dataset for picture structured understanding For text form picture data, data table type data, picture classification, and text2sql data, training datasets are constructed respectively as follows: For text form picture data, use the OCR model to extract the text data, and use the semantic relationship judgment model in NLP to generate the knowledge graph of the text data; For data table type data, use the API interface of office Excel for reverse construction. Based on a known data table, generate a data graph through the drawing API of office Excel; For picture classification and text2sql data, use a multi-modal large model to generate picture categories and the corresponding SQL data for the questions.
[0017] II. Joint training of multiple Agents In this step, the joint training of four Agents is adopted: according to the four Agents, construct the corresponding prompt data, mix the prompt data together for the training of the original multi-modal large language model, and obtain the trained multi-modal large language model.
[0018] In this step, the construction method of the prompt data is as follows; (2.1) For the picture classification Agent, construct the following prompt data: Question template: Please classify the following pictures; Picture information: Base64 string of the picture; Agent return result: Picture category, form picture, data graph or other pictures; For example: { "prompt": "Please classify the following pictures", "image_url": "data:image / jpeg;base64,iVBORw0KGgoAAAANSUhEUgAAAzAAAAQgCAIAAAC2Gy5ZAAAACXBIWXMAAA7EAAAOxAG...", "response": "Form picture / Data graph / Other pictures" } (2.2) Generate an Agent for the knowledge graph and construct the following prompt data: Question template: Please convert the following picture into a knowledge graph; Picture information: Base64 string of the picture; Agent return result: Knowledge graph of the picture text content, represented in JSON structure; For example: { "prompt": "Please convert the following picture into a knowledge graph", "image_url": "data:image / jpeg;base64,iVBORw0KGgoAAAANSUhEUgAAAzAAAAQgCAIAAAC2Gy5ZAAAACXBIWXMAAA7EAAAOxAG...", # Base64 string of the form picture "response": '{\n "Name": "Zhang San",\n "Gender": "Male",\n "Occupation": "Lawyer"\n}' } (2.3) Generate an Agent for the data table and construct the following prompt data: Question template: Please convert the following picture into a data table; Picture information: Base64 string of the picture; Agent return result: Table of the picture text content, represented in Markdown structure; For example: { "prmopt": "Please convert the following picture into a data table", "image_url": "data:image / jpeg;base64,iVBORw0KGgoAAAANSUhEUgAAAzAAAAQgCAIAAAC2Gy5ZAAAACXBIWXMAAA7EAAAOxAG...", "response": "| Month | Company Turnover Growth Rate | Company Net Profit Growth Rate | \n |---- | ---- | ---- | \n | January | 2% | 3% | \n | February | 3% | 4% | " } (2.4) Generate an Agent for SQL and use the following prompt data: Question template: Given the data table structure: (insert knowledge graph or data table structure), the user's input question is: (insert user question): Please generate a mysql query statement? The Agent returns the result: the query statement, where "()" represents a placeholder; For example: { "prompt": "Given the data table structure is: | Month | Company Turnover Growth Rate | Company Net Profit Growth Rate | \n | ---- | ---- | ---- | \n | January | 2% | 3% |, the data table name is data_table, the user input is: What is the company's business growth rate in January?, Please generate a mysql query statement", "response": " SELECT Company Turnover Growth Rate FROM data_table WHERE Month = 'January'" } Mix the prompts for the above four tasks together and use LoRA for fine-tuning to fine-tune the original multi-modal large language model. The specific method of fine-tuning is as follows: During the training process, based on the prompt data, perform low-rank adaptation training on the original multi-modal large language model, add side matrices A and B to both ends of the original multi-modal large language model for mild adjustment of the original weights of the model. The calculation method is:
[0019] Among them, is the weight of a certain layer of the original model, is the adjusted weight, The shape of is n, the dimension of the side matrix B is The dimension of the side matrix A is , m < n; among them, the matrix sizes of the side matrix A and the side matrix B are also much smaller than , making the number of trainable parameters small during the training process.
[0020] III. Multi-Agent System Construction Based on the above four types of Agents, a multi-Agent system is built for picture structured understanding and high-quality question answering. Specifically, the multi-Agent system includes the following: Picture Classification Agent: When the user inputs a multi-modal question, classify the picture into types: form pictures, table pictures, and other pictures. Among them, the multi-modal question includes pictures and text questions; Knowledge Graph Generation Agent: Identify form pictures, generate a knowledge graph of the picture content, and store it in a semi-structured database; Data Table Generation Agent: Identify table pictures, generate the table content of the picture content, and store it in a relational database; General Large Language Model: Process multi-modal question answering for other pictures and directly return the answer results; at the same time, be responsible for converting SQL data into natural language; Text2SQL Agent: For text questions, generate query statements, and query the above databases based on the semi-structured database and the relational database to obtain query data; Text2Sql Agent is fine-tuned on the basis of the general large language model using a large amount of open-source Text2SQL data to enhance its SQL generation ability. For example Figure 1 in, the user inputs a question related to the chart: "What is the company's revenue growth rate in January?", the TextSQL Agent will convert it into a query statement for the relational database "SELECT company turnover growth rate FROM data_table WHERE month = 'January'", and then send it to the database for query execution to obtain the query data "| turnover growth rate | | ---- | | 2% |" Output Module: Convert the query data into natural language and output it to the user. Combining the above example, the general large language model will continue to convert the query data into "The revenue growth rate is 2%" and output it to the user.
[0021] Embodiment 2 This embodiment provides a computer-readable storage medium with a computer program stored thereon. The computer program is executed by a processor to implement the construction method of the low-hallucination multi-Agent picture question answering system based on structured understanding provided in Embodiment 1. Those of ordinary skill in the art can understand that all or part of the steps to implement the method provided in Embodiment 1 can be completed by hardware related to the computer program. The above computer program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the method provided in Embodiment 1; and the above storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0022] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A construction method of a low-hallucination multi-Agent picture question answering system based on structured understanding, characterized in that, Including: Step S1: Construction of the training dataset for picture structured understanding: For text form pictures, data table pictures, picture classification, and text2sql data, construct the training datasets respectively; Step S2: Joint training of four Agents: According to the four Agents, construct the corresponding prompt data, mix the prompt data together for the training of the original multi-modal large language model, and obtain the trained multi-modal large language model; Step S3: Based on the trained multi-modal large language model, construct a multi-Agent system, and the multi-Agent system includes: Picture classification Agent: When the user inputs a multi-modal question, classify the picture into types: form picture, table picture, other pictures, where the multi-modal question includes pictures and text questions; Knowledge graph generation Agent: Recognize the form picture, generate the knowledge graph of the picture content, and store it in the semi-structured database; Data table generation Agent: Recognize the table picture, generate the table content of the picture content, and store it in the relational database; General large language model: Process the multi-modal question and answer of other pictures and directly return the answer result; at the same time, be responsible for converting SQL data into natural language; Text2SQL Agent: For text questions, generate query statements, and query the above databases based on the semi-structured database and the relational database to obtain query data; Output module: Convert the query data into natural language and output it to the user.
2. The construction method of the low-hallucination multi-Agent picture question answering system based on structured understanding according to claim 1, characterized in that The specific method of the step S1 is as follows: For text form picture data, use the OCR model to extract the text data, and use the semantic relationship judgment model in NLP to generate the knowledge graph of the text data; For data table data, use the API interface of office Excel for reverse construction. Based on a known data table, generate a data graph through the drawing API of office Excel; For picture classification and text2sql data, use the multi-modal large model to generate the picture category and the corresponding SQL data of the question.
3. The construction method of the low-hallucination multi-agent picture question-answering system based on structured understanding according to claim 2, characterized in that, In the step S2, the construction method of the prompt data is as follows; (2.1) For the picture classification Agent, construct the following prompt data: Question template: Please classify the following picture; Picture information: The base64 string of the picture; Agent return result: Picture category, form picture, data graph or other pictures; (2.2) For the knowledge graph generation Agent, construct the following prompt data: Question template: Please convert the following picture into a knowledge graph; Picture information: The base64 string of the picture; Agent return result: The knowledge graph of the picture text content, represented in json structure; (2.3) For the data table generation Agent, construct the following prompt data: Problem template: Please convert the following image into a data table; Image information: The base64 string of the image; Agent return result: A table of the text content of the image, represented in markdown format; (2.4) Generate an Agent for SQL, using the following prompt data: Problem template: Given the data table structure: (Insert knowledge graph or data table structure), the user's input question is: (Insert user question): Please generate a mysql query statement? Agent return result: The query statement.
4. The construction method of the low-hallucination multi-Agent picture question answering system based on structured understanding according to claim 3, wherein, In step S2, during the training process, use LoRA fine-tuning to fine-tune the original multi-modal large language model to obtain the trained multi-modal large language model.
5. The construction method of the low-hallucination multi-Agent picture question answering system based on structured understanding according to claim 4, wherein The specific method of the fine-tuning is as follows: During the training process, based on the prompt data, perform low-rank adaptation training on the original multi-modal large language model, add branch matrices A and B to both ends of the original multi-modal large language model for slight adjustment of the original weights of the model, and the calculation method is: , where is the weight of a certain layer of the original model, is the adjusted weight, has a shape of n, and the dimension of the side branch matrix B is , and the dimension of the side branch matrix A is , m < n.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by the processor to implement the construction method of the low-hallucination multi-Agent image question-answering system based on structured understanding as described in any one of claims 1 to 5.