Chart question-answering method and system based on large language model and multi-modal large language model
By generating structured tables for the chart and disassemblying them into sub-problems, combining the advantages of large language models and multimodal large language models, the problems of missing information and fragile in the chart Q&A are solved, and more accurate and reliable answers to chart questions are achieved.
Patent Information
- Application Number
- CN202411910444.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to fully utilize the cognitive ability of large language models in chart questions and answers, and the reasoning chain of multimodal large language models is fragile, especially when facing problems that require multi-step reasoning and positioning.
A graph question-and-answer method based on large language model and multimodal large language model is proposed. By generating a structured table in JSON format for the original chart, the problem is broken down into sub-problems, and the multimodal large language model is used for inference, and finally summarize and refine it through the large language model to obtain the final answer.
Improves the accuracy of answers to chart questions, reduces hallucinations, improves the reliability of results, and alleviates the shortcomings of a single model when solving chart questions and answers.
Smart Images

Figure CN120031124A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a diagram question-answering method and system based on a large language model and a multimodal large language model. Background Art
[0002] Visual chart analysis plays a decisive role in scientific research. People often ask complex reasoning questions to charts and analyze data in the process of answering questions.
[0003] The existing methods for solving chart question answering include methods based on large language models and methods based on multimodal large language models. The method based on large language models needs to convert the chart into a text table, prompting the large language model to complete the reasoning on the table, so as to indirectly complete the chart question answering. However, the rich visual clues on the chart, such as trends and colors, are difficult to reflect in the table, which makes it difficult to reflect the information missing problem and make it difficult to give full play to the cognitive ability of the large language model. The method based on the multimodal large language model uses the image and text data set to train a multimodal large language model. This model has a wider perceptual field of view and can directly receive charts and complete reasoning. However, there are significant distribution differences between the image and text pre-training data and the pure text pre-training data of the native large language model, which damages the original reasoning ability of the large language model, making some multimodal large language model reasoning chains fragile, and the effect is not good when facing problems that require multi-step reasoning and positioning. Summary of the invention
[0004] In view of the fact that the existing technology is difficult to fully utilize the cognitive ability of large language models, and some multimodal large language models have fragile reasoning chains, the present invention proposes a graph question and answer method and system based on a large language model and a multimodal large language model, which can alleviate the defects of a single model in solving graph question and answer, and provide more accurate answers to graph questions.
[0005] To achieve the above objectives, the technical solution of the present invention includes the following contents.
[0006] A graph question answering method based on a large language model and a multimodal large language model, the method comprising:
[0007] Generate a structured table in JSON format for the original chart;
[0008] Based on the structured table, the problem of the original graph is decomposed into n sub-problems q i ;
[0009] Based on the original graph and structured table, each sub-question q i Reasoning is performed to arrive at a final reasoned answer to the problem.
[0010] Furthermore, generating a structured table in JSON format for the original chart includes:
[0011] Generate structured tables for original charts based on the table generation model DEPLOT;
[0012] Convert the structured table into JSON format.
[0013] Further, based on the structured table, the problem of the original diagram is decomposed into n sub-problems q i ,include:
[0014] Extract complete row and column headers from structured tables;
[0015] The row header, the column header, the question and the manually constructed context example embedding question are input into the large language model to obtain n sub-questions q i .
[0016] Furthermore, when n=1, the original graph and the structured table are used to calculate each sub-problem q i Perform reasoning to obtain the final reasoning answer to the problem, including:
[0017] The sub-problem q i , the original chart and the structured table are input into a multimodal large language model for reasoning to obtain a final reasoning answer to the question.
[0018] Furthermore, when n≠1, the original graph and the structured table are used to calculate each sub-problem q i Perform reasoning to obtain the final reasoning answer to the problem, including:
[0019] The subproblem q j , the original graph and the structured table are input into the multimodal large language model for reasoning to obtain a sub-answer a j ; where j∈[1,n-1];
[0020] The subproblem q j With sub-answer a j Combine to get the intermediate reasoning step s j ;
[0021] The intermediate reasoning step s j With subproblem q n Splicing to obtain a first splicing result;
[0022] Inputting the first splicing result, the original chart, and the structured table into a multimodal large language model for reasoning to obtain a preliminary reasoning answer to the question;
[0023] The intermediate reasoning step s j , splicing the question and the preliminary reasoning answer to obtain a second splicing result;
[0024] Constructing a prompt word according to the second concatenation result and the manually constructed context example to guide the large language model to perform reasoning, and obtaining an output result of the large model; wherein the output result of the large model includes: verifying the answer or not giving the answer;
[0025] Based on the preliminary reasoning answer and the output result of the large model, a final reasoning answer to the question is obtained.
[0026] Furthermore, based on the preliminary reasoning answer and the output result of the large model, a final reasoning answer to the question is obtained, including:
[0027] When the preliminary reasoning answer is consistent with the verification answer, or the output result of the large model is that no answer is given, the preliminary reasoning answer is used as the final reasoning answer to the question;
[0028] In the case that the preliminary inference answer is inconsistent with the verification answer, the verification answer is used as the final inference answer to the question.
[0029] A graph question answering system based on a large language model and a multimodal large language model, the system comprising:
[0030] A sense source expansion module, used to generate a structured table in JSON format for the original chart;
[0031] The problem decomposition module decomposes the problem of the original diagram into n sub-problems q based on the structured table i , n is a positive integer;
[0032] The question reasoning module is used to solve each sub-question q based on the original graph and structured table. i Reasoning is performed to arrive at a final reasoned answer to the problem.
[0033] An electronic device, comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the diagram question answering method based on a large language model and a multimodal large language model as described above is implemented.
[0034] A computer-readable storage medium, characterized in that computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the graph question answering method based on a large language model and a multimodal large language model is implemented as described in any of the above.
[0035] A computer program product, characterized in that when the computer program product is run on a computer device, the computer device executes any of the above-mentioned graph question answering methods based on a large language model and a multimodal large language model.
[0036] Compared with the prior art, the present invention has at least the following beneficial effects.
[0037] 1) The present invention integrates enhanced perception sources from charts and tables, rather than allowing a large language model or a multimodal large language model to perceive information from a single chart or table, thereby improving the accuracy of the answer to the question.
[0038] 2) The present invention further refines and verifies the reasoning answers obtained based on the sub-questions, which helps to reduce hallucinations and improve the reliability of the results. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 Flowchart of the graph question answering method based on large language model and multimodal large language model. DETAILED DESCRIPTION
[0040] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present invention is described and illustrated below in conjunction with the accompanying drawings.
[0041] The present invention aims to first expand a single chart perception source into an enhanced perception source of chart plus table data, then decompose complex problems into simple sub-problems that are easy to solve based on the table based on the large language model, guide the multimodal large language model to use perception capabilities to initially solve the sub-problems based on the prompt word strategy, and finally input the generated sub-answers and sub-problems into the large language model for summary and refinement to obtain the final answer. In the specific implementation technology, the expansion of the chart into a chart plus a table is implemented using the existing chart conversion expert model, where the chart before the expansion and the chart after the expansion are the same, only the associated table data is added. The following is a specific technical solution:
[0042] The diagram question answering method based on the large language model and the multimodal large language model of the present invention is as follows: Figure 1 As shown, the steps include steps 1 to 3.
[0043] Step 1: Generate a structured table in JSON format for the original chart.
[0044] The present invention generates additional tables for charts based on the existing expert model DEPLOT to expand the perception source. When receiving the original chart input, a structured table is generated for the original chart based on the table generation model DEPLOT, and the table will be converted into JSON format and used as an enhanced perception source together with the original chart.
[0045] Step 2: Based on the structured table, the problem of the original diagram is decomposed into n sub-problems q i .
[0046] The present invention decomposes complex problems into several simpler sub-problems based on a large language model. When a question about a chart is received, the complete row and column headers are first extracted from the JSON table associated with the chart. These header information reflects part of the visual information of the chart so that the granularity of problem decomposition has a reference standard. Subsequently, the header and question with positioning information will be input into the large language model together with the manually constructed context example to prompt the large language model to decompose the current problem. In this step, the selectable large language model can be a large language model of any scale.
[0047] Step 3: For each sub-question q based on the original graph and structured table i Reasoning is performed to arrive at a final reasoned answer to the problem.
[0048] The present invention perceives enhanced perception sources including charts and tables based on a multimodal large language model, and infers the decomposed sub-problems.
[0049] In one embodiment, when the complexity of a sub-problem is low, that is, when the original problem is broken down into only one sub-problem, the sub-problem will be directly input into the multimodal large language model together with the enhanced perception source for reasoning to obtain the final reasoning answer to the problem.
[0050] In another embodiment, when the complexity of the sub-problems is high, that is, the original problem is broken down into multiple sub-problems, the final reasoning answer to the problem includes two sub-steps: generating a preliminary reasoning answer and summarizing and verifying.
[0051] Step 3.1: Preliminary reasoning answer generation.
[0052] When the complexity of the sub-problems is high, that is, when the original problem is broken down into multiple sub-problems, all sub-problems except the last sub-problem will be input into the multimodal large language model in sequence together with the enhanced perception source, and the multimodal large language model will generate sub-answers based on the prompt engineering prompts. The obtained sub-answers will be combined with the sub-problems to obtain intermediate reasoning steps. The intermediate reasoning steps are spliced with the last retained sub-problem, and then input into the multimodal large language model again together with the enhanced perception source to obtain a preliminary reasoning answer. In this step, the selectable multimodal large language model can be any multimodal large language model pre-trained on text recognition or chart datasets.
[0053] Step 3.2: Summarize and verify.
[0054] The present invention further summarizes and refines the obtained preliminary reasoning answers based on the large language model. In the summary and verification stage, the intermediate reasoning steps are spliced with the original question and the preliminary reasoning answers, and prompt words are constructed together with handwritten context examples to guide the large language model to further summarize and refine the obtained intermediate reasoning steps to obtain reasoning answers related to the original question. The large language model uses the reasoning answer as the verification answer and compares it with the preliminary reasoning answer. If they are consistent or the verification answer is not obtained according to the intermediate reasoning steps, the preliminary reasoning answer is retained. If they are inconsistent, the verification answer is retained. In this step, the selectable large language model can be a large language model of any size.
[0055] In summary, the present invention combines a large language model and a multimodal large language model in the form of a pipeline, and uses the cognitive ability of the large language model to break down complex problems into simple sub-problems, which are then input into the multimodal large language model together with charts and tables for preliminary reasoning. The obtained sub-answers will be combined with the sub-problems and input into the large language model for summary and refinement to obtain the final answer. The method proposed in the present invention combines the advantages of the large language model and the multimodal large language model to jointly solve the question and answer on the chart, alleviating the defects of a single model in solving chart question and answer.
[0056] In addition, the above embodiment uses a zero-sample learning-based solution to guide the multimodal large language model to perceive the required information from both charts and tables. This requires the multimodal large model to have certain chart understanding and long context understanding capabilities. This part can be replaced by instruction fine-tuning. By constructing the perception source data in the form of charts and tables, and adding corresponding questions to generate instruction data sets, fine-tune the open source model, and provide a more accurate local deployment solution.
[0057] The beneficial effects of the present invention are fully illustrated below in combination with the prior art solutions and the conventional chart question-answering dataset ChartQA.
[0058] Among them, the prior art uses the DOMINO solution. DOMINO proposes a dual system of graph question-answering reasoning consisting of two parts, which includes System-1 for visual information extraction and System-2 for decomposing questions and giving final answers based on reasoning. System-1 is composed of a trained DEPLOT model, and System-2 is composed of a trained large language model. When receiving the original question, System-2 selects from the manually constructed atomic query operations containing multiple questions, and then asks System-1 for query. At this time, System-1 generates an intermediate query result, which will be returned to System-2. Based on the result and the previous query, System-2 selects a suitable query from the atomic query operation again to ask questions about the graph. This process is repeated until System-2 believes that all information can be combined to infer the final answer. DOMINO needs to perform instruction fine-tuning and pre-training on System-1 and System-2 respectively. For System-1, a large number of question-answer pair data sets are generated using templates based on the manually constructed atomic query operations, which are used to train the DEPLOT model in System-1. For System-2, we collected 100 high-quality question decomposition and reasoning samples as training datasets to train the large language model in System-2.
[0059] As shown in Table 1, compared with the existing method DOMINO, the present invention achieves more beneficial average performance on ChartQA. On Augmented, a subset of ChartQA that only contains simpler questions generated by templates, DOMINO achieves better performance, which comes from the fact that DOMINO uses more data for additional supervised training and can effectively learn in the simpler Augmented. However, DOMINO lags far behind the present invention on Human, a subset of ChartQA that contains questions that are handwritten and require multi-step reasoning.
[0060] Table 1 Performance comparison of different methods on ChartQA
[0061]
[0062] The above description is only an explanation of a specific example of the present invention and does not impose any limitation on the present invention. Obviously, for those with professional knowledge in this field, once the content and principle of the present invention are understood, it is possible to make various modifications and changes in form and details without violating the original principle and structure of the present invention. However, these amendments and changes based on the idea of the present invention are still considered to be within the scope of protection of the claims of the present invention.
Claims
1. A graph question answering method based on a large language model and a multimodal large language model, characterized in that: The method comprises: generating a structured table in JSON format for the original chart; Based on the structured table, the problem of the original graph is decomposed into n sub-problems q i , n is a positive integer; Based on the original graph and structured table, each sub-question q i Reasoning is performed to arrive at a final reasoned answer to the problem.
2. The method according to claim 1, characterized in that The structured table in JSON format is generated for the original chart, including: Generate structured tables for original charts based on the table generation model DEPLOT; Convert the structured table into JSON format.
3. The method according to claim 1, characterized in that Based on the structured table, the problem of the original graph is decomposed into n sub-problems q i ,include: Extract complete row and column headers from structured tables; The row header, the column header, the question and the manually constructed context example embedding question are input into the large language model to obtain n sub-questions q i .
4. The method according to claim 1, characterized in that: When n=1, the original graph and the structured table are used to solve each sub-problem q i Perform reasoning to obtain the final reasoning answer to the problem, including: The sub-problem q i , the original chart and the structured table are input into a multimodal large language model for reasoning to obtain a final reasoning answer to the question.
5. The method according to claim 1, characterized in that When n≠1, the original graph and the structured table are used to solve each sub-problem q i Perform reasoning to obtain the final reasoning answer to the problem, including: The subproblem q j , the original graph and the structured table are input into the multimodal large language model for reasoning to obtain a sub-answer a j ; where j∈[1,n-1]; The subproblem q j With sub-answer a j Combine to get the intermediate reasoning step s j ; The intermediate reasoning step s j With subproblem q n Splicing to obtain a first splicing result; Inputting the first splicing result, the original chart, and the structured table into a multimodal large language model for reasoning to obtain a preliminary reasoning answer to the question; The intermediate reasoning step s j , splicing the question and the preliminary reasoning answer to obtain a second splicing result; Constructing a prompt word according to the second concatenation result and the manually constructed context example to guide the large language model to perform reasoning, and obtaining an output result of the large model; wherein the output result of the large model includes: verifying the answer or not giving the answer; Based on the preliminary reasoning answer and the output result of the large model, a final reasoning answer to the question is obtained.
6. The method according to claim 5, characterized in that Based on the preliminary reasoning answer and the output result of the large model, the final reasoning answer to the question is obtained, including: When the preliminary reasoning answer is consistent with the verification answer, or the output result of the large model is that no answer is given, the preliminary reasoning answer is used as the final reasoning answer to the question; In the case that the preliminary inference answer is inconsistent with the verification answer, the verification answer is used as the final inference answer to the question.
7. A graph question answering system based on a large language model and a multimodal large language model, characterized in that: The system includes: a sense source expansion module, used to generate a structured table in JSON format for the original chart; A problem decomposition module is used to decompose the problem of the original diagram into n sub-problems q based on the structured table. i , n is a positive integer; The question reasoning module is used to solve each sub-question q based on the original graph and structured table. i Reasoning is performed to arrive at a final reasoned answer to the problem.
8. An electronic device, characterized in that: The electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the graph question answering method based on a large language model and a multimodal large language model is implemented as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by the processor, the graph question answering method based on a large language model and a multimodal large language model is implemented as described in any one of claims 1 to 6.
10. A computer program product, characterized in that When the computer program product runs on a computer device, the computer device executes the graph question answering method based on a large language model and a multimodal large language model as described in any one of claims 1 to 6.
Citation Information
Cited By
Multi-agent cooperation method and system for domestic operating system
CN122086421A