An evaluation and testing method, medium, and equipment for a multimodal graph-based question-answering model.

By designing an evaluation and testing method for a multimodal graph question-answering model, and utilizing various evaluation indicators and judge models, the problem of incomplete evaluation in existing technologies is solved, and a comprehensive evaluation and accuracy judgment of the multimodal model on graph question-answering tasks is achieved.

CN119760369BActive Publication Date: 2025-11-14北京中科闻歌科技股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411808970.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-11-14
Estimated Expiration
2044-12-10

AI Technical Summary

Technical Problem

Existing benchmarking methods for multimodal graph question answering models lack comprehensiveness and are difficult to accurately evaluate the stability and performance of the models in various specific task scenarios, especially in the evaluation of subjective and objective task types.

Method used

An evaluation and testing method for a multimodal graph question-answering large model is designed. By acquiring different types of test datasets, including low-order and high-order tasks, multiple evaluation metrics and referee models (such as GPT4o) are used to match and score the model output results, ensuring the comprehensiveness and accuracy of the evaluation.

Benefits of technology

It enables a comprehensive evaluation of multimodal large models on graph question answering tasks, and can more accurately judge the accuracy and consistency of model output results, thereby improving the evaluation efficiency of the model's graph understanding and generation capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119760369B_ABST
    Figure CN119760369B_ABST
Patent Text Reader

Abstract

This invention relates to the field of large model evaluation, and particularly to an evaluation and testing method, medium, and device for a multimodal graph-based question-answering large model. It includes: inputting a judgment-type test dataset into the large model to be evaluated to obtain the output results of the judgment-type model. The question information in the judgment-type question-answer pairs includes the question text and a prompt indicating that the answer can only be positive or negative. The accuracy information of all fill-in-the-blank, selection, and judgment-type model output results is statistically analyzed to generate performance evaluation information for the large model to be evaluated. In this invention, considering the potential variability in the multimodal large model's adherence to instructions, the evaluation of low-order task performance uses three types of instructions—judgment questions (positive and negative), fill-in-the-blank questions, and selection questions—to ask questions to the model to be evaluated, thereby providing a more comprehensive evaluation of the large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large model evaluation, and in particular to an evaluation and testing method, medium, and equipment for a multimodal graph question-answering large model. Background Technology

[0002] With the rapid development of deep learning technology, the number of model parameters has grown exponentially, from millions to hundreds of billions or even trillions. This growth presents new challenges to computing resources, model training methods, and model performance evaluation. The application areas of large models are also constantly expanding, from natural language processing to computer vision and recommender systems, covering almost all AI application fields. The different task characteristics and requirements across various fields make it difficult for a single evaluation standard to meet the needs of all scenarios. As models become increasingly complex, how to improve their efficiency while ensuring performance has become a crucial issue.

[0003] In recent years, with the emergence of GPT4o, large-scale multimodal language models have demonstrated significant capabilities in multimodal understanding and generation. However, their understanding of chart data is limited. Multimodal chart question answering, as an important research direction in the field of multimodal artificial intelligence, aims to return the answer that best matches the given chart image and question description, thereby better helping users understand the content of the chart image or generating text descriptions related to the chart image. This task strives to improve the understanding and analysis capabilities of chart tasks through a fusion of visual and linguistic understanding of chart data. This type of task can be divided into two categories: subjective question answering and objective question answering. Objective question answering aims to extract the location, value, or classification of specific data points in the chart image, such as data retrieval, extreme values, and classification; subjective question answering requires the model to generate deeper interpretations and reasoning capabilities based on the content of the chart image, such as chart transformation, chart redrawing, and summarizing chart viewpoints.

[0004] While existing benchmarking methods for multimodal graph question answering models have made some progress in evaluating these models, they lack comprehensive testing method designs for both subjective and objective task types. Consequently, the models exhibit poor stability when facing various specific task scenarios. Furthermore, they suffer from a series of key problems, including overly simplistic benchmarks, limited task evaluation dimensions, and a single evaluation metric. These factors make accurately evaluating the graph understanding capabilities of large multimodal models challenging. Summary of the Invention

[0005] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows:

[0006] According to one aspect of the present invention, an evaluation and testing method for a large multimodal graph question-answering model is provided, the method comprising the following steps:

[0007] Obtain the corresponding test datasets for fill-in-the-blank, multiple-choice, and true / false questions for each type of low-level task. In the fill-in-the-blank test dataset, each chart-type test image corresponds to a set of fill-in-the-blank question-and-answer pairs. In the multiple-choice test dataset, each chart-type test image corresponds to a set of multiple-choice question-and-answer pairs. In the true / false test dataset, each chart-type test image corresponds to a set of true / false question-and-answer pairs. Low-level tasks focus on specific, detail-oriented queries, seeking or comparing precise data points in charts, and involve question-and-answer tasks related to direct factual information retrieval.

[0008] Input each chart-type test image and the corresponding question-answer pair information from the fill-in-the-blank test dataset into the large model to be evaluated, and obtain the fill-in-the-blank model output results for each chart-type test image.

[0009] The output of the fill-in-the-blank model corresponding to each chart-type test image is matched with the answer information in the corresponding question-answer pair and then input into GPT4o to generate accuracy information of the output of the fill-in-the-blank model corresponding to each chart-type test image.

[0010] Each chart-type test image and the question information from the question-answer pair in the judgment test dataset are input into the large model to be evaluated to obtain the output result of the judgment model for each chart-type test image. The question information in the question-answer pair in the judgment test dataset includes the text of the question itself and the hint that the answer can only be positive or negative.

[0011] Input each chart-type test image and the question information in the question-answer pair from the selection test dataset into the large model to be evaluated, and obtain the selection model output results corresponding to each chart-type test image.

[0012] Using regular expressions, the output results of the selection model or the judgment model corresponding to each chart-type test image are matched with the answer information in the corresponding question-answer pair to generate accuracy information for the output results of the selection model or the judgment model corresponding to each chart-type test image.

[0013] The accuracy information of the output results of all fill-in-the-blank, selection, and judgment models corresponding to each type of low-order task is statistically analyzed to generate the performance evaluation information of the large model to be evaluated for each type of low-order task.

[0014] Furthermore, the method also includes:

[0015] Obtain the advanced test dataset for each type of advanced task. Each chart type test image in the advanced test dataset corresponds to a set of advanced question-answer pairs. Advanced tasks focus on understanding the overall trend, pattern, or contextual summary of the chart.

[0016] Input each chart-type test image and the corresponding question information from the higher-order question-answer pair in the higher-order test dataset into the large model to be evaluated, and obtain the output results of the higher-order model corresponding to each chart-type test image.

[0017] The output of the high-order model corresponding to each chart-type test image and the corresponding answer information in the high-order question-answer pair are input into GPT4o for evaluation and scoring to generate accuracy information for the output of the high-order model corresponding to each chart-type test image. GPT4o evaluates and scores the results based on several dimensions, including the fluency of the generated answers, the ability to follow instructions, the length of the generated text, and the completeness of the results.

[0018] The accuracy information of the output results of all high-order models corresponding to each type of high-order task is statistically analyzed to generate the performance evaluation information of the large model to be evaluated for each type of high-order task.

[0019] Based on the performance evaluation information of each type of high-order task and each type of low-order task, the evaluation information of the large model to be evaluated is generated.

[0020] Furthermore, obtain the test dataset corresponding to each type of high-order or low-order task, including:

[0021] The self-instruct model is used to generate multiple seed questions for each chart class test image.

[0022] The Qwen2.5-0.5b model is used to score each seed problem, and the seed problem with the highest score is selected as the intermediate problem.

[0023] GPT4o is used to rewrite intermediate problems according to task type, generating multiple target problems corresponding to each chart-type test image.

[0024] The initial labeled test data, consisting of each target question and its corresponding chart-type test image, is input into the Qwen2-VL-72b model to generate the answer information corresponding to each initial labeled test data, thereby obtaining the target test data.

[0025] Furthermore, before using the self-instruct model to generate multiple seed questions corresponding to each chart class test image, the method also includes:

[0026] The initial chart-type test images obtained are all converted to RGB format to generate the first test image set.

[0027] Remove low-quality images from the first test image set to generate the second test image set. Low-quality images include low-resolution, over-compressed, blank or invalid images, and images with duplicate content.

[0028] Median filtering was used to remove noise from the images in the second test image set, generating a chart-type test image set in the test dataset.

[0029] Furthermore, before inputting each chart-type test image and the corresponding question-answer pair information from the fill-in-the-blank test dataset into the large model to be evaluated, in order to obtain the fill-in-the-blank model output results for each chart-type test image, the method also includes:

[0030] Obtain the large model to be evaluated by using either an API interface or model weights.

[0031] Furthermore, the types of chart test images include stacked bar charts, complex line charts, scatter plots, pie charts, regular line charts, grouped bar charts, regular bar charts, 3D bar charts, bubble charts, and mixed charts.

[0032] Furthermore, the categories of low-order tasks include reasoning, anomaly, distribution, correlation, range determination, sorting, filtering, data retrieval, extrema, and clustering.

[0033] Furthermore, advanced tasks include chart classification, opinion suggestions, chart redrawing, chart transformation, chart summarization, chart filling assistance, text generation, chart creation assistance, multi-chart Q&A, chart multi-turn dialogue, chart OCR recognition, and information location.

[0034] According to a second aspect of the present invention, a non-transitory computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the above-described evaluation and testing method for a multimodal graph question-answering large model.

[0035] According to a third aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described evaluation and testing method for a multimodal graph question-answering large model.

[0036] The present invention has at least the following beneficial effects:

[0037] In this invention, during the evaluation and testing of the multimodal graph question-answering model on low-order tasks (i.e., objective question-answering tasks), various types of instructions and evaluation metrics are used to quantitatively analyze the performance of the multimodal model on graph question-answering tasks, resulting in a more reasonable evaluation. Furthermore, given the potential variability in the multimodal model's adherence to instructions, the evaluation of its performance on low-order tasks includes three types of questions: true / false questions, fill-in-the-blank questions, and multiple-choice questions, presented from both positive and negative perspectives. This richer question variety allows for a more comprehensive evaluation of the multimodal model.

[0038] Furthermore, in judging the consistency between the model output and the corresponding labeled data, to avoid interference from unnecessary information caused by diverse expressions in the output of fill-in-the-blank models, GPT4o is used as the judge model to evaluate the consistency between the model output and the labeled answer information. This allows for a more accurate determination of whether the model output is accurate. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 This is a flowchart of an evaluation and testing method for a multimodal graph-based question-answering model provided in an embodiment of the present invention. Detailed Implementation

[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0042] As one possible embodiment of the present invention, such as Figure 1 As shown, an evaluation and testing method for a large multimodal graph question-answering model is provided, which includes the following aspects:

[0043] S1: Obtain the test dataset.

[0044] In this embodiment, chart question-answering tasks can be categorized into three main types: intrinsic tasks, low-order tasks, and high-order tasks. Intrinsic tasks refer to the capabilities inherent to the model itself, representing the primary tasks or objectives assigned during model design and training; they are the model's fundamental capabilities. High-order tasks focus on questions requiring understanding the overall context or summary of the chart, involving broader, goal-oriented queries aimed at understanding overall trends or patterns. In contrast, low-order tasks focus on specific, detail-oriented queries, seeking precise data points or comparisons within the chart, involving direct retrieval of factual information.

[0045] Based on the focus of self-defined tasks, high-order tasks, and low-order tasks in chart question answering, these tasks can be further classified. Self-defined tasks can be divided into two subcategories: security and self-awareness. Low-order tasks can be divided into ten subcategories: reasoning, anomaly, distribution, correlation, range determination, sorting, filtering, data retrieval, extreme values, and clustering. High-order tasks can be divided into twelve subcategories: chart classification, opinion suggestions, chart redrawing, chart transformation, chart summarization, chart filling assistance, text generation, chart creation assistance, multi-chart question answering, chart multi-turn dialogue, chart OCR recognition, and information location.

[0046] The evaluation method in this embodiment is mainly to evaluate the performance of the modal graph question answering model in low-order and high-order tasks. Therefore, when obtaining the test dataset, the corresponding test dataset is mainly obtained for the task categories included in the low-order and high-order tasks.

[0047] S1 includes:

[0048] S100: Collection and acquisition of chart data. Specifically, the types of chart test images in this embodiment include stacked bar charts, complex line charts, scatter plots, pie charts, regular line charts, grouped bar charts, regular bar charts, 3D bar charts, bubble charts, and mixed charts.

[0049] For image data in chart-based question-and-answer tasks, the acquisition of image data can be divided into two categories: chart-based images and non-chart-based images.

[0050] The acquisition channels for chart images include collecting them from the internet via keywords, using public datasets for chart question-and-answer tasks, and parsing chart images from academic papers or technical reports. This embodiment employs both public chart datasets and document parsing. The public chart datasets include chart data from 15 public datasets: ChartSumm, DVQA, PlotQA, ChartQA, FigureQA, ScigraphQA, TinyChart, Chart-to-Text, ChartBench, ChartLLama, MMC, ChartAssistant, OneChart, ChartX, and ChartInsights. Document parsing primarily involves extracting data from financial report-type PDF papers and reports. This allows for the acquisition of a sufficient number of chart images.

[0051] The acquisition channels for non-graphic images are mainly publicly available image data from both domestic and international sources, such as image data from publicly available recognition tasks, detection tasks, image generation tasks, and multimodal instruction data (e.g., COCO, LAION, LLAVA-Instruct). In addition, for images used in our proprietary tasks, we collect news articles or logos from various company websites.

[0052] After obtaining a large number of chart-like test images, it is necessary to clean and deduplicate this data to select high-quality and diverse images. The following are the corresponding steps:

[0053] S101: Convert all the obtained initial chart-type test images into RGB format to generate the first test image set.

[0054] This step involves converting and standardizing the image formats, unifying images of different formats into a standard format. The standard format includes the image's storage format and size, and converts all images to RGB format to avoid the influence of different color spaces.

[0055] S102: Remove low-quality images from the first test image set and generate the second test image set. Low-quality images include low-resolution, over-compressed, blank or invalid images, and images with duplicate content.

[0056] S103: Use median filtering to remove noise from the images in the second test image set, and generate a set of chart-type test images in the test dataset.

[0057] In S103, manual inspection can also be used to check for abnormalities such as extreme color values ​​(extreme color values ​​refer to colors in the image background that interfere with the icon's expression) and oversaturation in certain images. Finally, the filtered images are deduplicated to remove duplicate images, thereby ensuring the quality, accuracy, and diversity of the image dataset and improving the efficiency and effectiveness of the next step of instruction annotation.

[0058] After the above processing, the relevant chart images in the test dataset have been obtained. The next step is to generate corresponding instruction annotation information for each chart image. Specifically, in this embodiment, the instruction annotation information is generated by creating multiple question-and-answer pairs based on the content of the chart images. Each question-and-answer pair belongs to a set of annotation information, and these multiple pairs can be annotation information corresponding to multiple types of question-and-answer tasks. For example, the question information in multiple question-and-answer pairs corresponding to a certain chart image may include: 1. What type of chart is this? What are the labels on the x-axis? 3. What are the data labels for each element? 4. Is the length of the bar chart for "Thursday" abnormal?

[0059] Specifically, when annotating the cleaned image data with instructions, the following steps can be followed:

[0060] S104: Use the self-instruct model to generate multiple seed questions Q1 for each chart class test image.

[0061] In this step, you can also set corresponding prompt instructions for the self-instruct model to standardize and improve the effectiveness of the model's output information. For example, the prompt instruction here could be: "You are a professional chart content understanding expert. You can generate chart classification task questions based on the chart content. Be sure to accurately and meticulously analyze the key points: 1. The classification in the generated question must be consistent with the chart type in the image. 2. The generated question should be fluent and highly complex."

[0062] S105: Use the Qwen2.5-0.5b model to score each seed problem and take the seed problem with the highest score as the intermediate problem Q2.

[0063] After generating multiple seed questions for a specific task type, it's necessary to select the highest quality ones for subsequent expansion and rewriting. In this step, you can also set corresponding prompt instructions for the Qwen2.5-0.5b model to standardize and improve the effectiveness of the model's output. For example, the prompt instruction here could be: "You are a professional text content evaluation expert, able to select the highest quality text data based on the input text content. Be sure to evaluate it according to three aspects: text length, text fluency, and text complexity."

[0064] S106: Use GPT4o to rewrite the intermediate questions according to the task type, and generate multiple target questions Q3 corresponding to each chart-type test image.

[0065] This step expands and rewrites the problem based on the highest quality intermediate problem to increase its diversity and richness. In this step, you can also set corresponding prompt instructions for the GPT4o model to standardize and improve the effectiveness of the model's output. For example, the prompt instruction here could be: "You are a professional text content rewriting expert. Your task is to rewrite text content according to diversity, complexity, and fluency. Please note the following points in the rewritten text content: 1. Higher text complexity; 2. Better text fluency; 3. Higher text diversity."

[0066] S107: Input the initial labeled test data, consisting of each target question and the corresponding chart-type test image, into the Qwen2-VL-72b model to generate the answer information corresponding to each initial labeled test data, so as to obtain the target test data.

[0067] Each Q1 and its corresponding chart-type test image form the initial labeled test data. This data is then input into the Qwen2-VL-72b question-answering model to generate the corresponding answer A1. This creates a complete test dataset that includes chart-type test images and question-answer pairs. Thus, a comprehensive and high-quality benchmark dataset is constructed based on charts and task types.

[0068] Furthermore, the generated labeled instruction data (i.e., test data) needs to undergo quality control to ensure its accuracy, completeness, and consistency. Since both questions and answers are automatically generated using a large model during the instruction data labeling process, issues may arise such as mismatches between questions and image content, inconsistencies between questions and answers, and poor quality of answers and questions. Therefore, quality checks are necessary to remove substandard data and ensure the accuracy and consistency of the instruction data.

[0069] S2: Access the large model to be evaluated. S2 includes:

[0070] S201: Obtain the large model to be evaluated by using either the API interface or the model weights.

[0071] By using an API interface, evaluators do not need to understand the model's internal implementation or complex configuration; they only need to call the interface to interact with the model. Furthermore, through API access, the model and inference code can be automatically updated and optimized by the provider, eliminating the need for evaluators to manually update or replace their local models, ensuring that the latest version is always used.

[0072] By integrating models using weights, evaluators can simultaneously evaluate multiple models using the same inference code. Model weight integration is a common requirement in practice. Through efficient model integration methods, evaluators can switch between different models according to actual needs without completely modifying the existing evaluation system architecture. Furthermore, model integration allows multiple large models to be evaluated to be loaded locally, providing a clearer picture of the system resource consumption of each model during testing. This facilitates load balancing between models, preventing system performance degradation due to overload of any single model. Especially in high-concurrency scenarios, rationally allocating tasks to different models can effectively distribute computational pressure and improve system responsiveness.

[0073] S3: Evaluate and test the large model to be evaluated. This includes automated evaluation sets and manual evaluation.

[0074] S3 includes evaluation methods for low-order tasks. Specifically, low-order tasks include inference, anomaly, distribution, correlation, range determination, sorting, filtering, data retrieval, extrema, and clustering. The following steps are the specific steps for evaluating low-order tasks.

[0075] S301: Obtain the fill-in-the-blank, multiple-choice, and true / false test datasets for each type of low-level task. In the fill-in-the-blank test dataset, each chart-type test image corresponds to a set of fill-in-the-blank question-and-answer pairs. In the multiple-choice test dataset, each chart-type test image corresponds to a set of multiple-choice question-and-answer pairs. In the true / false test dataset, each chart-type test image corresponds to a set of true / false question-and-answer pairs. Low-level tasks focus on specific, detail-oriented queries, seeking or comparing precise data points in charts, involving direct factual information retrieval. The low-level task test data for each chart type includes fill-in-the-blank, true / false, and multiple-choice question types.

[0076] S302: Input each chart-type test image and the corresponding question-answer pair information from the fill-in-the-blank test dataset into the large model to be evaluated, so as to obtain the output results of the fill-in-the-blank model corresponding to each chart-type test image.

[0077] S303: Input the output of the fill-in-the-blank model corresponding to each chart-type test image and the answer information in the corresponding question-answer pair into GPT4o for matching, so as to generate the accuracy information of the output of the fill-in-the-blank model corresponding to each chart-type test image.

[0078] Evaluation of low-order tasks requires models to provide precise judgments or numerical values. However, due to the potential lack of conciseness in the output of large multimodal models, precise character matching may not accurately match the core fields in the answer, thus rendering it insufficient for evaluation. To avoid interference from unnecessary information caused by diverse expressions in fill-in-the-blank model outputs, GPT4o is used as the judge model to evaluate the consistency between the model output and the labeled answer information. By setting a tolerance level, answers falling within a specified range are considered correct. This allows for a more accurate determination of the accuracy of the model's output.

[0079] Specifically, you can also set corresponding prompt instructions for the GPT4o model to standardize and improve the effectiveness of the model's output. For example, the prompt instruction here could be: "You are a professional content parsing expert, capable of extracting answer-related content from text. Please note the following: 1. The parsing result should be the content appearing in the text. 2. The parsing result must be complete. 3. Only output the complete parsed answer; it should not contain other characters, punctuation marks, etc."

[0080] S304: Input each chart-type test image and the question information from the question-answer pair in the judgment class test dataset into the large model to be evaluated, and obtain the output result of the judgment class model corresponding to each chart-type test image. The question information in the question-answer pair in the judgment class test dataset includes the question text and the hint that the answer can only be positive or negative.

[0081] When evaluating true / false questions, providing ambiguous or neutral numerical answers is equivalent to the model not giving a clear response, which negatively impacts accuracy. Therefore, this invention introduces a two-sided decision design for each category of a given chart, explicitly requiring the multimodal large model to provide positive or negative answers to true / false questions, avoiding ambiguous statements.

[0082] Specifically, this can be achieved by adding hints to the question information in the judgment-based test dataset, indicating that the answer can only be positive or negative. For example, the question information in the question-answer pairs in the judgment-based test dataset could include questions such as: "Let's answer the following questions one by one: 1. What type of chart is this? What are the labels on the x-axis? 3. What are the data labels for each element? 4. Is the length of the bar chart for 'Thursday' abnormal?" and note: "You only need to answer 'yes' or 'no'."

[0083] Alternatively, a separate prompt message could be set for the large model to be evaluated. For example, the prompt message could be: "You are a professional text content analysis expert, capable of selecting the most suitable answer from the text content. Please note the following: 1. The selected answer must be the most relevant option, and cannot be a label other than the option labels. 2. For true / false questions, the specific result should be clearly given, and labels such as 'uncertain,' 'neutral,' 'moderate,' or 'neutral' should not appear."

[0084] S305: Input each chart-type test image and the question information in the question-answer pair from the selection test dataset into the large model to be evaluated, and obtain the selection model output result corresponding to each chart-type test image.

[0085] S306: Use regular expressions to match the output results of the selection model or the judgment model corresponding to each chart test image with the answer information in the corresponding question-answer pair to generate accuracy information of the output results of the selection model or the judgment model corresponding to each chart test image.

[0086] Since the large model being evaluated is limited by the conciseness of its output when responding to choice or judgment questions, it is not necessary to use an additional language model to determine whether the output is a specific answer, or to use regular expression matching to determine the correctness of the model's response.

[0087] S307: Statistically analyze the accuracy information of the output results of all fill-in-the-blank models, selection models, and judgment models corresponding to each type of low-order task, and generate the performance evaluation information of the large model to be evaluated for each type of low-order task.

[0088] Since there are multiple test data sets for each type of task, multiple judgments on the model's output results can be obtained. Ultimately, statistical analysis can be used to more objectively and accurately evaluate the large evaluation model.

[0089] Furthermore, S3 also includes evaluation methods for advanced tasks. Specifically, advanced tasks include chart classification, opinion suggestions, chart redrawing, chart transformation, chart summarization, chart completion assistance, text generation, chart creation assistance, multi-chart question and answer, multi-turn chart dialogue, chart OCR recognition, and information location. The following steps are the specific steps for evaluating advanced tasks.

[0090] S308: Obtain the advanced test dataset for each advanced task category. Each chart type test image in the advanced test dataset corresponds to a set of advanced question-answer pairs. Advanced tasks focus on understanding the overall trend, pattern, or contextual summary of the chart.

[0091] S309: Input each chart-type test image and the corresponding question information in the high-order question-answer pair from the high-order test dataset into the large model to be evaluated, so as to obtain the output results of the high-order model corresponding to each chart-type test image.

[0092] S310: Input the output of the high-order model corresponding to each chart-type test image and the answer information in the corresponding high-order question-answer pair into GPT4o for evaluation and scoring, so as to generate accuracy information of the output of the high-order model corresponding to each chart-type test image. GPT4o evaluates and scores the results from several dimensions, including the fluency of the generated answer, the ability to follow instructions, the length of the generated text, and the completeness of the result.

[0093] For higher-order tasks, since the answers are more subjective and lack fixed standard answers, it is necessary to design a referee model (GPT4o) to evaluate the performance of the model under evaluation. To improve the accuracy of the referee model's evaluation, it needs to score the generated answers by setting different prompt instructions based on the differences in each task type. For example, the prompt instruction could be: "You are a professional expert in evaluating chart question-answering models. You need to evaluate the model's answer quality in detail from different dimensions based on **charts**, **questions**, **human reference answers**, and **model answers**. Output an evaluation quality score, which corresponds to an integer value between [0-5]. A higher score indicates better text quality, and a lower score indicates worse text quality." In this embodiment, the score range is [0, 5], with a higher score indicating better answer quality. The scoring criteria are based on the fluency of the generated answer, the ability to follow instructions, the length of the generated text, and the completeness of the result.

[0094] S311: Statistically analyze the accuracy information of the output results of all high-order models corresponding to each type of high-order task, and generate the performance evaluation information of the large model to be evaluated for each type of high-order task.

[0095] S312: Generate evaluation information for the large model to be evaluated based on the performance evaluation information of each type of high-order task and each type of low-order task.

[0096] In this embodiment, the chart question-and-answer is divided into high-order tasks and low-order tasks, and different evaluation methods are designed according to the task type. This can meet and conform to the evaluation requirements of actual tasks, and the benchmark testing method has good generalization.

[0097] Furthermore, the evaluation of large evaluation models can also include human evaluation.

[0098] Specifically, human evaluation methods primarily assess the quality and accuracy of the model's output through manual evaluation, thereby evaluating its performance in graph question answering tasks. This evaluation method is often used to supplement automated evaluation, ensuring that the model's performance in complex, subjective, or ambiguous scenarios is reasonably evaluated.

[0099] The manual evaluation in this invention adopts a direct evaluation method, which analyzes the quality of the text generated by the model and scores it manually. The scoring range is [0, 5], with higher scores indicating better text quality. The scoring criteria are as follows: 0 points indicate incorrect reasoning process and incorrect answer; 1 point indicates accurate reasoning process, incorrect answer, and non-compliance with instructions; 2 points indicate incorrect reasoning process or non-compliance with instructions, and accurate answer; 3 points indicate correct answer and correct reasoning process, but incomplete reasoning content, incoherent text, overly brief text description, and non-compliance with instructions; 4 points indicate correct answer, correct and complete reasoning process, fluent text, and generally good compliance with instructions; 5 points indicate correct answer, correct and complete reasoning process, fluent text, good compliance with instructions, and good text quality.

[0100] Benchmarking methods for multimodal large-scale graph question answering models aim to design and optimize testing methods to comprehensively evaluate the model's performance in processing graph data and natural language questions. By combining multiple information sources such as images, text, and data, this invention explores how to effectively measure the model's understanding, reasoning, and text generation capabilities, including the accuracy of data extraction, analysis, and inference from graphs, as well as the accurate answers to natural language questions. Therefore, this invention attempts to introduce a novel and comprehensive task evaluation method and test dataset to evaluate the performance of multimodal large-scale models on graph question answering tasks, providing strong technical support for improving the performance of multimodal large-scale models in the graph question answering field.

[0101] This invention primarily relates to benchmarking techniques for large-scale multimodal chart question-answering models. It utilizes a series of standardized test datasets and task testing methods to evaluate the model's performance in chart understanding and question-answering tasks. These benchmarks cover various chart types (such as bar charts, line charts, pie charts, scatter plots, etc.) and related natural language questions. The tests assess the model's ability to extract information from charts, understand context, perform reasoning, and ultimately generate accurate answers.

[0102] By combining low-order and high-order tasks with human evaluation, a more comprehensive and accurate evaluation report on the performance of the large model on various tasks can be generated. This not only helps researchers reveal the model's strengths and limitations but also provides theoretical basis and practical guidance for further optimization of multimodal AI systems.

[0103] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.

[0104] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0105] In an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided.

[0106] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: entirely hardware implementations, entirely software implementations (including firmware, microcode, etc.), or implementations combining hardware and software aspects, collectively referred to herein as “circuits,” “modules,” or “systems.”

[0107] An electronic device according to this embodiment of the invention. The electronic device is merely an example and should not be construed as limiting the functionality or scope of the embodiments of the invention.

[0108] Electronic devices are manifested in the form of general-purpose computing devices. Components of an electronic device may include, but are not limited to: at least one processor, at least one memory, and buses connecting different system components (including memory and processor).

[0109] The memory stores program code that can be executed by a processor, causing the processor to perform the steps described in the "Exemplary Methods" section above, according to various exemplary embodiments of the present invention.

[0110] The storage may include readable media in the form of volatile storage, such as random access memory (RAM) and / or cache memory, and may further include read-only memory (ROM).

[0111] The storage may also include programs / utilities having a set (at least one) of program modules, including but not limited to: an operating system, one or more applications, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.

[0112] A bus can represent one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus that uses any of the various bus architectures.

[0113] The electronic device can also communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., routers, modems, etc.). This communication can be performed via input / output (I / O) interfaces. Furthermore, the electronic device can communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter. The network adapter communicates with other modules of the electronic device via a bus. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0114] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0115] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible embodiments, various aspects of the present invention may also be implemented as a program product comprising program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of the present invention described in the "Exemplary Methods" section above.

[0116] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0117] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0118] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0119] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0120] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0121] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0122] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An evaluation and testing method for a multimodal graph-based question-answering model, characterized in that, The method includes the following steps: Obtain the fill-in-the-blank test dataset, selection test dataset, and judgment test dataset corresponding to each type of low-level task; in the fill-in-the-blank test dataset, each chart test image corresponds to a set of fill-in-the-blank question-and-answer pairs; in the selection test dataset, each chart test image corresponds to a set of selection question-and-answer pairs; in the judgment test dataset, each chart test image corresponds to a set of judgment question-and-answer pairs; the low-level tasks are question-and-answer tasks that focus on specific, detail-oriented queries, seeking or comparing precise data points in charts, and involving direct factual information retrieval. Input each chart-type test image and the corresponding question-answer pair information from the fill-in-the-blank test dataset into the large model to be evaluated, so as to obtain the output results of the fill-in-the-blank model corresponding to each chart-type test image. The output results of the fill-in-the-blank model corresponding to each chart-type test image are matched with the answer information in the corresponding question-answer pair and then input into GPT4o to generate accuracy information of the output results of the fill-in-the-blank model corresponding to each chart-type test image. Input each chart-type test image and the question information in the question-answer pair from the judgment-type test dataset into the large model to be evaluated, and obtain the output result of the judgment-type model corresponding to each chart-type test image. The question information in the question-answer pairs in the judgment-type test dataset includes the question text and a prompt that the answer can only be positive or negative. Input each chart-type test image and the question information in the question-answer pair from the selection test dataset into the large model to be evaluated, and obtain the output results of the selection model corresponding to each chart-type test image. Using regular expressions, the output results of the selection model or the judgment model corresponding to each chart test image are matched with the answer information in the corresponding question-answer pair to generate accuracy information of the output results of the selection model or the judgment model corresponding to each chart test image. The accuracy information of the output results of all fill-in-the-blank, selection, and judgment models corresponding to each type of low-order task is statistically analyzed to generate the performance evaluation information of the large model to be evaluated for each type of low-order task.

2. The method according to claim 1, characterized in that, The method further includes: Obtain the advanced test dataset corresponding to each type of advanced task; each chart type test image in the advanced test dataset corresponds to a set of advanced question-answer pairs; the advanced task is a task that focuses on understanding the overall trend or pattern or context summary of the chart; Input each chart-type test image and the corresponding question information in the high-order question-answer pair from the high-order test dataset into the large model to be evaluated, so as to obtain the output results of the high-order model corresponding to each chart-type test image. The output of the high-order model corresponding to each chart-type test image and the answer information in the corresponding high-order question-answer pair are input into GPT4o for evaluation and scoring to generate accuracy information of the output of the high-order model corresponding to each chart-type test image; GPT4o evaluates and scores from several dimensions such as the fluency of the generated answer, the ability to follow instructions, the length of the generated text, and the completeness of the result. The accuracy information of the output results of all high-order models corresponding to each type of high-order task is statistically analyzed to generate the performance evaluation information of the large model to be evaluated for each type of high-order task. Based on the performance evaluation information of each type of high-order task and the performance evaluation information of each type of low-order task, the evaluation information of the large model to be evaluated is generated.

3. The method according to claim 2, characterized in that, Obtain the test dataset corresponding to each type of high-order or low-order task, including: Use the self-instruct model to generate multiple seed questions for each chart class test image; The Qwen2.5-0.5b model is used to score each seed problem, and the seed problem with the highest score is selected as the intermediate problem. The intermediate questions were rewritten using GPT4o according to the task type to generate multiple target questions corresponding to each chart-type test image. The initial labeled test data, consisting of each target question and its corresponding chart-type test image, is input into the Qwen2-VL-72b model to generate the answer information corresponding to each initial labeled test data, thereby obtaining the target test data.

4. The method according to claim 3, characterized in that, Before using the self-instruct model to generate multiple seed questions for each chart class test image, the method further includes: The initial chart-type test images obtained are all converted to RGB format to generate the first test image set; Remove low-quality images from the first test image set to generate a second test image set; the low-quality images include low-resolution, over-compressed, blank or invalid images, and images with duplicate content. Median filtering was used to remove noise from the images in the second test image set, generating a chart-type test image set in the test dataset.

5. The method according to claim 1, characterized in that, Before inputting each chart-type test image and the corresponding question-answer pair from the fill-in-the-blank test dataset into the large model to be evaluated, in order to obtain the output results of the fill-in-the-blank model corresponding to each chart-type test image, the method further includes: Obtain the large model to be evaluated by using either an API interface or model weights.

6. The method according to claim 1, characterized in that, The types of chart test images include stacked bar charts, complex line charts, scatter plots, pie charts, regular line charts, grouped bar charts, regular bar charts, 3D bar charts, bubble charts, and mixed charts.

7. The method according to claim 1, characterized in that, The categories of low-level tasks include reasoning, anomaly, distribution, correlation, range determination, sorting, filtering, data retrieval, extrema, and clustering.

8. The method according to claim 2, characterized in that, Advanced tasks include chart classification, opinion suggestions, chart redrawing, chart transformation, chart summarization, chart filling assistance, text generation, chart creation assistance, multi-chart Q&A, chart multi-turn dialogue, chart OCR recognition, and information location.

9. A non-transitory computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements an evaluation and testing method for a multimodal graph question-answering large model as described in any one of claims 1 to 8.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements an evaluation and testing method for a multimodal graph question-answering large model as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Question and answer pair evaluation data generation method and device, computer equipment and storage medium

    CN116775843A

  • Evaluation method and system for large model content security capability

    CN118035711A