General chart multi-modal model based on pre-training and instruction fine-tuning

Through pre-training and instruction fine-tuning methods, a large-scale chart dataset ChartBench is constructed, which solves the difficulty of chart models in explaining the relationship between charts and structured texts, and achieves efficient generalization in multi-task chart tasks.

WO2025138696A1PCT designated stage expired Publication Date: 2025-07-03SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/103526
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-27
Filing Date
2024-07-04
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing chart understanding models have difficulties in accurately interpreting the relationship between charts and structured texts, and the training data lacks annotations of visual elements and mathematical reasoning, resulting in poor generalization and requires fine-tuning for specific tasks.

Method used

Through a two-stage method of pre-training and instruction fine-tuning, a large-scale chart dataset ChartBench is constructed, and the ChartAssistant model is used to pre-train the chart-to-table translation task, and then multi-task instruction adjustment is performed to improve the generalization ability of the model.

Benefits of technology

Without the need for specific task fine-tuning, the ChartAssistant model performs well in multiple chart-related tasks, improving the generalization and task adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024103526_03072025_PF_FP_ABST
    Figure CN2024103526_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present invention is a general chart multi-modal model based on pre-training and instruction fine-tuning. The general chart multi-modal model based on pre-training and instruction fine-tuning comprises: acquiring a sample image and a corresponding instruction and response thereof; on the basis of the sample image and the corresponding instruction and response thereof, performing pre-training of a chart-to-table translation task, so as to obtain a basic general chart multi-modal model; constructing a large chart data set by means of instruction tracking data which is collected from various chart-related tasks; on the basis of the large chart data set, performing multi-task instruction tuning on the basic general chart multi-modal model, so as to obtain a target general chart multi-modal model. The general chart multi-modal model based on pre-training and instruction fine-tuning in the present invention improves model generalization.
Need to check novelty before this filing date? Find Prior Art

Description

A general graph multimodal model based on pre-training and instruction fine-tuning Technical Field

[0001] The present invention relates to the field of computer science and technology, and in particular to a general graph multimodal model and device based on pre-training and instruction fine-tuning. Background Art

[0002] To pursue general graph reasoning and understanding, pre-trained visual language models for graph-related tasks have been proposed. For example, Matcha and UniChart, both of which have been fine-tuned through multi-task instructions and task-specific fine-tuning, have demonstrated strong performance on multiple downstream tasks. Matcha, pre-trained on mathematical reasoning and graph data extraction tasks, has demonstrated strong performance in graph question answering and summarization. UniChart, through multi-task instructions for graph question answering, graph summarization, graph-to-table translation, and open-ended graph question answering, has become a versatile model for a variety of graph tasks.

[0003] Chart comprehension is challenging due to complex visual markup (lines, bars, and symbols), implicit numerical information, and complex spatial relationships between elements (axes and labels). Interpreting charts requires expertise, spatial reasoning, and numerical understanding. Advanced general-purpose multimodal models such as GPT-4V(ision), trained on natural images, struggle with chart-related tasks due to their specific complexity and unique relationships. Although recent multimodal literacy models have achieved impressive results on various document-level tasks, they still face difficulties in accurately answering questions related to charts.

[0004] In summary, existing models fail to explicitly align diagrams with related structured text tables, which is crucial for interpreting the relationships between elements in a diagram. Furthermore, existing training data lacks image-text annotations designed to improve the model's understanding of visual elements and mathematical reasoning, as well as annotations for specialized diagram types such as boxplots. Consequently, existing diagram-based models generalize poorly and require task-specific fine-tuning to achieve promising results on a variety of downstream tasks.

[0005] Summary of the Invention

[0006] In view of this, the present invention provides a general graph multimodal model and device based on pre-training and instruction fine-tuning to solve the above problems.

[0007] A first aspect of the present invention provides a general chart multimodal model based on pre-training and instruction fine-tuning, comprising: obtaining sample images and their corresponding instructions and responses; performing pre-training on a chart-to-table translation task based on the sample images and their corresponding instructions and responses to obtain a basic general chart multimodal model; constructing a large chart dataset by using instruction tracking data collected from various chart-related tasks; and performing multi-task instruction adjustment on the basic general chart multimodal model based on the large chart dataset to obtain a target general chart multimodal model.

[0008] In another implementation of the present invention, the loss function of pre-training is expressed as:

[0009] in, For the input graph, For the corresponding instruction, is the current prediction tag, is all previous response tokens.

[0010] In another implementation of the present invention, a large-scale chart dataset is constructed by collecting instruction tracing data from various chart-related tasks, including: collecting instruction tracing data from various chart-related tasks; adjusting the instruction tracing data into a unified format to construct a large-scale chart dataset, wherein each question or instruction is associated with an image and its corresponding answer.

[0011] In another embodiment of the present invention, the large chart dataset includes: instruction tracking data on various topics related to chart-to-table conversion, thought chain annotations for generating chart mathematical question-answering tasks, chart reference question-answering tasks, and professional types of charts such as radar charts and box plots.

[0012] In another implementation of the present invention, the training loss function of multi-task adjustment is expressed as:

[0013] where Ω is the instruction trace dataset from all tasks in the large graph dataset, and θ is the learnable weights initialized from the checkpoints in the pre-training phase.

[0014] The present invention proposes a universal chart multimodal model based on pre-training and instruction fine-tuning. The present invention proposes ChartBench, a large-scale chart dataset with the most types and the largest number. A novel training strategy is used to train ChartAssistant on ChartBench, achieving excellent results without the need for fine-tuning. Compared with previous methods, this method first uses image-to-table pre-training and then uses multi-tasks for instruction fine-tuning, which improves the generalization of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. By reading the detailed description of the embodiments below, the advantages and benefits of the solutions will become clear to those skilled in the art. The drawings are only for the purpose of illustrating preferred embodiments and are not to be considered as limiting the present invention. In the drawings:

[0016] FIG1 is a flowchart illustrating steps for training a general graph multimodal model based on pre-training and instruction fine-tuning according to an embodiment of the present invention.

[0017] FIG2 is a schematic diagram of a general chart multimodal model according to an embodiment of the present invention.

[0018] FIG3 is a schematic block diagram of a general chart multimodal model training process according to an embodiment of the present invention. DETAILED DESCRIPTION

[0019] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and detailedly described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments in the embodiments of the present invention should fall within the scope of protection of the embodiments of the present invention.

[0020] FIG1 is a flowchart of a general graph multimodal model based on pre-training and instruction fine-tuning provided by an embodiment of the present invention. As shown in FIG1 , this embodiment mainly includes the following steps:

[0021] S101: Obtain a sample image and its corresponding instructions and responses.

[0022] S102. Based on the sample images and their corresponding instructions and responses, pre-training of the chart-to-table translation task is performed to obtain a basic general chart multimodal model.

[0023] Exemplarily, ChartAssistant (a general chart multimodal language model via chart-to-table pre-training and fine-tuning with multi-task instructions) is first pre-trained on the chart-to-table translation task. This task involves parsing a chart and generating a table, which has similarities to dense captioning of natural images, allowing the model to interpret the elements and relationships in the chart. Similar to the role of image captions in training multimodal models, chart-to-table conversion helps align the chart with its structured text.

[0024] Preferably, each image X V There are corresponding instructions X q and response Y q, these image-text pairs are fed into the model, and the goal is to minimize the cross entropy loss of predicting the next tag. Given a graph The goal is to convert the chart into a table in text form In the instruction The superscript c2t represents the instruction-following data conversion task from a graph to a table. Through pre-training in the equation, the graph is aligned with its structured text table, enabling the model to understand the elements in the graph and their relationships.

[0025] S103. Build a large graph dataset by collecting instruction trace data from various graph-related tasks.

[0026] S104 , performing multi-task instruction adjustment on the basic general chart multimodal model based on the large chart dataset to obtain a target general chart multimodal model.

[0027] For example, after pre-training, ChartBench (a large chart dataset) is used for multi-task instruction tuning. This two-stage training approach enables ChartAssistant to achieve excellent performance on a range of chart-related tasks without the need for task-specific fine-tuning. As shown in Figure 2, the model can handle a range of tasks such as summarization, question answering, math question answering, and image-to-table conversion.

[0028] It should be understood that multi-task instruction tuning is a technique used to enhance the ability of language models to understand and follow natural language instructions. It has successfully improved the zero-shot and few-shot generalization capabilities of models such as InstructGPT and FLAN-T5 in NLP tasks.

[0029] In practice, the datasets from five chart-related tasks are organized into a unified format, where each question or instruction is associated with an image and its corresponding answer. The data for each task is then randomly mixed in a certain proportion, followed by end-to-end multi-task instruction fine-tuning. At this stage, all instruction trace data from the five tasks is placed in ChartBench, and a single model can be used to solve all tasks. Through multi-task instruction fine-tuning, ChartAssistant demonstrates strong performance on all tasks without the need for task-specific fine-tuning, which is necessary for previous chart-based models.

[0030] The present invention proposes a universal chart multimodal model based on pre-training and instruction fine-tuning. The present invention proposes ChartBench, a large-scale chart dataset with the most types and the largest number. A novel training strategy is used to train ChartAssistant on ChartBench, achieving excellent results without the need for fine-tuning. Compared with previous methods, this method first uses image-to-table pre-training and then uses multi-tasks for instruction fine-tuning, which improves the generalization of the model.

[0031] In another implementation of the present invention, the loss function of pre-training is expressed as:

[0032] in, For the input graph, For the corresponding instruction, is the current prediction tag, is all previous response tokens.

[0033] In another implementation of the present invention, a large-scale chart dataset is constructed by collecting instruction tracing data from various chart-related tasks, including: collecting instruction tracing data from various chart-related tasks; adjusting the instruction tracing data into a unified format to construct a large-scale chart dataset, wherein each question or instruction is associated with an image and its corresponding answer.

[0034] In another embodiment of the present invention, the large chart dataset includes: instruction tracking data on various topics related to chart-to-table conversion, thought chain annotations for generating chart mathematical question-answering tasks, chart reference question-answering tasks, and professional types of charts such as radar charts and box plots.

[0035] As an example, we first construct ChartBench by collecting instruction trace data from various diagram-related tasks. To address the limitations of existing diagram-based benchmarks, we introduce several modifications to improve the quality of data annotation: we add instruction trace data on various topics involving diagram-to-table conversion, which we find helpful in aligning diagrams and related structured text; we generate thought chain annotations for diagram math question-answering tasks to improve mathematical reasoning; we create diagram reference question-answering tasks to enhance understanding of visual elements and their relationships; and we include specialized types of charts such as radar charts and boxplots to improve generalization capabilities. Overall, compared to previous benchmarks, ChartBench contains a larger corpus of instruction trace data, encompasses a wider range of diagram-related tasks and types, and has more comprehensive data annotations.

[0036] In another implementation of the present invention, the training loss function of multi-task adjustment is expressed as:

[0037] where Ω is the instruction trace dataset from all tasks in the large graph dataset, and θ is the learnable weights initialized from the checkpoints in the pre-training phase.

[0038] In another implementation of the present invention, as shown in Table 1, experiments were conducted on datasets such as ChartQA. The experimental results show that the performance of the present invention is better than the previous technologies Matcha and Unichart.

[0039] Table 1. Comparison of experimental results of this method with other methods in the prior art

[0040] Thus far, specific embodiments of the present invention have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing may be advantageous.

[0041] It should be noted that all directional indications in the embodiments of the present invention (such as up, down, left, right, back, etc.) are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0042] In the description of the present invention, the terms "first" and "second" are used solely to facilitate description of different components or names and should not be construed as indicating or implying a sequential relationship, relative importance, or implicitly specifying the quantity of the technical features being described. Therefore, features specified as "first" or "second" may explicitly or implicitly include at least one of such features.

[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0044] It should be noted that although the specific embodiments of the present invention are described in detail in conjunction with the accompanying drawings, this should not be construed as limiting the scope of protection of the present invention. Within the scope described by the claims, various modifications and variations that can be made by those skilled in the art without creative effort still fall within the scope of protection of the present invention.

[0045] The examples of the embodiments of the present invention are intended to briefly illustrate the technical features of the embodiments of the present invention so that those skilled in the art can intuitively understand the technical features of the embodiments of the present invention, and are not intended to improperly limit the embodiments of the present invention.

[0046] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A general chart multi-modal model based on pre-training and instruction fine-tuning, characterized in that including: Obtain sample images and their corresponding instructions and responses; Based on the sample images and their corresponding instructions and responses, perform pre-training for the chart-to-table translation task to obtain a basic general chart multi-modal model; Construct a large chart dataset through instruction trace data collected from various chart-related tasks; Based on the large chart dataset, perform multi-task instruction adjustment on the basic general chart multi-modal model to obtain a target general chart multi-modal model.

2. The model according to claim 1, wherein, The loss function of the pre-training is expressed as: Among them, For the input chart, Is the corresponding instruction, is the current predicted token, All previous response tokens.

3. The model according to claim 1, wherein The constructing of the large chart dataset through instruction trace data collected from various chart-related tasks includes: Collect instruction trace data from various chart-related tasks; Adjust the instruction trace data into a unified format to construct a large chart dataset, where each question or instruction is associated with an image and its corresponding answer.

4. The model according to claim 3, characterized in that, The large chart dataset includes: instruction trace data covering various topics related to chart-to-table conversion, thought chain annotations for generating chart math Q&A tasks, chart reference Q&A tasks, and professional types of charts such as radar charts and box plots.

5. The model according to claim 1, characterized in that, The training loss function for the multi-task adjustment is expressed as: Where Ω is the instruction trace dataset for all tasks in the large chart dataset, and θ is the learnable weight initialized from the checkpoint in the pre-training stage.

Citation Information

Patent Citations

  • Image-based table restoration model training method and table restoration method

    CN116152833A

  • Deep learning-based chart extraction method and system

    CN116563872A

  • General chart multi-modal model based on pre-training and instruction fine tuning

    CN117786407A

  • Apparatus for analyzing chart data using artificial intelligence and method thereof

    KR102560770B1

  • Automated data analytics methods for non-tabular data, and related systems and apparatus

    US20230067026A1