Chart processing method based on multi-modal large model

By obtaining and filtering the context key information of the chart, building the target problem associated with the initial problem, and inputting a multimodal big model to generate answers, the problem of comprehensive expression of graph views in subjective question-and-answer tasks is solved, and the accuracy and efficiency of the output results are improved.

CN120088802AActive Publication Date: 2025-06-03北京中科闻歌科技股份有限公司
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510250211.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-03
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The lack of comprehensive solutions in subjective question-and-answer tasks, especially in the comprehensive expression of graphical views, has led to the impact of the accuracy and efficiency of model output results.

Method used

By obtaining context-critical information for the target chart, filtering information associated with the initial question type and target chart type entered by the user, and building the target question, inputting a multimodal big model to generate answers.

Benefits of technology

The accuracy and efficiency of the model output results are improved. By filtering out the context information related to the initial problem, the interference of redundant information is reduced and the complexity of model reasoning is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088802A_ABST
    Figure CN120088802A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric digital data processing, in particular to a chart processing method based on a multi-modal large model. The method comprises the following steps: acquiring context key information of a target chart; obtaining a question type of an initial question according to the initial question input by a user; screening information associated with the question type of the initial question and the type of a target chart from context key information of the target chart according to the question type of the initial question and the type of the target chart; constructing a target question corresponding to the initial question according to the screened information associated with the question type of the initial question and the type of a target chart and the initial question; and inputting the target question and the target chart into a multi-modal large model, and determining the output of the multi-modal large model as an answer corresponding to the initial question. The accuracy and efficiency of the model output result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electrical digital data processing, and particularly to a chart processing method based on a multimodal large model. Background Art

[0002] The multimodal chart question-answering task is an important research direction in the field of multimodal artificial intelligence, aiming to solve natural language problems through the visual and language fusion understanding of chart data. This task includes two major categories: objective question-answering and subjective question-answering. Objective question-answering focuses on extracting specific data points, numerical values, or classification information, while subjective question-answering requires the model to have high-level interpretation and reasoning abilities, such as trend summarization, category analysis, and comprehensive expression of chart viewpoints. However, current research mainly focuses on objective tasks, and there is still a lack of a comprehensive solution for subjective question-answering tasks, with the following problems: There is a lack of a context information extraction mechanism for task requirements, resulting in inaccurate context information input to the model. For example, when the context information input to the model includes noise and redundant information, it will interfere with the reasoning process of the model; when the context information input to the model lacks information related to the task, it will increase the complexity of the model's reasoning, ultimately affecting the accuracy and efficiency of the model's output results. Summary of the Invention

[0003] The object of the present invention is to provide a chart processing method based on a multimodal large model to improve the accuracy and efficiency of the model's output results.

[0004] According to the present invention, there is provided a chart processing method based on a multimodal large model, the method comprising the following steps:

[0005] S100, obtaining the context key information of the target chart; the context key information of the target chart includes the title information, legend information, and axis information of the target chart.

[0006] S200, obtaining the question type of the initial question according to the user input.

[0007] S300, screening the information associated with the question type of the initial question and the type of the target chart from the context key information of the target chart according to the question type of the initial question and the type of the target chart; the information associated with the question type of the initial question and the type of the target chart includes at least one of the title information, legend information, and axis information.

[0008] S400, constructing a target question corresponding to the initial question according to the screened information associated with the question type of the initial question and the type of the target chart and the initial question;

[0009] S500. Input the target problem and the target chart into the multimodal large model, and determine the output of the multimodal large model as the answer corresponding to the initial problem.

[0010] The present invention has at least the following beneficial effects compared with the prior art:

[0011] The present invention first obtains the title information, legend information, and axis information of the target chart, screens the information associated with the problem type of the initial problem and the type of the target chart from this information according to the problem type of the initial problem input by the user and the type of the target chart, and constructs a target problem corresponding to the initial problem based on this information and the initial problem; compared with the initial problem, the target problem of the present invention also includes the screened information associated with the problem type of the initial problem and the type of the target chart, and this information can be used to assist the multimodal large model to quickly understand the target chart, and the redundant information not associated with the problem type of the initial problem is screened out from this information. Therefore, the multimodal large model can quickly understand the target chart according to the information attached to the target problem, and then quickly and accurately answer the initial problem, achieving the purpose of improving the accuracy and efficiency of the output result. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0013] Figure 1 It is a flowchart of a chart processing method based on a multimodal large model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0015] According to this embodiment, as Figure 1 shown, a chart processing method based on a multimodal large model is provided, and the method includes the following steps:

[0016] S100. Obtain the context key information of the target chart; the context key information of the target chart includes the title information, legend information, and axis information of the target chart.

[0017] In this embodiment, the target chart includes a title, a legend, and coordinate axes. The title information, legend information, and coordinate axis information of the target chart are all known information. Among them, the title, as the theme of the chart, provides the overall background semantics for the model, such as the data source, time range, or analysis direction; the legend contains the correspondence between categories and colors and is an important information source for classification inference tasks. For example, the legend can illustrate the specific category or grouping represented by a certain color; the coordinate axis information can provide the unit and range of the data. For example, the x-axis represents the year and the y-axis represents the percentage change, which can directly guide numerical queries or trend descriptions.

[0018] S200. Obtain the question type of the initial question according to the initial question input by the user.

[0019] As an optional specific implementation manner, the question type of the initial question is one of the numerical query type, data point comparison type, trend summary type, semantic description type, and category inference type.

[0020] As a preferred specific implementation manner, S200 includes:

[0021] S210. Obtain the question type field list A; A = {A 1 , A 2 , …, A i , …, A n}, where A i is the i-th record included in A, the value range of i is from 1 to n, and n is the preset number of question types; A i = (T i , (d i,1 , d i,2 , …, d i,r(i) , …, d i,R(i) ))), where T i is the question type included in A i , d i,r(i) is the r(i)-th field included in A i , the value range of r(i) is from 1 to R(i), and R(i) is the number of fields included in A i ; the question type of any d i,r(i) is T i ; the T i included in different A i are different, and the fields with a smaller position serial number in (d i,1 , d i,2 , …, d i,r(i) , …, d i,R(i) ) have a higher priority than the fields with a larger position serial number, and the fields with the same position serial number included in different A i have the same priority.

[0022] In this embodiment, A is a pre-constructed list, and each T i corresponding field is known. For example, when T i is of the trend summary type, T i corresponding fields include trend summary, trend analysis, trend generalization, trend description, trend movement summary, trend movement analysis, trend movement generalization, and trend movement description, etc.

[0023] In this embodiment, the priority of different fields corresponding to the same problem type is related to the number of times the corresponding field is used when the user asks questions of the corresponding problem type within the historical time period. The more times a certain field is used when the user asks questions of a certain problem type, the higher the position serial number of this field in the record belonging to this problem type. It should be understood that d i,r(i) is at the position serial number r(i) in A i . For example, d i,1 is at the position serial number 1 in A i , and d i,1 is the field with the highest position serial number in A i .

[0024] S220, match the fields included in A with the initial problem in the order of decreasing priority of the fields until a matching field is obtained.

[0025] In this embodiment, the fields are matched with the initial problem in the order of decreasing priority of the fields. Since the fields with higher priority correspond to the fields with higher occurrence probabilities, therefore, matching the fields with the initial problem in the order of decreasing priority of the fields is beneficial to improving the efficiency of obtaining the matching fields.

[0026] In this embodiment, the priorities of the fields with the same position serial number included in different A i are the same. In the process of matching the fields included in A with the initial problem in the order of decreasing priority of the fields, it can be carried out in the order of first the field with the highest priority included in A 1 , then the field with the highest priority included in A 2 , …, and finally the field with the highest priority included in A n to perform the first round of matching. If all the matches fail, then perform the second round of matching in the order of the field with the second highest priority included in A 1 , then the field with the second highest priority included in A 2 , …, and finally the field with the second highest priority included in A n , and so on, until a matching field is obtained. It should be understood that if there is content in the initial problem that is the same as a certain field, then it is determined that this field is the matching field.

[0027] S230. Determine the problem type included in the record to which the matched field belongs as the problem type of the initial problem.

[0028] Based on S210 - S230, the problem type of the initial problem can be obtained quickly and accurately.

[0029] S300. Screen the information associated with the problem type of the initial problem and the type of the target chart from the context key information of the target chart according to the problem type of the initial problem and the type of the target chart; the information associated with the problem type of the initial problem and the type of the target chart includes at least one of title information, legend information, and axis information.

[0030] As a preferred specific implementation manner, S300 includes:

[0031] S310. Obtain the problem type - related information list B; B = {B 1 , B 2 , …, B j , …, B m}, where B j is the j - th record included in B, and the value range of j is from 1 to m, where m is the number of records included in B; B j =(T j , U j , E j ), where T j is the problem type included in B j , U j is the chart type included in B j , different B j include different T j or different U j or both T j and U j are different, and E j is the set of information types corresponding to T j and U j in B j , and E j includes at least one of the title information type, legend information type, and axis information type.

[0032] In this embodiment, B is a pre - constructed list, and the corresponding E j for each T j and U j is known, and E j can be determined according to experience. It should be understood that if E j only includes the title information type and the axis information type, it means that the multimodal large model is answering questions about charts of type U j regarding the type Tj When dealing with the problem of j only includes legend information type and axis information type, indicating that the multimodal large model, when answering questions of type U j for a chart of type T j When dealing with the problem of j includes title information type, legend information, and axis information type at the same time, indicating that the multimodal large model, when answering questions of type U j for a chart of type T j When dealing with the problem of , it is necessary to pay attention to title information, legend information, and axis information at the same time.

[0033] As a specific implementation, the question types include numerical query type, data point comparison type, trend summary type, semantic description type, category reasoning type, etc.

[0034] S320, match the question type of the initial question and the type of the target chart with the question types and chart types included in B.

[0035] As a specific implementation, sequentially match the question type of the initial question and the type of the target chart with the B 1 、B 2 、…、B n included in B until the question type and chart type included in a certain record included in B are the same as the question type of the initial question and the type of the target chart respectively, and this record is the matching record.

[0036] S330, determine the set of information types included in the matching record as the target set.

[0037] S340, if the target set includes the title information type, determine the title information in the context key information of the target chart as the information associated with the question type of the initial question and the type of the target chart; if the target set includes the legend information type, determine the legend information in the context key information of the target chart as the information associated with the question type of the initial question and the type of the target chart; if the target set includes the axis information type, determine the axis information in the context key information of the target chart as the information associated with the question type of the initial question and the type of the target chart.

[0038] In this embodiment, all the information associated with the question type of the initial question and the type of the target chart is taken as the information associated with the question type of the initial question and the type of the target chart obtained by screening.

[0039] Based on S310 - S340, the information screened out and associated with the problem type of the initial problem and the type of the target chart is relatively relevant to the initial problem, and can be used to assist the multi - modal large - model in quickly understanding the target chart. Moreover, the redundant information not associated with the problem type of the initial problem is screened out from this information.

[0040] S400, construct a target problem corresponding to the initial problem according to the information screened out and associated with the problem type of the initial problem and the type of the target chart, and the initial problem.

[0041] As a preferred specific implementation, S400 includes:

[0042] S410, obtain the type of the target chart.

[0043] Optionally, the type of the target chart is one of the bar chart type, line chart type, pie chart type, and scatter plot type.

[0044] S420, obtain a matching template text according to the type of the target chart and the problem type of the initial problem; the matching template text includes several sub - regions to be embedded, and the several sub - regions to be embedded include a sub - region for embedding the initial problem and a sub - region for embedding the information screened out and associated with the problem type of the initial problem and the type of the target chart.

[0045] In this embodiment, the template text is pre - constructed, and multiple template texts are pre - constructed. According to the type of the target chart and the problem type of the initial problem, a matching template text can be uniquely determined. If the problem type of the initial problem corresponds to E j Composed of the title information type, legend information type, and axis information type, then there are 4 sub - regions to be embedded in this matching template text, which are sub - regions for embedding title information, legend information, axis information, and the initial problem respectively; if the problem type of the initial problem corresponds to E j Composed of the title information type and axis information type, then there are 3 sub - regions to be embedded in this matching template text, which are sub - regions for embedding title information, axis information, and the initial problem respectively.

[0046] S430, embed the information screened out and associated with the problem type of the initial problem and the type of the target chart, and the initial problem into the matching template text, and determine the embedded result as the target problem corresponding to the initial problem.

[0047] In this embodiment, it is known which information or initial question each sub-region to be embedded in the matched template corresponds to, and the embedding can be performed according to this corresponding relationship during the embedding process. For example, if there is a specific corresponding relationship between a sub-region to be embedded in the matched template and the title information, then the title information of the target chart can be embedded in this sub-region to be embedded during the embedding process.

[0048] S500, input the target question and the target chart into the multimodal large model, and determine the output of the multimodal large model as the answer corresponding to the initial question.

[0049] In this embodiment, the whole composed of the target question and the target chart is the prompt of the multimodal large model.

[0050] Those skilled in the art know that any multimodal large model in the prior art falls within the protection scope of the present invention; optionally, the multimodal large model is Qwen2-VL or InternLM-XComposer-2.5.

[0051] As a first optional specific implementation manner, there is no need to specially train the multimodal large model for the chart question-answering task, and the multimodal large model can directly generate an answer according to the target question and the target chart. As a second optional specific implementation manner, the multimodal large model is enhanced trained to improve the adaptability of the multimodal large model to the chart question-answering task.

[0052] In this embodiment, the title information, legend information, and axis information of the target chart are first obtained, and the information associated with the question type of the initial question and the type of the target chart is screened from these information according to the question type of the initial question input by the user and the type of the target chart, and the target question corresponding to the initial question is constructed based on this information and the initial question; compared with the initial question, the target question in this embodiment also includes the screened information associated with the question type of the initial question and the type of the target chart, and this information can be used to assist the multimodal large model to quickly understand the target chart, and the redundant information not associated with the question type of the initial question is screened out from this information. Therefore, the multimodal large model can quickly understand the target chart according to the information attached to the target question, and then quickly and accurately answer the initial question, achieving the purpose of improving the accuracy and efficiency of the output result.

[0053] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.

Claims

1. A diagram processing method based on a multimodal large model, characterized in that: The method comprises the following steps: S100, obtaining context key information of a target chart; the context key information of the target chart includes title information, legend information and coordinate axis information of the target chart; S200, obtaining the question type of the initial question according to the initial question input by the user; S300, filtering information associated with the question type of the initial question and the type of the target chart from context key information of the target chart according to the question type of the initial question and the type of the target chart; the information associated with the question type of the initial question and the type of the target chart includes at least one of title information, legend information and coordinate axis information; S400, constructing a target question corresponding to the initial question based on the information associated with the question type of the initial question and the type of the target graph obtained by screening and the initial question; S500, inputting the target question and the target graph into the multimodal large model, and determining the output of the multimodal large model as the answer corresponding to the initial question.

2. The multimodal graph processing method based on contextual key information according to claim 1, characterized in that: S200 includes: S210, obtain a question type field list A; A={A1,A2,…,A i ,…,A n }, A i is the i-th record included in A, the value range of i is 1 to n, and n is the number of preset question types; A i =(T i ,(d i,1 ,d i,2 ,…,d i,r(i) ,…,d i,R(i) )), T i A i Types of questions included, i,r(i) A i The r(i)th field is included, and the value range of r(i) is 1 to R(i), and R(i) is A i The number of fields to include, either i,r(i) The question types are all T i Different A i Included T i Different, (d i,1 ,d i,2 ,…,d i,r(i) ,…,d i,R(i) ) has a higher priority than the field with the position number later. i Fields with the same position number have the same priority; S220, matching the fields included in A with the initial question in descending order of priority of the fields until a matching field is obtained; S230: Determine the question type included in the record to which the matched field belongs as the question type of the initial question.

3. The multimodal graph processing method based on contextual key information according to claim 1, characterized in that: S300 includes: S310, obtaining a list B of information related to the question type; B={B1, ​​B2,…, B j ,…,B m }, B j is the jth record included in B, where j ranges from 1 to m, and m is the number of records included in B; j =(T j ,U j ,E j ), T j For B j Types of questions included, j For B j Included chart types, different B j Included T j Different or U j Different or T j and U j All are different, E j For B j Included with T j and U j A set of information types with corresponding relationships, E j At least one of the following information types: title information type, legend information type and axis information type; S320, matching the question type of the initial question and the type of the target chart with the question type and chart type included in B; S330, determining a set of information types included in the matching records as a target set; S340, if the target set includes a title information type, the title information in the context key information of the target chart is determined as information associated with the question type of the initial question and the type of the target chart; if the target set includes a legend information type, the legend information in the context key information of the target chart is determined as information associated with the question type of the initial question and the type of the target chart; if the target set includes an axis information type, the axis information in the context key information of the target chart is determined as information associated with the question type of the initial question and the type of the target chart.

4. The multimodal graph processing method based on contextual key information according to claim 1, characterized in that: S400 includes: S410, obtaining the type of the target chart; S420, obtaining a matching template text according to the type of the target chart and the type of the initial question; the matching template text includes a plurality of sub-regions to be embedded, the plurality of sub-regions to be embedded including a sub-region for embedding the initial question and a sub-region for embedding screened information associated with the type of the initial question and the type of the target chart; S430, embedding the screened information associated with the question type of the initial question and the type of the target chart and the initial question into the matching template text, and determining the embedded result as the target question corresponding to the initial question.

5. The multimodal graph processing method based on contextual key information according to claim 1, characterized in that: The multimodal large model is Qwen2-VL.

6. The multimodal graph processing method based on contextual key information according to claim 1, characterized in that: The multimodal large model is InternLM-XComposer-2.

5.

7. The multimodal graph processing method based on contextual key information according to claim 1, characterized in that: The type of the target chart is one of a bar chart type, a line chart type, a pie chart type, and a scatter chart type.

8. The multimodal graph processing method based on contextual key information according to claim 1, characterized in that: The question type of the initial question is one of a numerical query type, a data point comparison type, a trend summary type, a semantic description type, and a category reasoning type.

Citation Information

Patent Citations

  • Data output method and device

    CN112597276A

  • Data processing method, device and equipment based on multi-modal model

    CN116204726A

  • Data processing system for acquiring target task data set

    CN116561390A

  • Information processing method and device, equipment, storage medium and computer program product

    CN118626626A

  • Data processing method and device for knowledge questions and answers, medium and equipment

    CN118690027A