Chart key information extraction method based on multi-modal large model
Through the multimodal large model, the problem of lack of targeted chart information extraction in the existing technology is solved, and more efficient and accurate information extraction is achieved, improving the operation efficiency and inference accuracy of the model.
Patent Information
- Application Number
- CN202510250209.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The prior art lacks targetedness in the extraction of chart information, resulting in the introduction of redundant information, affecting the operation efficiency of the model and the accuracy of reasoning.
The graph key information extraction method based on the multimodal large model is adopted. By obtaining the initial prompt text, determining the target information type and position information, building the target prompt text, and inputting it into the multimodal large model, the key information of the chart is accurately extracted.
It improves the accuracy and efficiency of the extraction of key information in the chart, reduces the impact of redundant information, and enhances the operating efficiency and inference accuracy of the model.
Smart Images

Figure CN120088801A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and particularly to a method for extracting key information from charts based on a multi-modal large model. Background Art
[0002] The multi-modal chart understanding task is an important research direction in the field of artificial intelligence, and its goal is to accurately extract context information from charts to support the reasoning and decision-making of models. Currently, the OCR-based technology is one of the most commonly used enhancement methods in the multi-modal chart understanding task. The OCR technology provides all potential relevant information for the model by comprehensively extracting the text in the chart. However, this comprehensive extraction strategy has significant limitations in practical applications. First, OCR lacks pertinence in information extraction and often introduces a large amount of redundant information. For example, in a task that only requires the background information of the coordinate axis, parsing data points one by one not only increases the computational complexity but may also interfere with the actual task. These redundant information seriously affects the running efficiency of the model, especially when dealing with complex reasoning tasks, it interferes with the judgment of the model. How to accurately extract the key information of the chart according to the task requirements is an urgent problem to be solved. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for extracting key information from charts based on a multi-modal large model to accurately extract the key information of the chart according to the task requirements.
[0004] According to the present invention, there is provided a method for extracting key information from charts based on a multi-modal large model, and the method includes the following steps:
[0005] S100, obtaining an initial prompt text; the initial prompt text includes a title prompt text, a legend prompt text, and an axis prompt text. The title prompt text includes text for characterizing the extraction of the title and text for characterizing the output title information format. The legend prompt text includes text for characterizing the extraction of the legend and text for characterizing the output legend information format. The axis prompt text includes text for characterizing the extraction of the axis and text for characterizing the output axis information format.
[0006] S200, obtaining a target information type according to the type of the target chart and the type of the question input by the user; the target information type includes at least one of a title information type, a legend information type, and an axis information type.
[0007] S300, obtaining the position information of the target information corresponding to the target information type in the target chart.
[0008] S400. Construct a target prompt text based on the prompt text corresponding to the target information type in the initial prompt text and the position information of the target information corresponding to the target information type in the target chart.
[0009] S500. Input the target prompt text and the target chart into a multi-modal large model, and obtain the key information related to the question input by the user for the target chart according to the output of the multi-modal large model.
[0010] The present invention has at least the following beneficial effects compared with the prior art:
[0011] The initial text of the present invention includes prompt texts corresponding to the title, legend, and coordinate axes. The prompt texts include texts for characterizing the extraction of corresponding contents and texts for characterizing the formats that should be met when outputting the corresponding contents. The multi-modal large model can achieve structured output of the corresponding partial contents according to the prompt texts therein; on this basis, the present invention determines which prompt texts in the initial text are used as part of the target prompt text according to the type of the target chart and the type of the question input by the user. This part can characterize the extraction requirements and output formats of key information with a relatively high correlation with the question input by the user, which is beneficial to improving the accuracy of finally extracting key information and expressing key information in the present invention; and the target prompt text of the present invention also includes the position information of the target information corresponding to the target information type in the target chart. Based on this position information, the multi-modal large model can quickly locate, which is beneficial to improving the efficiency of the information output by the large model. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0013] Figure 1 It is a flowchart of a method for extracting key information of a chart based on a multi-modal large model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0015] According to this embodiment, as Figure 1As shown in the figure, a method for extracting key information from charts based on a multimodal large model is provided. The method includes the following steps:
[0016] S100, obtaining an initial prompt text; the initial prompt text includes a title prompt text, a legend prompt text, and an axis prompt text. The title prompt text includes text for characterizing the extraction of the title and text for characterizing the output title information format. The legend prompt text includes text for characterizing the extraction of the legend and text for characterizing the output legend information format. The axis prompt text includes text for characterizing the extraction of the axis and text for characterizing the output axis information format.
[0017] In this embodiment, the title is the theme of the chart, and the title provides the overall background semantics for the model, such as the source of the data, the time range, or the direction of the analysis. The legend contains the correspondence between categories and colors and is an important information source for classification inference tasks. For example, the legend can illustrate the specific category or grouping represented by a certain color. The axis information can provide the unit and range of the data. For example, the x-axis represents the year, and the y-axis represents the percentage change, which can directly guide numerical queries or trend descriptions.
[0018] In this embodiment, the initial prompt text is pre-constructed.
[0019] As a preferred specific implementation, the legend prompt text includes: identifying the legend type and its corresponding color, in the following format: {Type 1 Name (Color)}|{Type 2 Name (Color)}|…|{Type k Name (Color)}|…|{Type p Name (Color)}, where k is a variable, and the value range of k is from 1 to p, and p is the number of types in the legend. Based on this preferred specific implementation, the multimodal large model can accurately output legend information. As a specific implementation, the output of the multimodal large model for the legend part is: Energy Imports (Grey)|Agricultural Exports (Green).
[0020] As a preferred specific implementation, the axis prompt text includes: extracting the x-axis label and its values, in the following format: The x-axis represents: {x-axis label}, and its values are: {value 1}|{value 2}|…|{value f x}|…|{value g x}; extracting the y-axis label and its values, in the following format: The y-axis represents: {y-axis label}, and its values are: {value 1}|{value 2}|…|{value f y}|…|{value g y}; where f x and f y are both variables, the value range of f x is from 1 to g x , g x is the number of data corresponding to the x-axis, fy The value range is 1 to g y , g y is the number of data corresponding to the y-axis. Based on this preferred specific implementation, the multimodal large model can accurately output the coordinate axis information. As a specific implementation, the output of the multimodal large model for the coordinate axis part is: the x-axis represents: year, its value is 1990|1995|2000; the y-axis represents: proportion, its value is 0%|50%|100%.
[0021] S200, obtaining a target information type according to a type of a target chart and a type of a question input by a user; the target information type includes at least one of a title information type, a legend information type and a coordinate axis information type.
[0022] In this embodiment, the target chart includes a title, a legend and a coordinate axis. As a specific implementation, the type of the target chart is one of a bar chart type, a line chart type, a pie chart type and a scatter chart type.
[0023] As a specific implementation, the type of the question input by the user is one of a numerical query type, a data point comparison type, a trend summary type, a semantic description type, and a category reasoning type. As a preferred specific implementation, the process of obtaining the type of the question input by the user includes:
[0024] S201, obtain question type field list A; A={A 1 ,A 2 ,…,A i ,…,A n}, A i is the i-th record included in A, the value range of i is 1 to n, and n is the number of preset question types; A i =(T i ,(d i,1 ,d i,2 ,…,d i,r(i) ,…,d i,R(i) )), T i A i Types of questions included, i,r(i) A i The r(i)th field is included, and the value range of r(i) is 1 to R(i), and R(i) is A i The number of fields to include, either i,r(i) The question types are all T i Different A i Included T i Different, (d i,1 ,d i,2 ,…,d i,r(i) ,…,d i,R(i))The priority of the field with a smaller position serial number in A is higher than that of the field with a larger position serial number, and the priorities of fields with the same position serial number included in different A i are the same.
[0025] In this embodiment, A is a pre-constructed list, and each T i corresponding field is known. For example, when T i is of the trend summary type, T i corresponding fields include trend summary, trend analysis, trend generalization, trend description, trend movement summary, trend movement analysis, trend movement generalization, and trend movement description, etc.
[0026] In this embodiment, the priority of different fields corresponding to the same problem type is related to the number of times the corresponding field is used when the user asks questions of the corresponding problem type within the historical time period. The more times a certain field is used when the user asks questions of a certain problem type, the more forward the position serial number of this field in the record belonging to this problem type. It should be understood that d i,r(i) in A i has a position serial number of r(i). For example, d i,1 in A i has a position serial number of 1, and d i,1 is the field with the most forward position serial number in A i .
[0027] S202. Match the fields included in A with the question input by the user in the order of decreasing field priority until a matching field is obtained.
[0028] In this embodiment, when matching the fields with the question input by the user in the order of decreasing field priority, since the fields with higher priorities correspond to fields with higher occurrence probabilities, therefore, matching the fields with the question input by the user in the order of decreasing field priority is beneficial to improving the efficiency of obtaining matching fields.
[0029] In this embodiment, the priorities of fields with the same position serial number included in different A i are the same. During the process of matching the fields included in A with the question input by the user in the order of decreasing field priority, it can be carried out in the order of first matching the field with the highest priority included in A 1 , then the field with the highest priority included in A 2 , …, and finally the field with the highest priority included in A n for the first round of matching. If all fail to match, then in the order of the field with the second highest priority included in A 1 , then the field with the second highest priority included in A 2 , …, and finally the field with the second highest priority included in A nPerform a second-round match in the order of the fields with the highest priority included, and so on until a matching field is obtained. It should be understood that if the content identical to a certain field exists in the question input by the user, then that field is determined to be the matching field.
[0030] S203. Determine the question type included in the record to which the matching field belongs as the type of the question input by the user.
[0031] Based on S201 - S203, the type of the question input by the user can be obtained quickly and accurately.
[0032] As a preferred specific embodiment, S200 includes:
[0033] S210. Obtain the information type list C; C = {C 1 , C 2 , …, C h , …, C H}, where C h is the h-th record included in C, and the value range of h is from 1 to H, where H is the number of records included in C; C h =(C h,1 , C h,2 , C h,3 ), where C h,1 is the question type included in C h , C h,2 is the chart type included in C h , and different C h include different C h,1 or different C h,2 or different C h,1 and C h,2 are all different. C h,3 is the set of information types corresponding to C h and having a corresponding relationship with C h,1 and C h,2 , and C h,3 includes at least one of the title information type, the legend information type, and the coordinate axis information type.
[0034] In this embodiment, C is a pre-constructed list, and each C h,1 and C h,2 corresponding C h,3 is known, and C h,3 can be determined according to experience. It should be understood that if C h,3 only includes the title information type and the coordinate axis information type, it means that the key information for the multimodal large model to answer questions of type C h,1 about charts of type C h,2 is the title information and the coordinate axis information; if C h,3only includes legend information type and axis information type, indicating that for the multimodal large model in answering questions of type C h,1 in the chart of type C h,2 the key information is legend information and axis information; if C h,3 includes title information type, legend information type and axis information type, indicating that for the multimodal large model in answering questions of type C h,1 in the chart of type C h,2 the key information is title information, legend information and axis information.
[0035] As a specific implementation, the question types include numerical query type, data point comparison type, trend summary type, semantic description type, category reasoning type, etc.
[0036] S220, match the type of the target chart and the type of the question input by the user in C.
[0037] As a specific implementation, sequentially match the type of the target chart and the type of the question input by the user with the C 1 、C 2 、…、C H included in C until the question type and chart type included in a certain record included in C are respectively the same as the question type input by the user and the type of the target chart, and this record is the matching record.
[0038] S230, determine the information type included in the matching record as the target information type.
[0039] Based on S210 - S230, the target information type can be determined quickly and accurately.
[0040] S300, obtain the position information of the target information corresponding to the target information type in the target chart.
[0041] It should be understood that the target information corresponding to the title information type is the title information, the target information corresponding to the legend information type is the legend information, and the target information corresponding to the axis information type is the axis information. As a specific implementation, the position information of the title, legend and axis in the target chart is known in advance and is the information input in advance.
[0042] S400, construct the target prompt text according to the prompt text corresponding to the target information type in the initial prompt text and the position information of the target information corresponding to the target information type in the target chart.
[0043] As a specific implementation, S400 includes:
[0044] S410, if the target information type includes the title information type, append the title prompt text and the position information of the target information corresponding to the title information type in the target chart to the preset specified text; if the target information type includes the legend information type, append the legend prompt text and the position information of the target information corresponding to the legend information type in the target chart to the preset specified text; if the target information type includes the axis information type, append the axis prompt text and the position information of the target information corresponding to the axis information type in the target chart to the preset specified text.
[0045] In this embodiment, the preset specified text is pre-constructed. Optionally, it includes three parts. The first part is the beginning of the specified text, which is fixed; the second part is blank and used to write the appended content; the third part is the end of the specified text, which is fixed. When writing to the second part, the order of title first, then legend, and finally axis can be followed (if there is content related to title, legend, and axis at the same time).
[0046] It should be understood that if the target information type includes the title information type and the legend information type, append the title prompt text, the position information of the target information corresponding to the title information type in the target chart, the legend prompt text, and the position information of the target information corresponding to the legend information type in the target chart to the preset specified text; if the target information type includes the title information type, the legend information type, and the axis information type, append the title prompt text, the position information of the target information corresponding to the title information type in the target chart, the legend prompt text, the position information of the target information corresponding to the legend information type in the target chart, the axis prompt text, and the position information of the target information corresponding to the axis information type in the target chart to the preset specified text; and so on.
[0047] S420, determine the specified text after appending as the target prompt text.
[0048] Based on S410 - S420, the target prompt text can represent the extraction requirements, output format, and position information of key information with a relatively high relevance to the problem input by the user.
[0049] S500, input the target prompt text and the target chart into the multi-modal large model, and obtain the key information related to the problem input by the user for the target chart according to the output of the multi-modal large model.
[0050] Those skilled in the art know that any multi-modal large model in the prior art falls within the protection scope of the present invention; optionally, the multi-modal large model is Qwen2-VL or OpenLLaMA.
[0051] As a first optional specific implementation, there is no need to perform new training on the multi-modal large model specifically for the chart extraction task, and the multi-modal large model can directly output according to the target prompt text and the target chart. As a second optional specific implementation, the multi-modal large model is enhanced trained to improve the adaptability of the multi-modal large model to the chart extraction task.
[0052] As an optional specific implementation, the output of the multi-modal large model is determined as the key information related to the question input by the user for the target chart.
[0053] As a preferred specific implementation, obtaining the key information related to the question input by the user for the target chart according to the output of the multi-modal large model includes:
[0054] S510, extract the core keywords in the question input by the user to obtain a set of core keywords.
[0055] As a specific implementation, first tokenize the question input by the user, perform part-of-speech tagging on each word, and determine the words with the preset type of part-of-speech as the core keywords; optionally, the preset type of part-of-speech is nouns and verbs.
[0056] S520, perform sentence splitting on the output of the multi-modal large model to obtain a set of sentences.
[0057] S530, obtain the relevance of each sentence in the set of sentences to the set of core keywords; where the relevance of any sentence in the set of sentences to the set of core keywords is the number of times the core keywords in the set of core keywords appear in the sentence.
[0058] In this embodiment, the number of times the core keywords in the set of core keywords appear in the sentence is the sum of the number of times each core keyword in the set of core keywords appears in the sentence.
[0059] S540, sort the sentences in the output of the multi-modal large model in descending order of relevance.
[0060] In this embodiment, the higher the relevance of a certain sentence in the output of the multi-modal large model to the set of core keywords, the higher the relevance of the sentence to the question input by the user, and the higher the probability that the sentence is key information.
[0061] S550, determine the sorted result as the key information related to the question input by the user for the target chart.
[0062] Based on S510 - S550, this embodiment realizes the sorting of the information in the output of the multi-modal large model, so that the information with a stronger relevance to the question input by the user in the output of the multi-modal large model is presented more prominently.
[0063] The initial text of this embodiment includes a title, a legend, and prompt text corresponding to the coordinate axes. The prompt text includes text for characterizing the extraction of corresponding content and text for characterizing the format that should be followed when outputting the corresponding content. Based on the prompt text therein, the multimodal large model can achieve structured output of corresponding parts of the content; on this basis, this embodiment determines which prompt texts in the initial text are used as part of the target prompt text according to the type of the target chart and the type of the question input by the user. This part can characterize the extraction requirements and output format of key information that is highly relevant to the question input by the user, which is beneficial to improving the accuracy of finally extracting key information and expressing key information in this embodiment; moreover, the target prompt text of this embodiment also includes the position information of the target information corresponding to the target information type in the target chart. Based on this position information, the multimodal large model can quickly locate, which is beneficial to improving the efficiency of the large model in outputting information.
[0064] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. A method for extracting key information from a chart based on a multimodal large model, characterized in that: The method comprises the following steps: S100, obtaining initial prompt text; the initial prompt text includes title prompt text, legend prompt text and coordinate axis prompt text, the title prompt text includes text for representing the extracted title and text for representing the output title information format, the legend prompt text includes text for representing the extracted legend and text for representing the output legend information format, and the coordinate axis prompt text includes text for representing the extracted coordinate axis and text for representing the output coordinate axis information format; S200, acquiring a target information type according to a type of a target chart and a type of a question input by a user; the target information type includes at least one of a title information type, a legend information type, and a coordinate axis information type; S300, obtaining position information of target information corresponding to the target information type in a target graph; S400, constructing a target prompt text according to the prompt text corresponding to the target information type in the initial prompt text and the position information of the target information type corresponding to the target information type in the target chart; S500, inputting the target prompt text and the target chart into the multimodal large model, and obtaining key information of the target chart related to the question input by the user according to the output of the multimodal large model.
2. The method for extracting key information from a chart based on a multimodal large model according to claim 1, characterized in that: The key information related to the user input question to obtain the target graph based on the output of the multimodal large model includes: S510, extracting core keywords from the question input by the user to obtain a core keyword set; S520, performing sentence processing on the output of the multimodal large model to obtain a sentence set; S530, obtaining the relevance between each sentence in the sentence set and the core keyword set; wherein the relevance between any sentence in the sentence set and the core keyword set is the number of times the core keyword in the core keyword set appears in the sentence; S540, sorting the sentences in the output of the multimodal large model in descending order of relevance; S550: Determine the sorted results as key information of the target chart related to the question input by the user.
3. The method for extracting key information from a chart based on a multimodal large model according to claim 1, characterized in that: S200 includes: S210, obtain information type list C; C={C1, C2, …, C h ,…,C H }, C h is the hth record included in C, where h ranges from 1 to H, and H is the number of records included in C; C h =(C h,1 ,C h,2 ,C h,3 ), C h,1 C h Types of questions included, C h,2 C h Included chart types, different C h Included C h,1 Different or C h,2 Different or C h,1 and C h,2 All are different, C h,3 C h Included with C h,1 and C h,2 A set of information types with corresponding relationships, C h,3 At least one of the following information types: title information type, legend information type and axis information type; S220, matching the type of the target chart and the type of the question input by the user in C; S230: Determine the information type included in the matching record as the target information type.
4. The method for extracting key information from a chart based on a multimodal large model according to claim 1, characterized in that: The legend prompt text includes: identifying the legend type and its corresponding color, in the following format: {1st type name (color)}|{2nd type name (color)}|…|{kth type name (color)}|…|{pth type name (color)}, where k is a variable, the value range of k is 1 to p, and p is the number of types in the legend.
5. The method for extracting key information from a chart based on a multimodal large model according to claim 1, characterized in that: The axis prompt text includes: extract the x-axis label and its value, the format is as follows: x-axis represents: {x-axis label}, its value is: {value 1}|{value 2}|…|{value f x }|…|{value g x }; Extract the y-axis label and its value in the following format: y-axis representation: {y-axis label}, its value: {value 1}|{value 2}|…|{value f y }|…|{value g y }; where f x and f y are variables, f x The value range is 1 to g x , g x is the number of data corresponding to the x-axis, f y The value range is 1 to g y , g y is the number of data corresponding to the y-axis.
6. The method for extracting key information from a chart based on a multimodal large model according to claim 1, characterized in that: S400 includes: S410, if the target information type includes a title information type, append the title prompt text and the position information of the target information corresponding to the title information type in the target chart to the preset designated text; if the target information type includes a legend information type, append the legend prompt text and the position information of the target information corresponding to the legend information type in the target chart to the preset designated text; if the target information type includes a coordinate axis information type, append the coordinate axis prompt text and the position information of the target information corresponding to the coordinate axis information type in the target chart to the preset designated text; S420, determining the appended designated text as the target prompt text.
7. The method for extracting key information from a chart based on a multimodal large model according to claim 1, characterized in that: The type of the question input by the user is one of a numerical query type, a data point comparison type, a trend summary type, a semantic description type, and a category reasoning type.
Citation Information
Patent Citations
Data output method and device
CN112597276A
Method and system for obtaining target document by identifying document typesetting structure
CN113553800A
Test question disassembling method and system based on test paper image, storage medium and equipment
CN113610068A
Collaborative manufacturing enterprise-oriented unstructured chart data analysis method
CN114936279A
Chart question-answering method and system based on multi-modal large model, medium and equipment
CN117390165A