Chart question and answer data set generation method and device and storage medium
By combining the content of the paper and chart information, using large language models to generate chart Q&A data sets in the scientific field, the problem of insufficient applicability of general data sets in the scientific field is solved, and the scientificity and accuracy of chart selection and question-Addition generation are achieved, and more comprehensive analysis is supported by scientific researchers.
Patent Information
- Application Number
- CN202510532894.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
The existing general chart question and answer data sets are insufficiently applicable in the scientific field, making it difficult to deal with complex scientific charts, especially highly specialized and data-complex charts in the basic science field, and lack targeted support.
Combining the paper content in the target field and multiple chart information, the chart with the highest correlation with the research conclusions is determined through large language model analysis, a question-and-answer pair is generated, and a professional chart question-and-answer dataset is constructed, including advanced reasoning and summary questions, to ensure the scientificity and accuracy of chart selection and question-and-answer generation.
It improves the accuracy of chart selection, expands the application scope of chart Q&A, supports more comprehensive analysis by scientific researchers, ensures that the question and answer are closely related to scientific research conclusions, and improves the scientific rigor and practicality of the data set.
Smart Images

Figure CN120448492A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of graph question answering technology, and in particular to a method, device, and storage medium for generating a graph question answering dataset. Background Art
[0002] In recent years, Chart Question Answering (CQA), as an important branch of the field of visual question answering, has gradually become a research hotspot for cross-modal understanding. Existing technologies mainly focus on methods for constructing chart question answering datasets in general fields, with typical representatives such as PlotQA and DVQA. Such datasets focus on simple statistical charts commonly found in the humanities (such as bar charts and line charts), and evaluate the model's ability to extract basic information from charts through structured question forms (such as information extraction, counting, and pattern recognition). Although these datasets have promoted the development of visual reasoning technology in general scenarios, the simplicity of their data structure and question design limits their applicability in scientific fields. Scientific charts (such as those in astronomy) are usually highly specialized, involving complex data relationships, multi-dimensional parameters, and discipline-specific symbols. Existing general datasets lack targeted support for such charts, making it difficult for models to adapt to the in-depth needs of scientific research.
[0003] Therefore, there is an urgent need to develop specialized and highly relevant graph question answering datasets for basic science fields to fill the current technological gaps. Summary of the Invention
[0004] To overcome the problems existing in the related art, this specification provides a method, device and storage medium for generating a chart question and answer dataset.
[0005] According to a first aspect of the embodiments of this specification, a method is provided, comprising:
[0006] Combining the content of the target field's papers with the information of multiple charts in the papers, determining the chart with the highest correlation with the research conclusions of the papers as the target chart;
[0007] Based on the relationship between the target graph and the research conclusion of the paper, generating a question-answer pair corresponding to the target graph;
[0008] Based on the multiple question-answer pairs, a graph question-answering dataset in the target domain is created.
[0009] According to a method for generating a chart question-answering dataset provided in this specification, the method combines the content of a target field paper with information of multiple charts in the paper to determine the chart with the highest correlation with the research conclusion of the paper as the target chart, including:
[0010] For each of the papers, perform the following steps:
[0011] Analyze the content of the paper and generate research conclusions that reflect the central conclusions of the paper;
[0012] Analyze multiple figures in the paper and generate figure summaries describing the key information of the figures;
[0013] The chart summary of each chart is compared with the research conclusion of the paper, and the chart with the highest correlation with the research conclusion of the paper is determined as the target chart.
[0014] According to a method for generating a chart question-and-answer dataset provided in this specification, the chart abstract of each chart is compared with the research conclusion of the paper, and the chart with the highest correlation with the research conclusion of the paper is determined as the target chart, including:
[0015] Based on multiple large language models, the chart abstract of each chart is compared with the research conclusions of the paper, and the charts to be selected that have the highest correlation with the research conclusions of the paper are selected;
[0016] The to-be-selected chart with the highest number of votes among the screening results of the plurality of large language models is determined as the target chart.
[0017] According to a method for generating a chart question-answering dataset provided in this specification, the method of analyzing multiple charts in the paper and generating a chart summary describing key information of the charts includes:
[0018] All the figures and tables in the paper were extracted and rendered using image extraction technology;
[0019] The extracted charts are used to analyze multiple charts in the paper and generate chart summaries that describe key information of the charts.
[0020] According to a method for generating a graph question answering dataset provided in this specification, the method further includes analyzing multiple graphs in the paper for the extracted graphs to generate a graph summary describing the key information of the graphs.
[0021] Filter out non-text charts from the extracted charts;
[0022] For the filtered figures, multiple figures in the paper are analyzed to generate a figure summary describing the key information of the figure.
[0023] According to a method for generating a graph question-answer dataset provided in this specification, generating a question-answer pair corresponding to the target graph based on the association between the target graph and the research conclusion of the paper includes:
[0024] Based on the relationship between the target graph and the research conclusions of the paper, a large language model is used to generate multi-dimensional questions and answers corresponding to the questions;
[0025] Based on the questions and the answers, a plurality of question-answer pairs corresponding to the target graph are created.
[0026] According to a method for generating a graph question-answering dataset provided in this specification, the multi-dimensional question-answering questions include reasoning questions, which are obtained by reasoning based on the data and scientific background of the target graph;
[0027] After creating a plurality of question-answer pairs corresponding to the target graph, the method further includes:
[0028] generating a corresponding answer based on the reasoning question using a large language model without combining information from the target graph;
[0029] performing consistency evaluation on the answer generated by combining the information of the target chart with the answer generated without combining the information of the target chart;
[0030] Eliminate the corresponding question-answer pairs when the evaluation result is greater than the threshold.
[0031] According to a method for generating a chart question and answer dataset provided in this specification, the multi-dimensional question and answer questions also include summary questions, and the answers to the summary questions are summarized based on the research conclusions reflected in the target chart.
[0032] According to a second aspect of the embodiments of this specification, there is provided an apparatus, including:
[0033] A chart determination module is used to combine the content of a paper in a target field with information of multiple charts in the paper, and determine the chart with the highest correlation with the research conclusion of the paper as the target chart;
[0034] A question-answer generation module, configured to generate question-answer pairs corresponding to the target graph based on the relationship between the target graph and the research conclusions of the paper;
[0035] A dataset creation module is used to create a graph question-answering dataset in the target domain based on the multiple question-answer pairs.
[0036] According to a third aspect of the embodiments of this specification, a device is provided, including:
[0037] The system comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, any of the above-mentioned methods for generating a chart question and answer dataset is implemented.
[0038] According to a third aspect of the embodiments of this specification, a computer-readable storage medium is provided, including:
[0039] When the computer program is executed by a processor, it implements any of the above-mentioned methods for generating a graph question and answer dataset.
[0040] The technical solutions provided by the embodiments of this specification may have the following beneficial effects:
[0041] In the embodiments of this specification, the content of the target field paper and the information of multiple charts in the paper are combined to determine the chart with the highest correlation with the research conclusion of the paper as the target chart; by combining the content analysis of the paper, the accuracy of chart selection is improved without relying solely on the similarity of image features, ensuring that the final selected chart maximizes its support for the research conclusion. Then, based on the correlation between the target chart and the research conclusion of the paper, a question-and-answer pair corresponding to the target chart is generated; this not only expands the application scope of chart question-and-answer, but also can deeply explore the information in the chart, ensure the scientificity and accuracy of chart selection and question-and-answer generation, and create a chart question-and-answer dataset for the target field based on multiple question-and-answer pairs to support more comprehensive analysis by scientific researchers.
[0042] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the specification and, together with the description, serve to explain the principles of the specification.
[0044] Figure 1 It is a flow chart of a method according to an exemplary embodiment of the present specification.
[0045] Figure 2 It is a block diagram of a device according to an exemplary embodiment of the present specification.
[0046] Figure 3 This is a schematic diagram of a device according to an exemplary embodiment of the present specification. DETAILED DESCRIPTION
[0047] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this specification. Rather, they are merely examples of apparatus and methods consistent with certain aspects of this specification, as detailed in the appended claims.
[0048] The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this specification. As used in this specification and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0049] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information without departing from the scope of this specification. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."
[0050] This specification provides a method, device, and computer-readable storage medium for generating a graph question-and-answer dataset. The following describes the embodiments of this specification in detail with reference to the accompanying drawings. The features of the following embodiments and implementations may be combined with each other unless they conflict.
[0051] To expand chart question-answering capabilities in multidisciplinary contexts, some research has attempted to construct cross-domain datasets, such as CharXiv. This dataset combines visual encoders (such as SigLIP) with manual screening to select candidate charts from a pre-trained image library. Initial screening is performed based on cosine similarity (≥0.65) with the average visual embedding of charts in MathVista to cover a wide range of chart types. Subsequently, descriptive and reasoning questions are designed through manual annotation. Descriptive questions primarily assess the model's ability to extract and summarize essential information about the chart, covering tasks such as information extraction, enumeration, pattern recognition, counting, and combination. Each chart is assigned four descriptive questions, one of which is intentionally unanswerable. Reasoning questions assess the model's ability to reason visually and numerically. Questions are manually designed to ensure clear and unique answers and do not require advanced domain knowledge or additional information outside the chart. Answers to reasoning questions are categorized into four types: text within the chart, text outside the chart, numbers within the chart, and numbers outside the chart. Compared to other datasets, CharXiv emphasizes reasoning solely based on the visual and numerical information within the chart, avoiding the reliance on additional expertise.
[0052] Therefore, there is an urgent need to develop specialized and highly relevant graph question answering datasets for basic science fields to fill the current technological gaps.
[0053] In order to solve the above technical problems, this specification provides a method for generating a graph question-answering dataset.
[0054] It aims to improve the accuracy of chart selection by combining content analysis of the paper, without relying solely on the similarity of image features, to ensure that the final selected charts maximize their support for the research conclusions, and by using the relationship between the selected charts and the research conclusions of the paper, it can generate scientific research questions and answers that are more complex than traditional descriptive and inferential questions. This not only expands the application scope of chart questions and answers, but also can deeply explore the information in the charts, ensure the scientificity and accuracy of chart selection and question and answer generation, and support researchers in more comprehensive analysis.
[0055] like Figure 1 As shown, Figure 1 This is a flowchart of a method according to an exemplary embodiment of the present specification, comprising the following steps:
[0056] In step 102, the content of the paper in the target field is combined with the information of multiple charts in the paper, and the chart with the highest correlation with the research conclusion of the paper is determined as the target chart.
[0057] As an example, the target field described in the specification refers to the user-selected field in the research branch of basic science that is highly specialized, data-complex, and dependent on subject knowledge, such as astrophysics, cosmology, high-energy physics, etc. At present, all evaluation data sets lack content related to domain knowledge in a specific field, lack coverage of specialized charts, and are difficult to migrate directly. Therefore, a chart question-and-answer data set for a specific field is constructed to relate the data set generation process to specific domain knowledge, provide underlying support for intelligent scientific research assistants, automated knowledge discovery, and cross-modal model training, and promote the efficiency and depth of scientific research. The following takes the field of astronomy as an example for explanation. Other fields are basically the same as their embodiments and will not be repeated here.
[0058] In some embodiments, combining the content of a paper in the target field with information of multiple charts in the paper to determine the chart with the highest correlation with the research conclusion of the paper as the target chart includes:
[0059] For each of the papers, perform the following steps:
[0060] Step a11: analyzing the content of the paper and generating a research conclusion of the paper that reflects the central conclusion of the paper;
[0061] Step a12: analyzing multiple charts in the paper and generating chart summaries describing key information of the charts;
[0062] Step a13: Compare the chart abstract of each chart with the research conclusion of the paper, and determine the chart with the highest correlation with the research conclusion of the paper as the target chart.
[0063] First, obtain the required papers in your target field.
[0064] This manual uses a screening mechanism based on authoritative journals and paper citations to quickly identify papers with high reference value, thereby ensuring the quality of subsequent analysis results.
[0065] As an example, stratified sampling was conducted by sub-field and time period, excluding tool articles (such as instruments and software).
[0066] For example, we searched for astrophysics papers published in authoritative journals and conferences (e.g., Nature, ApJ, MNRAS) in the ADS database, with publication dates between April 2007 and July 2023. These papers are grouped into six sub-fields on arXiv: astro-ph.HE, astro-ph.EP, astro-ph.SR, astro-ph.CO, astro-ph.GA, and astro-ph.IM.
[0067] To reduce potential bias caused by publication dates, papers were further divided into time periods, and screening was performed based on each time period. For example, papers published between April 2007 and July 2023 were divided into three time periods: 2007 to 2012, 2013 to 2017, and 2018 to 2023. Within each subfield and time period, papers were selected based on defined criteria, including, but not limited to, citation counts and research value. For example, the top 1% of most cited papers were selected based on citation counts from the SAO / NASA Astrophysics Data System (ADS).
[0068] Furthermore, to improve the quality of selected papers, we communicated with astronomy experts and found that while articles describing observational instruments, software, observational projects, and observational data received a high number of citations, they did not publish significant scientific conclusions and therefore were not considered desirable. Therefore, we performed fuzzy matching on the keywords field, removing articles containing the terms catalog, survey, instrument, and software. This resulted in a selection of papers that would be used to construct a high-quality astronomy chart question-and-answer dataset.
[0069] Next, analyze the content of the paper and extract the research conclusions of the paper.
[0070] The content of the paper described in this description includes the title, abstract, and conclusion of the paper. The title, abstract, and conclusion of the paper are analyzed through natural language processing (NLP) technology, keywords and research hypotheses are identified, the central conclusion of the paper is extracted, and its main findings are synthesized at a high level to obtain the research conclusion of the paper.
[0071] As an example, the content of the paper is analyzed using a large language model, such as Claude 3.5, BERT, and GPT models.
[0072] Next, we automatically extract charts from the paper, analyze multiple charts in the paper, and generate chart summaries describing the key information of the charts for subsequent screening of charts that are highly relevant to the research conclusions.
[0073] As an example, the analysis of multiple figures in the paper generates a figure summary describing key information of the figures, including:
[0074] Step a121, extracting and rendering all the graphs in the paper using image extraction technology;
[0075] Step a122 , analyzing the extracted charts in the paper and generating a chart summary describing key information of the charts.
[0076] First, all the charts in the papers obtained above are extracted using automated image extraction technology. Among these charts, some have little relevance to the research value of the target field, so a preliminary selection of charts is required.
[0077] In some cases, the paper includes clearly non-informative figures and tables (e.g., figures in the appendix), which generally fail to summarize the core findings of the study and therefore need to be excluded.
[0078] Exemplarily, all charts in the paper are extracted and rendered using image extraction technology; non-text charts in the extracted charts are filtered out; and for the filtered charts, multiple charts in the paper are analyzed to generate chart summaries that describe key information of the charts.
[0079] In this example, automated paper screening and image extraction technology reduces manual intervention during the screening process, significantly improving the speed and accuracy of chart selection. Combined with the paper content, this improves the precision of chart selection, making each chart more valuable for research.
[0080] Next, multiple charts in the paper are analyzed to generate chart summaries that describe the key information of the charts.
[0081] As an example, in step a12, a large language model (e.g., Claude 3.5) is used to parse the diagram metadata, and each diagram and its title and related description content (e.g., " Figure 1 : The relationship between galaxy mass and star formation rate”) to produce a concise graphical summary that captures its key insights.
[0082] If the same paper contains multiple figures and tables, priority should be given to figures and tables that directly support or verify the conclusions of the paper (such as figures and tables containing key experimental data).
[0083] As an example, comparing the chart abstract of each chart with the research conclusion of the paper and determining the chart with the highest correlation with the research conclusion of the paper as the target chart includes:
[0084] Step a131: Based on multiple large language models, compare the chart abstract of each chart with the research conclusion of the paper, and select the charts to be selected that have the highest correlation with the research conclusion of the paper;
[0085] Step a132: Determine the to-be-selected chart with the highest number of votes in the screening of the plurality of large language models as the target chart.
[0086] For each paper, we compared the insights extracted from individual figures with the key conclusions of the paper, and selected one or two figures that best represented its main conclusions. This process was implemented using a large language model.
[0087] By combining this with content analysis of the paper's conclusions, we can identify the most relevant figures and tables, ensuring that they directly reflect the research findings or support the research hypotheses. This approach avoids simplistic reasoning based solely on the figures' content and effectively connects the figures' content to the research conclusions. By leveraging the inherent connection between figures and the paper's conclusions, we improve the accuracy of our figure selection and ensure that the final figures and tables provide the most support for the research conclusions.
[0088] Furthermore, in order to ensure the accuracy and fairness of the selection process, multiple large language models are used to independently compare the chart abstract of each chart with the research conclusions of the paper. Each model votes for the chart based on the relevance of the content, and the votes for each chart are finally counted. The charts with the highest number of votes are determined to be representative and retained to ensure a robust and fair selection process and avoid the subjective bias of a single model. In some examples, multiple large language models are used to independently compare the chart abstract of each chart with the research conclusions of the paper. Only charts determined to be representative by the majority of models are retained. The majority of models can be, but are not limited to, models that exceed more than half of the number of large language models.
[0089] Exemplarily, two large language models, GPT-4o and Claude 3.5, are used to implement the above process of determining representative charts. In this process, GPT-4o and Claude 3.5 are performed independently. GPT-4 and Claude 3.5 independently evaluate the relevance of the charts to the conclusions of the paper, and determine the target charts through consistency comparison to ensure objectivity.
[0090] It is understandable that when multiple charts are associated with the same conclusion, a comprehensive evaluation is required through dimensions such as data coverage and chart complexity.
[0091] Through the above embodiment, after analyzing the content of the paper and generating the research conclusions of the paper, analyzing multiple charts in the paper and generating chart summaries, and then integrating the correlation between the chart summaries and the research conclusions of the paper, the most representative target chart is selected by using the chart-paper conclusion matching and voting selection mechanism. This complete chain of thought (CoT) reasoning mechanism ensures the scientificity and accuracy of the chart selection.
[0092] In step 104, based on the association between the target graph and the research conclusion of the paper, a question-answer pair corresponding to the target graph is generated.
[0093] While identifying the target chart through the thought chain reasoning mechanism, a summary text is generated describing the overlap between the chart and the paper. This summary text refers to the correlation or consistency between the chart's content and the paper's research conclusions. It can be understood that the generated summary text includes how the chart supports the paper's research conclusions and the corresponding relationship between the chart and the research conclusions. For example, how the chart supports the paper's research conclusions can include: how the data trends, outliers, or key parameters in the chart verify the hypotheses proposed in the paper. The corresponding relationship between the chart and the research conclusions can be explained, for example: a certain curve in the chart directly reflects the "nonlinear relationship" mentioned in the paper, or a certain set of data points supports the conclusion that "galaxy mergers drive star formation." Therefore, by using text to explain the specific connection between the selected chart and the core conclusion of the paper, we avoid relying solely on the model voting results and lack of interpretability, provide background information for the subsequent question and answer generation, and design reasoning questions that are more closely aligned with the paper's conclusions.
[0094] Furthermore, based on the specific relationship between the target graph and the research conclusions of the paper, scientifically rigorous question-answer pairs are generated.
[0095] As an example, generating a question-answer pair corresponding to the target graph based on the association between the target graph and the research conclusion of the paper includes:
[0096] Step b11: Based on the relationship between the target graph and the research conclusions of the paper, generate multi-dimensional questions and answers corresponding to the questions through a large language model;
[0097] Step b12: Based on the questions and the answers, create multiple question-answer pairs corresponding to the target chart.
[0098] Based on the summary content of a chart, we can further infer possible related research questions and generate chart-question-answer pairs to support multiple rounds of scientific Q&A or data reuse. Unlike general scientific chart evaluation, constructing a domain-specific chart Q&A dataset requires a more precise and expert-based approach. The questions and answers must be closely related to established scientific principles and research results.
[0099] The input to the multimodal large language model is the summary text describing the overlap between the figure and the paper, as well as the figure itself. Using the multimodal large language model, a variety of questions and answers related to the figure are automatically generated, ensuring coverage of a variety of question formats.
[0100] Specifically, the multimodal large language model automatically generates questions of varying difficulty based on the summary of chart-conclusion correlations. It then combines chart data and the research conclusions of the paper to generate answers to the questions, which are then combined into question-answer pairs to meet the needs of in-depth scientific research analysis.
[0101] In some embodiments, multi-dimensional questions are generated through a large language model. This method covers multiple question categories, including high-level reasoning questions and summary questions.
[0102] Among them, high-level reasoning questions focus on logical analysis and in-depth reasoning that combines chart data with scientific background (such as "comparison of the strength of the correlation between HI mass and HI radius").
[0103] Example: How does the relationship between HImass and HIradius compare to the relationship between HI mass and stellar luminosity in terms ofcorrelation strength?
[0104] This type of question requires a deep understanding of the relevant research, integrating the data in the chart with its scientific context to draw accurate conclusions. It primarily assesses the model's ability to perform complex reasoning, examining how it integrates information, infers relationships, and draws scientifically sound conclusions.
[0105] To ensure that high-level reasoning questions are highly relevant to graphs, we then generate answers for all high-level reasoning questions without graphs using a large language model, compare these answers with the previously generated answers, and update the question-answer pairs to ensure that the reasoning questions are closely relevant to the graphs.
[0106] As an example, after creating the plurality of question-answer pairs corresponding to the target chart, the method further includes:
[0107] Step b14: generating a corresponding answer based on the reasoning question using a large language model without combining information from the target graph;
[0108] Step b15: performing consistency evaluation on the answer generated by combining the information of the target chart with the answer generated without combining the information of the target chart;
[0109] Step b16: Eliminate the question-answer pairs corresponding to evaluation results greater than the threshold.
[0110] By comparing the differences between answers generated without charts and answers generated with charts, we screen out high-level reasoning questions that must rely on chart information to be answered correctly, ensuring that these questions are highly relevant to the charts. For example, the answers corresponding to questions generated with information from the target chart and answers generated without information from the target chart are evaluated for consistency. For question-answer pairs corresponding to evaluation results greater than a threshold, this means that the answers without charts are highly consistent with the answers with charts, indicating that the question itself may not rely on chart information and can be answered by attempting. Such questions cannot reflect chart specificity, resulting in a lack of scientific research value in the generated question-answer pairs. For question-answer pairs corresponding to evaluation results less than or equal to a threshold, this means that large differences in answers indicate that the question must be combined with specific data in the chart, ensuring that the question-answer pairs closely revolve around the chart content and improving the scientific rigor of the data set.
[0111] Exemplarily, the answers corresponding to the questions generated by combining the information of the target graph are compared with the answers generated without combining the information of the target graph, and scored using the large model (ranging from 0 to 1), and all question-answer pairs with scores above 0.7 are removed to ensure that the reasoning questions are closely related to the graph.
[0112] Through the above-mentioned difference scoring mechanism, based on the consistency comparison of answers without charts and answers with charts, general questions are filtered out, and questions that must rely on chart information to answer are screened out, ensuring that the generated question-answer pairs are highly relevant to the chart content, supporting complex scientific reasoning tasks (such as hypothesis verification and trend analysis), thereby improving the scientific value and practicality of the dataset.
[0113] Summary questions focus on summarizing chart trends (such as "Summarize the trend-like characteristics") and core insights from a macro perspective.
[0114] Example: Summarize the trend-like or high-level characteristics of this chart.
[0115] These questions focus on extracting information from a macro perspective and high-level insights, rather than simply interpreting the chart's features. They primarily assess the model's ability to identify patterns in the chart, integrate the chart's intended message, and provide a concise yet comprehensive summary that conveys the chart's core message. Summary questions also need to be answered by summarizing the research conclusions reflected in the target chart.
[0116] Through the above-mentioned embodiments, by expanding the dimensions of question construction and leveraging the relationship between charts and research conclusions, the system can generate more complex scientific research questions and answers than traditional descriptive and inferential questions. This not only expands the application scope of chart questions and answers, but also enables in-depth exploration of the information contained in charts, supporting more comprehensive analysis by researchers.
[0117] In step 106 , a graph question-answering dataset of the target domain is created based on the plurality of question-answer pairs.
[0118] By automatically extracting professional charts in the field of astronomy from academic literature, combined with manual verification and question-answering expansion by domain experts, we construct a high-quality astronomy chart question-answering dataset to meet the needs of this field in intelligent question-answering and knowledge discovery.
[0119] This specification provides a method, device, and computer-readable storage medium for generating a chart question and answer dataset. By combining the content of a target field paper with the information of multiple charts in the paper, the chart with the highest correlation with the research conclusion of the paper is determined to be the target chart. By combining the content analysis of the paper, the accuracy of chart selection is improved without relying solely on the similarity of image features, ensuring that the final selected chart maximizes its support for the research conclusion. Then, based on the correlation between the target chart and the research conclusion of the paper, a question and answer pair corresponding to the target chart is generated. This not only expands the application scope of chart question and answer, but also can deeply explore the information in the chart, ensure the scientificity and accuracy of chart selection and question and answer generation, and create a chart question and answer dataset for the target field based on multiple question and answer pairs to support more comprehensive analysis by scientific researchers.
[0120] like Figure 2 As shown, Figure 2 1 is a block diagram of a device according to an exemplary embodiment of the present specification, the device including:
[0121] A chart determination module is used to combine the content of a paper in a target field with information of multiple charts in the paper, and determine the chart with the highest correlation with the research conclusion of the paper as the target chart;
[0122] A question-answer generation module, configured to generate question-answer pairs corresponding to the target graph based on the relationship between the target graph and the research conclusions of the paper;
[0123] A dataset creation module is used to create a graph question-answering dataset in the target domain based on the multiple question-answer pairs.
[0124] The implementation process of the functions and effects of each module / submodule / unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and the same technical effects can be achieved, so it will not be repeated here.
[0125] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this specification. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0126] Figure 3 The following is a schematic diagram of the physical structure of a computer device, such as Figure 3 As shown, the computer device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. The processor 810, the communications interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call logic instructions in the memory 830 to execute the method for generating a graph question and answer dataset.
[0127] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of this specification, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of this specification. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0128] On the other hand, this specification also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the chart question and answer dataset generation method provided by the above methods.
[0129] On the other hand, the present specification also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the method for generating a chart question and answer dataset provided by the above methods.
[0130] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0131] Other embodiments of the present invention will readily occur to those skilled in the art upon consideration of the present invention and practice of the invention claimed herein. This specification is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of this specification and include common knowledge or customary techniques in the art not claimed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present invention being indicated by the following claims.
[0132] It should be understood that the present description is not limited to the exact structure that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present description is limited only by the appended claims.
[0133] The above description is only a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this specification should be included in the scope of protection of this specification.
Claims
1. A method for generating a graph question answering dataset, characterized in that: The method comprises: Combining the content of the target field's papers with the information of multiple charts in the papers, determining the chart with the highest correlation with the research conclusions of the papers as the target chart; Based on the relationship between the target graph and the research conclusion of the paper, generating a question-answer pair corresponding to the target graph; Based on the multiple question-answer pairs, a graph question-answering dataset in the target domain is created.
2. The method for generating a graph question-answering dataset according to claim 1, wherein: The target chart is determined by combining the content of the target field paper with the information of multiple charts in the target field paper to have the highest correlation with the research conclusion of the target paper, including: For each of the papers, perform the following steps: Analyze the content of the paper and generate research conclusions that reflect the central conclusions of the paper; Analyze multiple figures in the paper and generate figure summaries describing the key information of the figures; The chart summary of each chart is compared with the research conclusion of the paper, and the chart with the highest correlation with the research conclusion of the paper is determined as the target chart.
3. The method for generating a chart question-answering dataset according to claim 2, wherein: The chart summary of each chart is compared with the research conclusion of the paper, and the chart with the highest correlation with the research conclusion of the paper is determined as the target chart, including: Based on multiple large language models, the chart abstract of each chart is compared with the research conclusions of the paper, and the charts to be selected that have the highest correlation with the research conclusions of the paper are selected; The to-be-selected chart with the highest number of votes among the screening results of the plurality of large language models is determined as the target chart.
4. The method for generating a graph question-answering dataset according to claim 2, wherein: The method includes analyzing multiple charts in the paper and generating a chart summary describing the key information of the charts, including: All the figures and tables in the paper were extracted and rendered using image extraction technology; For the extracted charts, multiple charts in the paper are analyzed to generate chart summaries describing the key information of the charts.
5. The method for generating a graph question-answering dataset according to claim 4, wherein: The method further comprises analyzing the extracted charts and graphs in the paper to generate a chart summary describing the key information of the charts. Filter out non-text charts from the extracted charts; For the filtered figures, multiple figures in the paper are analyzed to generate a figure summary describing the key information of the figure.
6. The method for generating a graph question-answering dataset according to claim 1, wherein: Generating a question-answer pair corresponding to the target graph based on the association between the target graph and the research conclusion of the paper includes: Based on the relationship between the target graph and the research conclusions of the paper, a large language model is used to generate multi-dimensional questions and answers corresponding to the questions; Based on the questions and the answers, a plurality of question-answer pairs corresponding to the target graph are created.
7. The method for generating a graph question-answering dataset according to claim 6, wherein: The multi-dimensional question-answering questions include reasoning questions, which are obtained by reasoning based on the data and scientific background of the target chart; After creating a plurality of question-answer pairs corresponding to the target graph, the method further includes: generating a corresponding answer based on the reasoning question using a large language model without combining information from the target graph; performing consistency evaluation on the answer generated by combining the information of the target chart with the answer generated without combining the information of the target chart; Eliminate the corresponding question-answer pairs when the evaluation result is greater than the threshold.
8. The method for generating a graph question-answering dataset according to claim 6, wherein: The multi-dimensional question-and-answer questions also include summary questions, and the answers to the summary questions are summarized based on the research conclusions reflected in the target chart.
9. A computer device, characterized in that: The method comprises a memory, a processor and a chart question and answer dataset generation program stored in the memory and executable on the processor, wherein when the processor executes the chart question and answer dataset generation program, the steps of the chart question and answer dataset generation method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a chart question and answer dataset generation program, which, when executed, implements the steps of the chart question and answer dataset generation method according to any one of claims 1 to 8.
Citation Information
Cited By
Scientific chart question and answer data generation method and system based on multi-modal large model
CN121615791A