Knowledge graph construction method and retrieval system based on chart image

Through deep learning technology, visual elements and semantic relationships are extracted from chart images and knowledge graph representation is established, which solves the problems of loss and poor adaptability of chart knowledge mining information in the existing technology, and realizes efficient and automated chart analysis and retrieval.

CN120011573APending Publication Date: 2025-05-16HANGZHOU DIANZI UNIVERSITY SHANGYU INSTITUTE OF SCIENCE & ENGINEERING CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411972074.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art ignores the visual coding and semantic meaning of the chart in the mining of chart image knowledge, resulting in information loss, and traditional methods are time-consuming and labor-intensive and difficult to adapt to new chart types.

Method used

Using deep learning-based methods, visual elements and their semantic relationships are extracted from chart images through technologies such as convolutional neural networks, and a unified knowledge graph representation is established to support tasks such as chart retrieval and question answering.

Benefits of technology

It realizes high-quality knowledge representation of chart images, improves the degree of automation of chart analysis, and enhances the accuracy and efficiency of chart retrieval and question answering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011573A_ABST
    Figure CN120011573A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph construction method based on chart images and a retrieval system. The method comprises the following steps of: 1, establishing a division system of entity types and semantic relationship types among entities in a chart image; 2, identifying a chart category of the chart image, extracting character information in the chart image, and dividing a plurality of entities in the chart image; and 3, establishing a semantic relationship among different entities, and establishing a knowledge graph according to the entities extracted from the chart image and the semantic relationship among the different entities. The invention provides a unified and explainable chart image representation method, which can effectively capture visual elements and semantic relationships in the chart image and support downstream tasks such as semantic perception chart retrieval and chart question and answer, thereby improving the ability of chart image analysis and application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of chart data visualization and knowledge mining, and specifically relates to a knowledge graph construction method and retrieval system based on chart images. Background Art

[0002] With the widespread application of data visualization, a large number of chart images such as bar charts, pie charts, and line charts have emerged. Mining knowledge from these chart images has become increasingly important, which is of great significance for downstream tasks such as chart retrieval and knowledge graph completion. However, most existing chart knowledge mining methods focus on converting chart images into raw data, often ignoring the visual encoding and semantic meaning of charts, which may lead to information loss in many downstream tasks.

[0003] The development of machine learning algorithms has provided new solutions for knowledge mining of chart images. Traditional chart image processing methods usually rely on manually designed features and templates to identify chart elements and their relationships, but this method is not only time-consuming and labor-intensive, but also difficult to adapt to new chart types. In recent years, deep learning-based methods have begun to become the mainstream method for chart image analysis. These methods use neural networks to automatically extract features from raw chart images, reducing the need for manual feature engineering and have strong generalization capabilities when dealing with complex chart relationships.

[0004] Data visualization plays an important role in quickly acquiring knowledge and gaining insights into patterns. Applying visual analysis techniques to chart images can improve the understanding and expression of chart images. Common chart visualizations include chart element recognition, character recognition, and classification. Visual analysis techniques express the importance and relevance of elements through the size of chart elements, line thickness, and other methods. Simplified visualization methods of charts are suitable for large-scale chart images, which have the characteristics of huge data volume, multidimensional attributes, and complex structures. Some work is devoted to the fusion research of chart images to promote the improvement and development of the entire chart image knowledge mining.

[0005] The embedding representation of chart images is a technique that maps the visual elements and semantic relationships in chart images into vector space. It converts the data of chart images into numerical features that can be processed by machine learning algorithms, which is convenient for subsequent data mining, reasoning and prediction tasks. Typical chart image embedding techniques can be divided into two categories: visual feature extraction models and semantic matching models. Visual feature extraction models use technologies such as convolutional neural networks to extract visual features from chart images, while semantic matching models measure the credibility of chart image content by matching the latent semantics of chart elements with the relationships contained in the vector space representation. These embedding methods have been proven to be very effective in tasks such as chart retrieval and chart question answering. However, most methods encode chart elements and relationships in independent embedding vectors, lacking attention to the embedding of individual chart elements.

[0006] In summary, existing research mainly focuses on the visualization and processing of chart images. In contrast, this method applies interactive visual analysis to the element identification and relationship mining of chart images to promote the high-quality construction of chart image knowledge graphs. Although chart images are very effective in representing structured data, their structural characteristics make operations on raw chart images not intuitive or easy to expand. Chart image embedding representation technology can map the elements and relationships of chart images into vector space, providing a more convenient way of data processing and analysis. The present invention designs a new embedding method based on chart elements, which focuses more on capturing the relational features of chart elements in order to more intuitively reveal the characteristics of chart images. Summary of the invention

[0007] The purpose of the present invention is to provide a method for constructing a knowledge graph based on chart images and a retrieval system; the method can effectively extract visual elements and their semantic relationships from chart images and convert them into a unified knowledge graph representation to facilitate downstream tasks such as chart retrieval and chart question answering. The method can provide a highly unified, expressive and interpretable chart image knowledge representation.

[0008] In a first aspect, the present invention provides a method for constructing a knowledge graph based on a chart image, which comprises the following steps: Step 1: Establish a classification system for entity types in chart images and semantic relationship types between entities.

[0009] Step 2: Identify the chart category of the chart image, extract text information in the chart image, and divide multiple entities in the chart image.

[0010] Step 3: Establish semantic relationships between different entities, and build a knowledge graph based on the entities extracted from the chart image and the semantic relationships between different entities.

[0011] Preferably, the entity types in the chart image include visual elements, visual element attribute values, data variables, data variable values ​​and visual insights.

[0012] Preferably, the visual element comprises a graphic mark; and the visual element attribute value comprises a height and a line color of the graphic mark.

[0013] Preferably, the data variables include the coordinate axes and the title of the legend in the chart image; the data variable values ​​include the numerical labels on the coordinate axes and the legend. The data variable values ​​are obtained by extracting text information from the chart image.

[0014] Advantageously, the visual insight comprises semantic information inferred from the original chart image.

[0015] Preferably, the entity types in the chart image include visual elements, visual element attribute values, data variables, data variable values ​​and visual insights.

[0016] Preferably, the visual element comprises a graphic mark; and the visual element attribute value comprises a height and a line color of the graphic mark.

[0017] Preferably, the data variables include the coordinate axes and the title of the legend in the chart image; the data variable values ​​include the numerical labels on the coordinate axes and the legend. The data variable values ​​are obtained by extracting text information from the chart image.

[0018] Advantageously, the visual insight comprises semantic information inferred from the original chart image.

[0019] Preferably, the entity relationships in the chart image include visual attribute correspondence relationships, data variable correspondence relationships, visual encoding mapping relationships and visual insight correspondence relationships.

[0020] Preferably, the text information in the chart image is extracted by an optical character recognition method. The chart category of the chart image is identified by a convolutional neural network.

[0021] In a second aspect, the present invention provides a chart image retrieval system, which includes a chart recognition module, a text extraction module, a knowledge graph construction module and a retrieval execution module. The chart recognition module is used to identify the type of chart in the image. The text extraction module is used to extract text information in the chart image; the knowledge graph construction module is used to execute the aforementioned knowledge graph construction method based on chart images; the retrieval execution module is used to match the knowledge graph with the retrieval information input by the user to obtain a matching result.

[0022] In a third aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the memory stores the computer program; and the processor executes the aforementioned method for constructing a knowledge graph based on a chart image.

[0023] In a fourth aspect, the present invention provides a readable storage medium storing a computer program; when the computer program is executed by a processor, it is used to implement the aforementioned method for constructing a knowledge graph based on a chart image.

[0024] Compared with the prior art, the present invention has the following beneficial effects: The present invention proposes a novel knowledge graph representation method, which can characterize the visual elements and their relationships in a chart image in a unified and expressive way; designs a unified framework, which can automatically convert chart images into the proposed knowledge graph representation, and improves the automation of chart image analysis; through case studies and downstream applications, it proves the effectiveness of the proposed method in representing chart images, which can promote downstream tasks such as chart retrieval and chart question answering; through quantitative evaluation, it verifies the effectiveness of the two basic building blocks of the chart to knowledge graph conversion framework - object recognition and optical character recognition, which provides strong technical support for knowledge mining of chart images. In summary, the present invention provides a brand-new solution for the knowledge representation and analysis of chart images by constructing a knowledge graph, which has broad application prospects and practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a process framework diagram for converting a chart image into a knowledge graph in the present invention.

[0026] Figure 2 It is a schematic diagram of the relationship types between different entities in the present invention.

[0027] Figure 3 Schematic diagram of the knowledge graph of four different types of charts in the present invention; part (a) corresponds to a bar chart, part (b) corresponds to a line chart, part (c) corresponds to a pie chart, and part (d) corresponds to a scatter plot.

[0028] Figure 4 An example result diagram of a graph search using the knowledge graph constructed by the present invention on a graph image; (a) corresponds to the result of the data variable search with the highlighted query keyword. (b) corresponds to the result of combining the data variable and visual insight search with the related entities and relationships displayed below the graph.

[0029] Figure 5 This is a result graph of knowledge search and answering using the knowledge graph constructed using the present invention on a chart image. DETAILED DESCRIPTION

[0030] In conjunction with the accompanying drawings, the diagram image representation method based on the knowledge graph of the present invention is further described below, and its application to downstream tasks such as diagram retrieval and diagram question answering is further described. The specific steps are as follows: like Figure 1 As shown, a method for constructing a knowledge graph based on a chart image includes the following steps: Step 1) In order to convert the chart image into a unified knowledge graph representation, the present invention first defines entity types. The chart includes some or all of the following elements: chart title, x-axis title, y-axis title, legend title, x-axis label, y-axis label, legend label, and graphic mark. Based on these elements, this embodiment constructs five types of entities.

[0031] The five types of entities include visual elements (VE), visual element attribute values ​​(VEPV), data variables (DV), data variable values ​​(DVV), and visual insights (VI). Specifically, visual elements (VE) are key components of charts, including graphical markers (bars, lines, etc.) and axes, such as entries in a bar chart, sectors in a pie chart, or data points in a line chart. Visual element attribute values ​​(VEPV) include bar heights and line colors, which are usually encoded with different data values ​​and are therefore considered entities in the knowledge graph. Data variables (DV) include x-axis titles, y-axis titles, and legend titles, which contain key semantic information related to data values. Data variable values ​​(DVV) include x-axis labels, y-axis labels, and legend labels, which convey specific data values ​​for categorical and continuous data and are also basic entities for correctly understanding charts. Visual insights (VI). The present invention regards visual insights that are easily perceived in charts as different entity types. By grouping these elements into separate entities, the relationships between different elements in the knowledge graph can be modeled more effectively.

[0032] After defining the five entity types, this embodiment further defines four types of semantic relationships between different entities in the graph: visual attribute correspondence, data variable correspondence, visual encoding mapping, and visual insight correspondence. A complete list of these relationships can be found in Figure 2 The specific instructions are as follows:

[0033] Visual attribute correspondence: This relationship is intended to represent the attribute value of a visual element. It connects a visual element to a visual element attribute value, such as "the color of the bar is blue."

[0034] Data variable correspondence: It focuses on specific instances of data variables, for example, "England" and "United States" can be values ​​of the "country" variable. We connect data variable values ​​with data variables to indicate the relationship between variables and values, such as "2011 is the value of year".

[0035] Visual encoding mapping relationship: The visual encoding mapping expresses the semantic information conveyed by the underlying data and models the mapping from visual element attribute values ​​to data variables or data variable values, such as "the color of the bar is blue, representing England".

[0036] Visual Insight Correspondence: It is used to represent complete visual insights, usually describing the underlying features between different data variables and data variable values. We establish relationships between existing visual insights and corresponding data variables / data variable values, such as "the trend of GDP is rising".

[0037] Step 2): In the process of converting a chart image into a knowledge graph, the chart image must first be classified. In order to verify the accuracy of the chart type classification, this embodiment designs an experiment to explore the performance of the chart classification model. Specifically, this embodiment adopts a widely used deep learning model, the convolutional neural network (CNN), specifically the ResNet50 network, to classify chart images. First, the input chart image is divided into four categories: bar chart, line chart, pie chart and scatter chart. In the experiment, the present invention uses a ResNet model pre-trained on the Imagenet dataset and fine-tunes it using the Adam optimizer, with the learning rate set to 0.0005. In this way, the chart classification model can accurately classify the chart image into one of the above four types. The experimental results show that on the dataset of the present invention, the model achieved an accuracy of 87.2%. In addition, the present invention also compares other models to verify the consistency of this result. These experimental results further confirm the effectiveness and accuracy of ResNet50 in the chart classification task.

[0038] Step 3): Figure 3 As shown, after completing the graph parsing, the present invention performs entity classification and relationship construction based on the extracted content, and then establishes the final graph semantic network.

[0039] In the entity classification stage, the present invention preliminarily divides the marked elements in the chart into five different entity types. Based on the entity categories defined in step 1, the graphic marks in the chart are regarded as multiple entities VE, and the attribute values ​​of each mark (such as color and size) are classified as entity types VEPV. As for DV and DVV, the title of the chart component usually describes the variable, while the component label represents the specific variable value. Therefore, according to the role of the text in the chart, the text content in the chart is divided into two categories: DV and DVV. Finally, the visual insights cover the key semantic information in the chart that is difficult to obtain directly from the original data. The present invention combines the attribute values ​​of the graphic marks into data tuples (for example, ⟨height, index, color>). Further, the color and index in the tuple are converted into rows and columns, and the height is used as the data value, and then these tuples are converted into a table format as the input required by QuickInsights. Next, the insights in the chart are extracted, and each extracted visual insight is treated as an independent entity.

[0040] In the relationship construction stage, the present invention adopts a rule-driven approach to construct four relationships: visual attribute correspondence, data variable correspondence, visual encoding mapping, and visual insight correspondence. Visual attribute correspondence describes the label attribute, such as the height of the bar in a bar chart; data variable correspondence processes variable instances, such as the x label value is an instance of the x title variable; visual encoding mapping is used to establish the association between the label attribute and the data variable, such as determining the association by calculating the distance or color similarity; visual insight correspondence is used to establish the association between the data variable and the extracted visual insight, such as the correlation between the x-axis and the y-axis title.

[0041] In order to evaluate the effectiveness of the framework of the present invention in constructing a knowledge graph-based chart representation, the present invention quantitatively evaluates its core components, object recognition and optical character recognition (OCR). Specifically, the method selects a subset containing rich annotation information of chart elements from PlotQA as a corpus, which contains three types of charts: bar charts, line charts, and scatter charts. Considering that the main difference between scatter charts and line charts is whether the points are connected, the method replaces the dot charts with more common pie charts and scatter charts. A batch of charts with annotated information are automatically generated using the matplotlib package, and the final corpus contains four types of charts, each type of up to 10,000. This embodiment randomly selects 70% of each chart type as a training set, 15% as a validation set, and the remaining 15% as a test set to evaluate the performance of object recognition and OCR.

[0042] In the object recognition evaluation part, the present invention uses several key indicators, including mean average precision (mAP), precision (Precision) and recall (Recall). The evaluation results are shown in Table 1 below.

[0043] Table 1 Object recognition evaluation results

[0044] Table 1 shows that this embodiment performs well in identifying various types of charts, with mAP50, precision, and recall scores generally exceeding 0.9; mAP50-95 scores mostly exceeding 0.7. Although the recognition scores of line charts and scatter plots are relatively low, the mAP remains above 0.7, demonstrating the reliability and accuracy of the system on these types.

[0045] For optical character recognition evaluation, the present invention evaluates the recognition accuracy of each category of text by comparing the predicted text content with the annotated text content. The results show that the OCR accuracy of most text categories exceeds 70%. Although OCR errors will not affect the construction of the graph structure, they may affect the accuracy of the extracted variable names, values ​​and other elements, thereby affecting the results of downstream tasks. In the future, the accuracy of OCR can be improved by introducing more comprehensive training data and improving the model.

[0046] In order to demonstrate the practicality of the present embodiment in semantic-aware chart retrieval, the present invention designs a chart retrieval method, which can better meet the user's retrieval needs and highlight the flexibility and user-friendliness of the present embodiment in managing chart databases. Specifically, the method enhances retrieval efficiency and user interaction experience by formatting user input and applying conditional filtering. The user needs to provide the chart type, DV / DVV, encoding relationship, and existing visual insights in sequence. First, the user specifies the chart type, key entities, and their relationships based on the required sentence order; then, irrelevant charts are filtered out according to the specified chart type; then, a variable dictionary is built using the pre-extracted DV and DVV to match the variables input by the user, further narrowing the scope of the target chart; finally, the chart knowledge graph that meets the conditions is traversed to match the encoding relationship. In order to prove that this method is superior to the keyword-based retrieval method, the present invention conducted ten retrieval experiments and invited three experts in the field of visualization to evaluate the retrieval results. As Figure 4 As shown in the figure, the experimental results show that the time consumption and user rating of this method are comparable to those of the keyword method when retrieving variables, while in terms of semantic retrieval, the average score of the present invention exceeds 4.0, which is significantly higher than the keyword method, and the time consumption is only increased by a few seconds. This shows that the chart retrieval method supported by chartKG can more accurately meet the diverse retrieval needs of users without sacrificing time efficiency.

[0047] In order to verify the effectiveness of this embodiment in the visual question answering (VQA) task, the present invention designs two experiments. First, the entities and relationships involved in the question are located by using a graph matching method, and then the structured representation of the knowledge graph is used to enhance the interpretability of the VQA model. Specifically, the method adopts the following steps:

[0048] Question type determination: Segment the user question text into words and determine the question type through keyword matching.

[0049] Entity relationship location: Use similarity matching in the knowledge graph to locate the entities and relationships involved in the question. For data comparison problems, the present invention also groups entities according to the location information in the question.

[0050] Knowledge search and reasoning: Apply specific knowledge search rules and reasoning rules to find answers based on the type of question. For example, a data sequence X can be obtained through the path from visual element (VE) to visual element property value (VEPV) to data value (DV). The grouping information of the data can be obtained based on the path from VE to VEPV to DVV, where DVV represents the order or category of the visual elements in the diagram.

[0051] It can be seen that Figure 5 As shown, this embodiment verifies the effectiveness of the present invention in the VQA task through experiments, and demonstrates its advantages in accuracy and time efficiency by comparing with other methods. This embodiment not only improves the performance of VQA, but also enhances the interpretability of the model due to its structured representation.

[0052] In the experiment, the present invention used the T5 model as a baseline and compared it with the method provided in this embodiment. The present invention also constructed a VQA dataset and trained and evaluated T5 and the method provided in this embodiment on the dataset. The experimental results show that the method provided in this embodiment is superior to the T5 model based on deep learning in terms of accuracy and time efficiency.

[0053] Through these experiments, the present invention demonstrates the potential of ChartKG in visual question answering tasks and proves its value in improving the performance and interpretability of VQA models. These results provide strong support for applying the method provided in this embodiment to a wider range of visual data analysis tasks in the future.

Claims

1. A method for constructing a knowledge graph based on a chart image, characterized in that: The specific steps are as follows: Step 1: Establish a classification system for entity types in graph images and semantic relationship types between entities; Step 2: Identify the chart category of the chart image, extract text information in the chart image, and divide multiple entities in the chart image; Step 3: Establish semantic relationships between different entities, and build a knowledge graph based on the entities extracted from the chart image and the semantic relationships between different entities.

2. The method for constructing a knowledge graph based on a chart image according to claim 1, characterized in that: The entity types in the chart image include visual elements, visual element attribute values, data variables, data variable values, and visual insights.

3. The method for constructing a knowledge graph based on a chart image according to claim 2, characterized in that: The visual element includes a graphic mark; the visual element attribute value includes a height and a line color of the graphic mark.

4. The method for constructing a knowledge graph based on a chart image according to claim 2, characterized in that: The data variables include the coordinate axes and the title of the legend in the chart image; the data variable values ​​include the numerical labels on the coordinate axes and the legend; the data variable values ​​are obtained by extracting text information in the chart image.

5. The method for constructing a knowledge graph based on a chart image according to claim 2, characterized in that: The visual insights include semantic information inferred from the original chart image.

6. The method for constructing a knowledge graph based on a chart image according to claim 1, characterized in that: The entity relationships in the diagram image include visual attribute correspondence relationships, data variable correspondence relationships, visual encoding mapping relationships, and visual insight correspondence relationships.

7. The method for constructing a knowledge graph based on a chart image according to claim 1, characterized in that: The text information in the chart image is extracted by an optical character recognition method; and the chart category of the chart image is identified by a convolutional neural network.

8. A diagram image retrieval system, characterized in that: It includes a chart recognition module, a text extraction module, a knowledge graph construction module and a retrieval execution module; the chart recognition module is used to identify the chart type in the image; the text extraction module is used to extract text information in the chart image; the knowledge graph construction module is used to execute the knowledge graph construction method based on chart image described in claim 1; the retrieval execution module is used to match the knowledge graph with the retrieval information input by the user to obtain a matching result.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: The memory stores computer programs; the processor executes a method for constructing a knowledge graph based on a chart image as described in any one of claims 1 to 7.

10. A readable storage medium storing a computer program; characterized in that: When the computer program is executed by a processor, it is used to implement a method for constructing a knowledge graph based on a chart image as described in any one of claims 1 to 7.