Question Answering for Data Visualization

By generating a feature map for data visualization and encoding query feature vectors, the problem of difficulty in searching data visualization content in traditional technologies is solved, and efficient search and interpretation of information within data visualization is achieved.

CN109960734BActive Publication Date: 2025-06-06ADOBE INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201811172058.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-12-22
Filing Date
2018-10-09
Publication Date
2025-06-06
Estimated Expiration
2038-10-09

AI Technical Summary

Technical Problem

Traditional computer search technology is difficult to access or utilize the underlying data contained in data visualization, resulting in users not being able to benefit from the desired information when searching for information.

Method used

By generating a feature map that characterizes the data visualization and encoding the query into a feature vector, the answer is determined based on the feature map and the query feature vector, and the answer is determined.

Benefits of technology

It realizes efficient search and interpretation of information contained in data visualization. Users can submit queries through the query bar and receive the expected answers, solving the problem that traditional technology cannot access data visualization content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN109960734B_ABST
    Figure CN109960734B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to question answering for data visualization. Systems and techniques are described for providing question answering using data visualizations, such as bar graphs. Such data visualizations are typically generated from collected data and provided within an image file that shows the underlying data and relationships between data elements. The described techniques analyze a query and the associated data visualizations and identify one or more spatial regions within the data visualizations where answers to the queries can be found.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to question answering for data visualization. Background Art

[0002] It is often desirable to present information using data visualizations such as bar graphs, pie charts, and various other forms of visualizations. Such data visualizations are advantageously able to convey large amounts and various types of information in a compact and convenient form. Furthermore, such data visualizations are versatile enough to illustrate information for audiences ranging from elementary school students to advanced professional users in nearly any field of endeavor.

[0003] Although data visualizations are specifically designed to convey large amounts of underlying data in a manner that is visually observable to human readers, conventional computer search techniques are often largely or completely unable to access or utilize the underlying data. For example, document files often include data visualizations as image files; e.g., embedded within a larger document file. There are many techniques for searching text within such document files, but such techniques will ignore the image files that display the data visualization.

[0004] Thus, for example, a user may submit a query using conventional techniques for desired information to be applied across multiple documents, and if the desired information is included in a data visualization that is included in one or more of the searched documents, the query will not be satisfied. Similarly, a user's search within a single document may be unsuccessful if the desired information is included within a data visualization. As a result, in these and similar scenarios, such a user may not benefit from the desired information even if the desired information is included in the available document(s). Summary of the invention

[0005] According to one general aspect, a computer program product is tangibly embodied on a non-transitory computer-readable storage medium and includes instructions. The instructions, when executed by at least one computing device, are configured to cause the at least one computing device to identify a data visualization (DV), and to generate a DV feature map representing the data visualization, including maintaining a correspondence of spatial relationships of features of a mapping of the data visualization within the DV feature map with corresponding features of the data visualization. The instructions are also configured to cause the at least one computing device to identify a query for an answer included within the data visualization, encode the query as a query feature vector, generate a prediction of at least one answer location within the data visualization based on the DV feature map and the query feature vector; and determine an answer from the at least one answer location.

[0006] According to another general aspect, a computer-implemented method includes receiving a query for a data visualization (DV), encoding the query as a query feature vector, and generating a DV feature map representing at least one spatial region of the data visualization. The computer-implemented method also includes generating a prediction of at least one answer location within the at least one spatial region based on a combination of the at least one spatial region of the data visualization and the query feature vector, and determining an answer to the query from the at least one answer location.

[0007] According to another general aspect, a system includes: at least one memory including instructions; and at least one processor operably coupled to the at least one memory and arranged and configured to execute the instructions. The instructions, when executed, cause the at least one processor to: generate a visualization training data set, the visualization training data set including a plurality of training data visualizations and visualization parameters, and query / answer pairs for the plurality of training data visualizations, and train a feature map generator to generate a feature map for each of the training data visualizations. The instructions, when executed, are also configured to train a query feature vector generator to generate a query feature vector for each of the queries in the query / answer pair, and train an answer position generator to generate an answer position within each of the training data visualizations for the answer to the corresponding query in the query / answer pair based on the output of the trained feature map generator and the trained query feature vector. The instructions, when executed, are also configured to input a new data visualization and a new query into the trained feature map generator and the trained query feature vector to obtain a new feature map and a new query feature vector, and to generate a new answer position within a new data visualization for the new query based on the new feature map and the new query feature vector.

[0008] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 is a block diagram of a system for question answering for data visualization;

[0010] Figure 2 It is shown Figure 1 A flowchart of an example operation of the system;

[0011] Figure 3 It is shown Figure 1 A block diagram of a more detailed example implementation of a system;

[0012] Figure 4 It is shown Figure 3 A flowchart of an example operation of the system;

[0013] Figure 5 It is shown Figure 3 A block diagram of a more detailed example implementation of a system;

[0014] Figure 6 Shows that it can be combined Figures 1 to 5 a first example data visualization utilized by an example implementation of a system;

[0015] Figure 7 Shows that it can be combined Figures 1 to 5 a second example data visualization utilized with an example implementation of a system showing a result of an applied attention graph;

[0016] Figure 8 Shows that it can be combined Figures 1 to 5 a third example data visualization utilized by an example implementation of a system of; and

[0017] Fig. 9 is shown for use Figures 1 to 5 A flowchart and associated screenshots of a technique for re-capturing data and visualizing the data. DETAILED DESCRIPTION

[0018] This document describes systems and techniques for providing answers to questions using data visualizations. Such systems and techniques overcome the technical challenges of previous systems and techniques and improve the process(es) of performing related automated processing. For example, from within an application, a query can be submitted and a search can be performed on the content of a data visualization included in a document(s) of identification, rather than performing a text search (e.g., a string-based or semantic matching of document terms with query terms) only within or across documents(s). The described data visualization search techniques use more efficient, faster, more accurate, and more complete algorithms(s) than other algorithms that attempt to search any(s) of data visualizations. In addition, the data visualization search techniques provide new computer functions, such as requesting and finding data contained within (represented by) a data visualization (even if the data visualization is an image file that includes only pixel values ​​encoding an image of the data visualization) and providing search results that are identified in conjunction with processing of the data visualization regarding at least one query.

[0019] The systems and techniques provide a user interface within an application to enable users to submit queries, including various types of queries that human users typically consider when viewing and interpreting data visualizations. For example, such queries can include structural questions related to the structure of the data visualization (s) (such as how many bars are included in a bar graph), value questions (such as the specific value of a specific bar within a bar graph), comparison questions (such as which bar within a bar graph is the largest or smallest), and label questions (such as identifying and reading the contents of a label or other identifier of a specific bar of a bar graph).

[0020] For example, a user may provide a verbal query, such as, "Which South American country has the highest GDP?". The query may be provided with respect to a specific document or with respect to a collection of documents or with respect to any search engine that is tasked with searching through a defined search space. The query need not be directed to any particular data visualization or type of data visualization, and indeed, the user submitting the query may not be aware that a (relevant) data visualization even exists within the search space. Nevertheless, the user may be provided with the desired answer because the present technology enables finding, analyzing, and interpreting information that is implicit within the data visualization but not available to current search technology. As a result, the described technology is able to provide users with information that might otherwise be found by or available to these users.

[0021] The example user interface may thus be operable to generate and visually display the desired answer independently of the query's data visualization itself or in conjunction with (e.g., overlaying) the query's data visualization itself. In other examples, the described systems and techniques may be operable to capture data underlying (multiple) data visualizations and thereafter re-render the same or different data visualizations in an editable format.

[0022] As described in detail below, example techniques include using various types of machine learning and associated algorithms and models, where synthetic data sets are generated for the type of data visualization to be searched, including synthetic queries and answers. The synthetic data sets are then used as training data sets to train multiple models and algorithms.

[0023] For example, a model such as a convolutional neural network can be trained to generate a feature map for data visualization, and a long short-term memory (LSTM) network can be used to encode the query into a feature vector. The feature map and the query feature vector can then be combined using an attention-based model and associated techniques, such as by generating a visualization feature vector from the feature map and then concatenating the visualization feature vector and the query feature vector.

[0024] In this manner, or using additional or alternative techniques, the location of the expected answer within the image of the data visualization(s) can be predicted. One or more techniques can then be used to determine the answer from the predicted or identified answer location. For example, optical character recognition (OCR) can be used. In other examples, additional machine learning models and algorithms can be used to provide an end-to-end solution to read the identified answer from the data visualization (e.g., character by character).

[0025] In other example implementations, the described systems and techniques can be used to examine a data visualization and recover all data or all desired data used to generate the examined data visualization, e.g., the underlying source data for the data visualization. For example, by using the query / answer techniques described herein, a set of queries can be used to systematically identify source data. The source data can then be used for any desired purpose, including plotting the source data into different (types of) data visualizations. For example, a bar chart can be analyzed to determine the underlying source data, and then drawing software can be used to re-plot the source data into a pie chart or other type of data visualization.

[0026] Thus, for example, the systems and techniques described in this document facilitate searching across a corpus of documents or other files containing data visualizations to find information that would otherwise be missed by the user performing the search. In addition, intelligent document-based software can be provided that can provide documents for viewing, editing, and sharing, including finding and interpreting information for readers of the document(s) that might otherwise be missed or misunderstood by such readers.

[0027] Additionally, the systems and techniques described herein advantageously improve upon the prior art. For example, as described above, computer-based searches are improved, for example, by providing computer-based searches of data visualizations and the data included therein. Furthermore, the systems and techniques can be used to interpret data visualizations in a more automated and more efficient and faster manner.

[0028] In particular, for example, the described techniques do not rely on the classification of the actual content of the data visualizations used for training or the data visualizations being searched. For example, the described techniques do not have to classify one of two or more known answers to a pre-specified query, for example, as a "yes" or "no" answer to a question such as "Is the blue bar the largest bar in a bar graph?" Instead, the described techniques are able to find the answer or determine the information needed to answer the answer within the (multiple) spatial locations of the image of the data visualization itself, including first considering the entire area of ​​the data visualization being considered. Thus, the described techniques can answer newly received queries even when such queries (or associated answers) have not been seen before during the context of the training operation. For example, although the training data may include actual bar graphs (i.e., "real world" bar graphs representing actual data), or other data visualizations if desired, the training set may also include meaningless words and / or random content. As a result, it is very simple to generate a training data set of the desired type without having to create or compile existing actual data visualizations.

[0029] Figure 1 1 is a block diagram of a system 100 for question answering using data visualization. System 100 includes a computing device 102 having at least one memory 104, at least one processor 106, and at least one application 108. Computing device 102 can communicate with one or more other computing devices via a network 110. For example, computing device 102 can communicate with a search server 111 via network 110. Computing device 102 can be implemented as a server, a desktop computer, a laptop computer, a mobile device such as a tablet device or a mobile phone device, and other types of computing devices. Although a single computing device 102 is shown, computing device 102 can represent multiple computing devices that communicate with each other, such as multiple servers that communicate with each other to perform various functions over a network. In many of the following examples, computing device 102 is described or can be understood to represent a server.

[0030] The at least one processor 106 may represent two or more processors on the computing device 102 executing in parallel and utilizing corresponding instructions stored using the at least one memory 104. The at least one memory 104 represents at least one non-transitory computer-readable storage medium. Thus, similarly, the at least one memory 104 may represent one or more different types of memory utilized by the computing device 102. In addition to storing instructions that allow the at least one processor 106 to implement the application 108 and its various components, the at least one memory 104 may be used to store data.

[0031] The network 110 may be implemented as the Internet, but other different configurations may be used. For example, the network 110 may include a wide area network (WAN), a local area network (LAN), a wireless network, an intranet, a combination of these networks, and other networks. Of course, although the network 110 is shown as a single network, the network 110 may be implemented to include multiple different networks.

[0032] Application 108 can be accessed directly at computing device 102 by a user of computing device 102. In other implementations, application 108 can be run on computing device 102 as a component of a cloud network, where the user accesses application 108 from another computing device (e.g., user device 112) through a network such as network 110. In one implementation, application 108 can be a document creation or viewer application. In other implementations, application 108 can be a standalone application designed to work with a document creation or viewer application (e.g., running on user device 112). Application 108 can also be a standalone application used to search for multiple documents created by (multiple) document creation or viewer applications. In other alternatives, application 108 can be an application that runs at least partially in another application such as a browser application. Of course, application 108 can also be a combination of any of the above examples.

[0033] exist Figure 1 In the example of , user device 112 is shown as including display 114 having document 116 drawn therein. As described above, document 116 can be provided by a document reader application, which can include at least a portion of application 108 or can utilize capabilities of application 108 to benefit from the various visual search techniques described herein.

[0034] Specifically, document 116 is shown to include a query field 118 in which a user may submit a query related to one or more data visualizations (in Figure 1 120). For example, in a context where document 116 is a single, potentially large or lengthy document that a user is viewing, a query bar 118 may be provided to enable the user to submit one or more questions, requests, or other queries about the content of document 116. In these and similar examples, such a lengthy document may include potentially large data visualizations, and the user may not be aware of the location or existence of a data visualization 120 containing the desired information within the context of the larger document 116.

[0035] Thus, using the various techniques described herein, a user may nonetheless submit a query for the desired information via the query bar 118 and may be provided with an identification of the data visualization 120, including the desired information as well as the specific requested information. Various techniques for providing the requested information are shown and described herein or will be apparent to those skilled in the art.

[0036] Furthermore, it should be appreciated that the techniques described herein can be applied across multiple documents or other types of content including data visualizations. For example, query bar 118 can be implemented as a standard search bar within a search application that is configured to search across multiple documents or other files. In such a context, a user of the search application can thus be provided with, for example, identified documents and included data visualizations, such as data visualization 120, to satisfy the user's search request.

[0037] Therefore, it should be appreciated that, as a non-limiting term, the term query should be understood to include any question, request, search or other attempt to locate a desired answer, search result, data or other information, and can be used interchangeably with them. Any such information can be included in any suitable type of digital content, including documents, articles, textbooks, or any other type of text files that can include and display images as described above. Of course, such digital content can also include a single image file that includes one or more images of one or more data visualizations such as data visualization 120. Therefore, such digital content can be included in a single file with any suitable format (such as, for example, .pdf, .doc, .xls, .gif or .jpeg file formats). In addition, such data visualization can be included in a video file, and the search techniques described herein can be applied to each frame of such a video file containing a data visualization such as data visualization 120.

[0038] exist Figure 1 In the example of , the data visualization 120 is shown as a simplified bar chart, where the first bar 122 has the label "ABC" and the second bar 124 has the label "XYZ". Of course, again, it should be appreciated that Figure 1 The simplified data visualization 120 should be understood as representing the Figure 1 These are non-limiting examples of many different types of data visualizations that may be utilized in the context of system 100 .

[0039] For example, many different types of bar charts can be used, some examples of which are referenced below. Figures 6 to 8In addition, other types of data visualizations can be used, such as pie charts, scatter plots, networks, flow charts, tree maps, Gantt charts, heat maps, and many other types of visual representations of data. Figure 1 Data visualization 120 and other data visualizations described herein should generally be understood to represent any suitable technique for encoding digital data using points, lines, bars, segments, or any other combination of visual elements and associated numerical values, along with appropriate labels for the elements and values ​​that together convey the underlying information in a graphical, schematic format.

[0040] As referenced above and described in detail herein, the system 100 enables a user of a user device 112 to submit queries and receive answers regarding a data visualization represented by a data visualization 120. For example, such queries may include structural queries (e.g., how many bars are included?), value queries (e.g., what is the value of XYZ?), or comparison questions (e.g., what is the largest bar in a bar graph?). As can be observed from the previous examples, some queries may be relatively specific to the type of data visualization 120 in question. However, in other examples, the system 100 may be configured to provide answers to natural language queries that are independent of the type of data visualization being examined and may even be submitted with respect to any specific data visualization that contains the desired information. For example, if the data visualization 120 represents a chart of gross domestic product (GDP) for multiple countries, a user may submit a natural language query such as "Which country has the highest GDP?" and thereby receive an identification of a specific country associated with a specific bar of the bar graph of the example data visualization 120.

[0041] To provide these and many other features and benefits, the application 108 includes a visualization generator 126 that is configured to provide sufficient information to a model trainer 128 to ultimately train elements of a data visualization (DV) query handler 130. In this manner, the system 100 enables the DV query handler 130 to input a previously unknown data visualization, such as the data visualization 120, and one or more associated queries, and thereafter provide a desired answer that has been extracted from the data visualization 120. More specifically, as shown, the visualization generator 126 includes a parameter handler 132, and the visualization generator 126 is configured to generate a visualization training data set 134 using parameters received through the parameter handler 132.

[0042] For example, continuing with the example where the data visualization 120 is a bar graph, the parameter handler 132 may receive parameters from a user of the application 108 related to the generation of a plurality of bar graphs to be included in the visualization training data set 134. Such parameters may thus include virtually any characteristics or constraints of the bar graphs that may be relevant for training purposes. As non-limiting examples, the range of such parameters may include the number of bars to be included, the values ​​of the X, Y, and / or Z axes for the various bar graphs to be generated, the characteristics of the key / legend of the bar graphs, and various other aspects.

[0043] The received parameters may also include queries, types of queries, or query characteristics that may be relevant to the training purpose. Likewise, such queries may be more or less specific to the type of data visualization being considered, or may simply include natural language questions or types of natural language questions that may be answered using the techniques described herein.

[0044] Using the received parameters, visualization generator 126 can generate visualization training data set 134, where visualization training data set 134 is understood to include a complete collection of data, data visualizations generated from the data, queries or query types that can be applied to the data and / or data visualizations, and answers or answer types that correctly respond to the corresponding queries or query types. In other words, visualization training data set 134 provides an internally consistent data set, where, for example, each query corresponds to specific data, where the data is represented using any corresponding data visualizations, and answers to the queries are associated with each of the queries, data, and data visualizations.

[0045] In other words, visualization training dataset 134 provides "ground truth" of known query / answer pairs that are related through the generated data and corresponding data visualizations. However, as described in detail below, visualization training dataset 134 need not include or represent any actual or correct real-world data. In fact, visualization training dataset 134 may include random or meaningless words and values ​​that are related to each other only in the sense that, as referenced, they are internally consistent with respect to the data visualizations generated in conjunction therewith.

[0046] For example, if data visualization 120 is used and included within visualization training dataset 134, it may be observed that the label "XYZ" for a bar has no obvious underlying relationship to the value 3, other than the fact that bar 124 has the value 3 within data visualization 120. As described in detail herein, DV query handler 130 may be configured to analyze the entirety (at least a portion) of data visualization 120 prior to attempting to answer a specific query, and thereby identify and utilize specific locations within data visualization 120 in order to predict where and how to extract desired information. Thus, the actual content or meaning of labels, values, or other words or numbers within data visualization training dataset 134 is not necessary to derive its desired characteristics within the context of visualization training dataset 134. Of course, such content, including the semantics, syntax, and meaning of the included words or images, may be considered if desired, including in the context of answering natural language queries. In any case, at least from the above description it can be observed that the described implementations of visualization generator 126 , model trainer 128 , and DV query handler 130 can be configured to provide desired query answers even when neither the query nor the answer in question is included in visualization training dataset 134 .

[0047] The ability of visualization generator 126 to synthetically generate visualization training data set 134 (rather than having to rely on, create, or identify real-world data visualizations) means that very large and comprehensive training data sets can be quickly and easily obtained for any of the various types of data visualizations and related parameters that users of application 108 may desire. As a result, sufficient training data can be provided to model trainer 128 to produce reliable, accurate, and efficient operation of DV query handler 130, as described below.

[0048] In the following description, the model trainer 128 is configured to use the visualization training data set 134 to provide training for one or more neural networks and related models or algorithms. Figure 1 In the example of , several examples of such neural networks are provided, each of which is configured to provide specific functionality regarding the operation of the DV query handler 130. Specifically, as shown, the model trainer 128 can be used to train a convolutional neural network (CNN) 136, a long / short term memory (LSTM) encoder 138, and an attention model 140. Specific examples and details regarding the CNN 136, the LSTM encoder 138, and the attention model 140 are provided below, and additional or alternative neural networks may also be utilized.

[0049] Typically, such neural networks provide a computational model used in machine learning that is composed of nodes organized into layers. Nodes may also be referred to as artificial neurons, or just neurons, and perform functions on provided inputs to produce a certain output value. Such neural networks typically require a training cycle to learn the parameters (e.g., weights) used to map inputs to specific outputs. As described above, the visualization training data set 134 provides "ground truth" to training examples that are used by the model trainer 128 to train various models 136, 138, 140.

[0050] The input to the neural network can be constructed in the form of one or more feature vectors. Typically, a feature vector is an array of numbers, the array having one or more dimensions. Such feature vectors allow, for example, digital or vector representation of words, images, or other types of information or concepts. By representing concepts digitally in this way, otherwise abstract concepts can be processed computationally. For example, when representing words as feature vectors, in the context of multiple words represented by corresponding feature vectors, a vector space can be defined in which vectors of similar words naturally appear close to each other in the vector space. In this way, for example, word similarity can be detected mathematically.

[0051] Using such feature vectors and other aspects of the neural network described above, the model trainer 128 can continue to perform training using training examples of the visualization training data set 134, including performing a series of iterative rounds of training, wherein the optimal weight values ​​of one or more mapping functions used to map input values ​​to output values ​​are determined. When determining the optimal weights, the model trainer 128 essentially makes predictions based on the available data, and then uses the available basic facts in conjunction with the visualization training data set 134 to measure errors and predictions. The function used to measure such error levels is generally referred to as a loss function, which is generally designed to sum over relevant training examples and increase the calculated loss if the prediction is incorrect, or reduce / minimize the calculated loss if the prediction is correct. In this way, the various models can be conceptually understood as being trained to learn from the mistakes made during the various iterations of prediction, so that when deployed in the context of the DV query handler 130, the referenced resulting trained model will be fast, efficient, and accurate.

[0052] exist Figure 1In the example of , the model trainer 128 includes a convolutional neural network (CNN) 136, which represents a specific type of neural network that is specifically configured to process images. That is, because such a convolutional neural network explicitly assumes that the input features are images, the attributes can be encoded into the CNN 136, which causes the CNN 136 to be more efficient than a standard neural network while reducing the number of parameters required for the CNN 136 relative to the standard neural network.

[0053] In more detail, the parameters of CNN 136 may include one or more learnable filters, each of which has spatial dimensions of width and height while extending through the entire depth of the input volume. For example, if the input to CNN 136 includes an image, the filter applied to the image may have an example size of 5×5×3, representing a width of 5 pixels, a height of 5 pixels, and a depth dimension of 3 corresponding to the color channels that may be included. To apply CNN 136, each of the one or more filters is passed (in other words, convolved) across the width and height of the filtered pixels of the input image. When the filter is convolved across the width / height and volume of the input image, a dot product or other appropriate calculation may be performed between the entries of the filter and each input position.

[0054] As referenced above with respect to neural networks, the parameters of one or more filters will be learned and adjusted over time so as to be activated in response to a desired type of visual feature (e.g., image edges, image orientation, color, or some other image aspect of the CNN 136 being trained). Thus, once the CNN 136 has been successfully trained, the result will be, for example, a set of parameterized filters in a corresponding plurality of layers, each producing a separate 2D feature map, which set of parameterized filters can then be compiled along the depth dimension to produce a total output feature map volume.

[0055] Further with respect to the model trainer 128, the LSTM encoder 138 refers to a specific type of recurrent neural network (RNN) designed to exploit information contained within a sequence of related information. That is, a recurrent neural network is described as recursive because the same or similar operation is performed on each element of the sequence, where each output depends on previous calculations.

[0056] For example, such a sequence may include sentences or questions included in each query of the visualization training dataset 134, because a typical sentence represents a sequence of related information, and subsequent portions of such a sequence / sentence may be more accurately inferred by considering earlier portions of the sentence. For example, this concept may be understood by considering that a sentence beginning "People from Germany are fluent in ... " may be used to more easily determine the word "German" than attempting to infer the word German in isolation from other text.

[0057] Thus, as can generally be observed, such recurrent neural networks rely on information derived earlier, which therefore must be stored over time in order to be useful later in the process. In practice, it may happen that there is a large distance between the current word and the information necessary to determine the current word (e.g., a large number of words within a sentence). Thus, the recurrent neural network may need to store more information than is practical or useful to store with respect to optimally determining the current information.

[0058] In contrast, LSTM 138 is specifically designed to avoid this type of long-term dependency problem and is able to selectively store useful information for a period of time long enough to successfully determine the information currently under consideration. Figure 1 In the example of , an input query encoded as a feature vector from the visualization training dataset 134 can be used to train the LSTM encoder 138, which minimizes a loss function relative to the ground truth answer included in the visualization training dataset 134 for the corresponding query. As a result, as described in detail below, the DV query handler 130 can be configured to receive new queries about newly determined data visualizations, such as queries received about the data visualization 120. The DV query handler 130 is therefore configured to encode the newly received query as a feature vector, even when the specific content of the received query is not necessarily or explicitly included in the visualization training dataset 134.

[0059] Finally, about Figure 1 For the model trainer 128 of the attention model 140, the attention model 140 represents a mechanism for producing an output that depends on a weighted combination of all input states rather than on a particular input state (e.g., rather than on the last input state). In other words, the weights of the weighted combination of all input states represent the extent to which each input state should be considered for each output. In practice, the weights can be normalized to sum to a value of 1 to represent a distribution over the input states. In other words, the attention model 140 provides a technique for looking at the entire input state before determining whether each input state should be considered important and to what extent it should be considered important.

[0060] In operation, the attention model 140 may be trained using the outputs of the CNN 136 and the LSTM encoder 138. In other words, the input to the attention model 140 during training may include the encoded feature maps of a specific data visualization of the visualization training dataset 134, while the output of the LSTM encoder 138 includes a feature vector encoded from a corresponding query of the same data visualization involving the CNN 136. As described herein, the visualization training dataset 134 also includes known answers to the considered query with respect to the considered data visualization, so that the assignment of weights / weight values ​​can be judged and the loss function of the attention model 140 can be minimized.

[0061] Of course, the examples of the model trainer 128 should be understood to be non-limiting, as various or alternative types of neural networks may be utilized. For example, multiple convolutional neural networks may be utilized, each convolutional neural network being trained to identify different image aspects of one or more input images. Other aspects of the model trainer 128 are provided below as examples, for example, with respect to Figures 3 to 5 , or it will be clear to those skilled in the art.

[0062] Once training has been completed, DV query handler 130 may be deployed to receive new data visualizations and associated queries. In operation, for example, DV query handler 130 may include DV feature map generator 142 that utilizes trained CNN 136 to process an input data visualization, such as data visualization 120. As described above, DV feature map generator 142 is thus configured to generate a feature map representing a feature vector of data visualization 120.

[0063] Meanwhile, the query feature vector generator 144 may be configured to utilize the LSTM encoder 138 trained by the model trainer 128. In other words, the received query may be encoded into a corresponding query feature vector by the query feature vector generator 144. For example, a query such as “What is the value of XYZ?” may be received.

[0064] Then, answer location generator 146 can be configured to use trained attention model 140 to combine the feature map of DV feature map generator 142 representing data visualization 120 and the query feature vector generated by query feature vector generator 144 to thereby predict one or more pixel locations within data visualization 120 where an answer to the encoded query feature vector is most likely to be found. More specifically, as described in more detail below, for example, with respect to Figure 3, the attention map generator 148 may be configured to provide the attention weight distribution provided to the attention map generator 148 for multiplication with the feature map produced by the DV feature map generator 142 to thereby enable the feature weight generator 150 to generate an attention-weighted feature map for the data visualization 120 with respect to the received query. The resulting attention-weighted feature map may then be processed to determine a single representative feature vector for the data visualization 120, wherein this resulting DV feature vector may be combined with the query feature vector of the query feature vector generator 144 using the answer vector generator 152 to thereby generate a corresponding answer vector.

[0065] As a result, answer generator 154 may receive the generated answer vector and may proceed to utilize one or more predicted answer locations to thereby output the desired answer. Specifically, for example, answer generator 154 may include boundary generator 156 configured to determine a boundary around a particular location and number of pixels where the answer is predicted to occur. For example, in Figure 1 In the example of , a boundary 157 around label XYZ may be determined by boundary generator 156 .

[0066] In some examples, border 157 may represent an implicit or contextual identification of a location within data visualization 120 used internally by system 100. In other example implementations, view generator 158 of answer generator 154 may be configured to explicitly show a visual border (e.g., a rectangular box) around label XYZ within data visualization 120.

[0067] Similarly, the view generator 158 can be configured to draw various other types of elements that are overlaid on top of the data visualization 120 in order to convey the desired answer. For example, as shown, if the query received is "Which bar has the value 3?", the boundary generator 156 and the view generator 158 can be configured to draw and draw a visible rectangular box as a boundary 157. In another example, if the query is "What is the value of XYZ?", the boundary generator 156 and the view generator 158 can be configured to draw a dashed line 159 that shows that the level of the bar 124 corresponding to the label XYZ is equal to the value 3.

[0068] Of course, the manner in which answers or potential answers are provided to a user of user device 112 may vary significantly, depending on the context of a particular implementation of system 100. For example, a user submitting a query to a search application may receive a link to a document containing data visualizations that include the desired answer, perhaps with some identification of the individual data visualizations within the linked document(s). In other example implementations, such as when a user searches within a single document 116, the determined answer may be provided simply as text output, e.g., in conjunction with an answer bar displayed adjacent to query bar 118. Other techniques for providing answers are provided as example implementations described below, or will be apparent to those skilled in the art.

[0069] Described in more detail below (e.g., with respect to Fig. 9 ), source data generator 160 can be configured to capture all or a specified portion of the data underlying any particular data visualization, such as data visualization 120. That is, as described above, the normal process for authoring or creating a particular data visualization is to compile or collect the data using one of the many available software applications (such as Excel, MatLab, or many other spreadsheet or computing software applications) that are equipped to generate data visualizations, and then utilize the features of the specific software application to configure and draw the desired data visualization. As described herein, such software applications typically draw the resulting data visualization in the context of an image file, such that the original source data is only indirectly accessible, such as through visual observation of the generated data visualization.

[0070] Using the system 100, the source data generator 160 can apply a structured set of queries to the data visualization 120 to thereby incrementally acquire and determine the source data from which the data visualization 120 is generated. Once this source data has been determined, the source data can be used for any conventional or future use of data of related types, some examples of which are provided below. For example, a user of the user device 112 can convert the data visualization 120 from a bar chart format to any other suitable or available type or format of data visualization. For example, the user can simply select or indicate the data visualization 120, and then request that the data visualization 120 (bar chart) be redrawn as a pie chart.

[0071] In various implementations, the application 108 may use or utilize a number of available tools or resources that may be implemented internally or externally to the application 108. For example, in Figure 1, optical character recognition (OCR) tool 162 may be utilized by answer generator 154 to assist in generating the determined answer. For example, as described above, a query such as “Which bar has a value of 3?” may have answer XYZ in the example of data visualization 120. Answer generator 154 is configured to use the answer vector provided by answer location generator 146 and possibly in conjunction with boundary generator 156 to predict the location within data visualization 120 where answer XYZ occurs.

[0072] However, in Figure 1 In the example of , as described, the answer location includes a plurality of identified pixels within the image of data visualization 120, rather than the actual (e.g., editable) text of the label itself. Therefore, it should be appreciated that OCR tool 162 can be used to perform directional optical character recognition within boundary 157 to thereby output text XYZ as an answer to their received query.

[0073] In other example aspects of the system 100, a collection 164 of source data queries may be pre-stored in available memory. In this manner, the source data generator 160 may utilize the selected / appropriate source data queries in the stored source data queries 164 to analyze a specific corresponding type of data visualization, and thereby output the result source data 166 for the examined data visualization. For example, the application 108 may be configured to search for multiple types of data visualizations, such as bar charts and pie charts. Thus, the source data queries 164 may include a first set of source data queries specifically designed to recover source data for bar charts. The source data queries 164 may also include a second set of source data queries specifically designed to recover source data for pie chart-type data visualizations.

[0074] Likewise, in Figure 1 In the example of , similar to OCR tool 162, visualization generator 168 represents a suitable software application for utilizing source data 166 to generate a new or modified data visualization therefrom. For example, as in the implementation referenced above, data visualization 120 may initially be viewed as a bar chart as shown, and in response to a user request, source data generator 160 may apply a set of suitable source data queries from source data queries 164 to thereby recover the source data of the observed bar chart within source data 166. Thereafter, visualization generator 168, such as Excel, MatLab, or other suitable software applications or portions thereof, may be configured to plot the recovered source data from source data 166 as a pie chart.

[0075] exist Figure 1In the example of , when receiving a data visualization, such as receiving data visualization 120 at DV feature map generator 142, OCR tool 162 (or a similar OCR tool) can be used to predetermine all text (e.g., a list of text) within data visualization 120, and their corresponding positions, including but not limited to bar labels, numbers, axis labels, titles, and information in legends. The position for each text item can be defined using a bounding rectangle containing the text item, such as described herein with respect to bounds 157.

[0076] Thus, inputs to the neural network(s) of system 100, or to system 100 as a whole, may include inputs derived from a list of such text items present in the data visualization. In this manner, the resulting text items may be used to facilitate, enhance, or verify a desired answer, for example, by providing OCR output earlier. For example, when answer position generator 146 or answer generator 154 operates to output a predicted answer position and / or answer, the result may be compared to previously determined text items. Furthermore, once an answer position is predicted, related text items of the previously identified text item may be identified and output as the desired answer.

[0077] Figure 2 It is shown Figure 1 Flowchart 200 of example operations of the system 100. Figure 2 In the example of , operations 202 to 212 are shown as separate, sequential operations. However, it should be appreciated that in various implementations, additional or alternative operations or sub-operations may be included, and / or one or more operations or sub-operations may be omitted. In addition, the following may occur: any two or more operations or sub-operations may be performed in a partially or completely overlapping or parallel manner, or in a nested, iterative, cyclic or branching manner.

[0078] exist Figure 2 In the example of , a data visualization (DV) (202) is identified. For example, Figure 1 The DV feature map generator 142 package may identify the data visualization 120. As described, the DV feature map generator 142 may specifically identify the data visualization 120 in response to a user selection thereof, or may identify the data visualization 120 in the context of searching for multiple data visualizations included within one or more documents or other types of content files. As also described above, the identified data visualization may be included in a separate image file, or in an image file that is embedded within or included in any other type of file, including text, image, or video files.

[0079] A DV feature map representing the data visualization may be generated, including maintaining correspondence of spatial relationships of features of the data visualization mapped within the DV feature map with corresponding features of the data visualization (204). For example, the DV feature map generator 142 may execute the trained CNN 136 to generate the referenced feature map. As described above, it may be trained to apply a set of filters on the data visualization 120 to obtain a collection of feature vectors that together form a feature map for an image of the data visualization 120. Using, for example, the following reference Figure 3 As described in more detail, the trained CNN 136 can be configured to output a feature map having feature vectors having spatial locations and orientations corresponding to the original image.

[0080] A query for an answer included within the data visualization may be identified (206). For example, such a query may be received via a query bar 118 drawn in conjunction with the document 116 or in conjunction with a search application configured to search across multiple documents, as described herein.

[0081] The query may be encoded as a query feature vector (208). For example, query feature vector generator 144 may be configured to implement trained LSTM encoder 138 to output an encoded version of the received query as the described query feature vector.

[0082] A prediction of at least one answer location within the data visualization may be generated based on the DV feature map and the query feature vector (210). For example, answer location generator 146 may be configured to implement trained attention model 140, including input DV feature map generator 142 and the output of query feature vector generator 144. Thus, by considering the query feature vector in conjunction with the DV feature map, answer location generator 146 may determine one or more pixel locations within the original image of data visualization 120 where a desired answer may be determined or derived.

[0083] Thus, an answer may be determined from at least one answer position (212). Figure 1 Answer generator 154 may output or identify one or more image portions within the image of data visualization 120 where the answer may be found, including, in some implementations, drawing additional visual elements (e.g., border 157 or dashed line 159 as a value line) that overlap data visualization 120. In other examples, answer generator 154 may output answer text read from within the predicted answer location, possibly using OCR tool 162, in order to provide the answer.

[0084] Figure 3 It is shown Figure 11 is a block diagram of a more detailed example implementation of the system 100. As shown, an input image 302 of a data visualization (e.g., a bar graph) is passed through a trained CNN 304 to produce a resulting feature map 306, as described above. Figure 3 In a specific example, at least one CNN filter described above can be convolved over the input image 302 so that a feature vector, e.g., a 512-dimensional feature vector, is identified at each position in the 14×14 grid. In other words, at each such position, an encoding of the image 302 is obtained that captures relevant aspects of the content included in the image at that position.

[0085] In the traditional CNN approach, spatial information about different image regions of a digital image may be lost or reduced. Figure 3 In the example of , only the convolutional layers of the VGG-16 network are used to encode the digital image 302. The resulting feature map 306 thus has dimensions 512×14×14, which corresponds to the defined dimensions of the original input image 302, and preserves the spatial information from the original input image 302, as described in more detail below. For example, as indicated by arrow 308, the upper right portion of the input image 302 corresponds directly to the upper right portion of the feature map 306.

[0086] In more detail, some CNNs contain fully connected layers that lose all spatial information. Figure 3 In an example of , convolutional layers and pooling layers are used, where pooling layers generally refer to layers used between consecutive convolutional layers that are used to reduce the spatial size of the representation (e.g., reduce the parameters / computations required in the network). This approach preserves at least some of the spatial information (e.g., some spatial information may be reduced). During decoding, unpooling layers can be used to offset or restore at least some of any spatial resolution that is lost during the above-described type of process.

[0087] Further in Figure 3 , an input question 310 shown as a sample query "Which country has the highest GDP?" may be encoded as part of the process of question characterization 312. For example, as described above with respect to Figure 1 As described, a trained LSTM model may be used to perform question characterization 312 to encode the input question 310 .

[0088] The feature map 306 and the query feature vector encoding the input question 310 may then be input to an attention prediction process 314. Conceptually, the attention prediction 314 outputs an attention map 316 that indicates where attention should be paid within the input image 302 to determine the answer to the input question 310. That is, Figure 3As shown in , the attention map 316 includes a distribution of attention weights at specified locations of the attention map 316. For example, as described above, each attention weight of the attention weight distribution can be assigned a value between 0 and 1 to obtain a standardized distribution across this range of values. At the same time, each location of the feature map 306 represents a feature vector representing the location of the input image 302. Then, by multiplying the attention map 316 with the feature map 306 during the multiplication process 318, the distribution of attention weights of the attention map 316 can be efficiently applied to each feature vector in the feature vectors of the feature map 306.

[0089] For example, with reference to the position of the input image 302 and feature map 306 specified by arrow 308, the following may occur: the corresponding feature vector is assigned an attention weight value of 0, so that the resulting map 320 of the attention-weighted features has a value (importance) of 0 at that position with respect to the current query / answer pair. Therefore, it is predicted that no answer will be found at that position. On the other hand, if a weight of 1 is assigned to the reference to the feature vector from the distribution of attention weights in the attention map 316, then the attention-weighted features corresponding thereto within the map 320 will very likely include information relevant to the predicted answer.

[0090] Finally, about the figure Figure 3 , an image 322 may be generated that illustrates the relative importance levels of various image regions of the original input image 302 with respect to the input question 310. For example, such an image may be explicitly plotted, as shown below in Figure 7 In other examples, the image 322 need not be explicitly generated and drawn for the user, but can simply be used by the application 108 to generate the requested answer.

[0091] Figure 4 It is shown Figure 1 and Figure 3 400 of a more detailed example implementation of a system. Figure 4 In the example of , a synthetic data set is generated (402). For example, visualization generator 126 may receive necessary parameters, such as received via parameter handler 132, and may thereafter generate visualization training data set 134. As described above, the necessary parameters received at parameter handler 132 may vary depending on various factors. For example, as described above, different types of data visualization or different types of desired queries may require different parameters. Parameters may be used to parameterize the generation of desired ranges or values, including random number generators and random word generators.

[0092] Then, the synthetic data set (e.g., visualization training data set 134) can be used to train the desired neural network model, including, for example, CNN, LSTM, and attention models (404). For example, model trainer 128 can be used to train models 136, 138, and 140.

[0093] Subsequently, a data visualization and a query may be received (406). For example, a new data visualization may be received directly from a user or indirectly through a search application. In some implementations, the search process may be configured to make an initial determination regarding the presence or inclusion of data visualizations of a type for generating a visualization training data set 134. For example, such a configured search application may perform a scan of the searched documents to generate a visualization training data set 134, such as a bar chart. Figure 1 The invention also provides a method for determining the presence or inclusion of at least one bar graph in a scene of the invention. In this regard, various techniques may be used. For example, a text search may include a search for the word bar graph, bar graph, etc. In other example implementations, an initial image recognition process may be performed to identify easily recognizable features of a data visualization of a relevant type, such as a vertical axis that may be included within a bar graph.

[0094] Subsequently, a DV feature map may be generated from the data visualization, the DV feature map comprising an array of DV feature vectors (408). For example, Figure 1 The DV feature map generator 142 and / or Figure 3 The convolutional layer in the CNN 304 is used to obtain, for example, a feature map 306.

[0095] A query feature vector may be generated ( 410 ). For example, the query feature generator 144 and / or the question characterization process 312 may be used to obtain a corresponding query feature vector for, for example, the input question 310 .

[0096] An attention map can then be generated from the DV feature map and the query feature vector (412). For example, the attention map generator 148 and / or the attention prediction process 314 can be used to generate an attention map, such as Figure 3 of interest Figure 316. As described above and with respect to Figure 3 As illustrated, the attention map 316 includes a distribution of normalized attention weights, each attention weight being assigned to a particular location of the attention map.

[0097] Therefore, the attention map can be multiplied with the DV feature map to obtain a weighted DV feature vector (414). For example, the feature weight generator 150 can be used to perform Figure 3 A multiplication process 318 is performed to obtain a map 320 of feature vectors of attention weights for the feature vectors of the feature map 306 .

[0098] In addition, Figure 4 In the example of , the weighted DV feature vectors can be subjected to weighted averaging to obtain a composite DV feature vector (416). By generating such a composite DV feature vector for a weighted distribution of feature vectors, it becomes straightforward to concatenate the composite DV feature vector with the original query feature vector to thereby obtain a joint feature / query representation (418) that is used for prediction of the answer position.

[0099] For example, in the above Figure 3 In the example given, where the feature map 306 has a size of 512×14×14, the resulting vector of size 512×14×14 (which can be rewritten as 196×512) retains spatial information, such that each of the 196 vectors of size 512 comes from a specific part of the image, and thus retains the spatial information in the input image. Figure 4 Instead of creating a single vector as described, 196 different numbers that sum to 1.0 can be used to predict the relative importance for each of the 196 regions. For example, if for a given problem and image, only (hypothetical) region 10 is important, then the composite vector will have a value of 1.0 at that position and 0 everywhere else. In this case, the resulting single vector is just a copy of the 10th vector of the earlier 196*512 dimensional vector. If the 9th position is determined to be 70% important and the 8th position is 30% important, then the resulting single vector will be calculated as (0.7*9th vector+0.3*8th vector). Figure 5 yes Figure 1 and Figure 3 A block diagram of an alternative implementation of a system. Figure 1 and Figure 3 In the example, one or more answer positions are predicted. Figure 1 As described, answer generator 154 can be configured to utilize the predicted answer position to extract the desired answer therefrom. For example, boundary generator 156 can be configured to identify a boundary around the predicted answer position, so OCR tool 162 can be used to simply read the desired answer from the predicted position.

[0100] exist Figure 5 In the example of , further refinement steps for this type of bounding box prediction may be included to provide an end-to-end question / answer process for bar graphs that is independent of the availability of OCR tool 162. For example, Figure 5 The system 500 can be configured to decode the information within the predicted answer position to provide the desired answer. For example, the answer can be decoded one character at a time, or, in the case of common answers such as "yes" or "no", the answer can be decoded as a single identified token.

[0101] In addition, an input image 502 of data visualization is received. Figure 5 In the example of FIG. 5 , system 500 also inputs input question 504 (shown as a random sample question, “Which months are the hottest in Australia?”).

[0102] As already described, the query 504 can be processed as part of the question encoding process at the question encoder 506, while the input image 502 can be processed by the convolutional layers of the trained CNN 508. Figures 1 to 4 The various types and aspects of attention prediction 510 described may be implemented using the output of the question encoder 506 and the trained CNN 508 to obtain a predicted coarse position 512.

[0103] In other words, the coarse position 512 represents a relatively high level or first pass prediction at the desired answer position. The question encoder 506 may then also output the encoded question (eg, query feature vector) to a refinement layer 514 which outputs a finer position 516 .

[0104] exist Figure 5 In the example of , it is assumed that the CNN 518 for text recognition has been trained by the model trainer 128 as part of the training process using the visualization training dataset 134. The resulting text recognition process can be further refined by the RNN decoder 520 using the original query feature vector provided by the question encoder 506.

[0105] exist Figure 5 , output 522 represents the output of RNN decoder 520. For example, for a particular label (e.g., "USA"), the RNN will make a prediction at every location in the image for the possible location of the label. However, these locations may be spaced very close to each other so that, for example, several locations may be on the same letter. Thus, in this example, the output may be a number of "U", then a number of "S", and so on. Output 522 represents the return of the output, followed by a transition to a new letter (represented by the dash(es) in 522). For example, the output may be "UUU-SS-AAAA" to convey "USA", generally in Figure 5 is shown as "xx-xx-xx-xxxx".

[0106] Further in Figure 5In , elements 524, 526, 528 are example loss functions that can be used to determine whether the answer has been correctly predicted during training. For example, 524 can represent a least squares (L2) loss, and the coarse position 512 is compared with the "ground truth" position. That is, at each pixel in the image, the coarse position prediction will be tested for matching the ground truth position. For 526, a similar process will be followed, but the fine position 516 is used. CTC loss 528 represents a loss designed for words to see if the characters in the word are correct, such as from output 522. Any errors that occur in the coarse position, fine position, or RNN output will be revealed in the corresponding loss function, which can then be combined into a single loss 530. The single loss 530 can then be used to correct any errors that occur in the neural network.

[0107] Figure 6 is an example bar chart that may be generated and included in the visualization training data set 134. As shown, Figure 6 The example bar graph of includes bar 602, bar 604, bar 606, and bar 608. Each of bars 602 to 608 is associated with a corresponding respective label 610, 612, 614, and 616. As described above, and as Figure 6 As shown in the example of , the various tags 610 to 616 may be generated randomly and need not specifically consider any semantic or contextual meaning of the terms used.

[0108] Thus, label 610 is shown as the word "elod", label 612 is shown as the word "roam", label 614 is shown as the word "yawl", and 616 is shown as the word "does". In this example, in conjunction with the example query shown, i.e., "What is the largest bar?", and the correct answer, i.e., label 610 of "elod", Figure 6 A bar chart of can be included in the visualization training data set 134. As can be observed again, the illustrated bar chart, example questions, and corresponding answers are all internally consistent and thus based on Figure 6 provides suitable ground truth answers for training purposes, even though the label “elod” is a nonsense word.

[0109] In addition, a bounding box 618 may be generated to predict the correct answer as part of the training process so that the training process can assess its appropriate size. For example, the generated bounding box 618 must be large enough to ensure that the entire label 610 is captured, but not so large as to include other text or content.

[0110] therefore, Figure 6An example of a comparison or inference type question / answer is shown, such as determining the largest included bar. Of course, similar types of comparison / inference questions can be utilized, such as questions related to finding the smallest bar, or a comparison of a single bar relative to another single bar (e.g., whether the second bar is greater than or less than the fourth bar).

[0111] In these and similar examples, the answer may be provided as one or more of the individual tags (e.g., one of tags 610 to 616), or the answer may be provided with respect to an identification of such a bar, such as using "the second bar" as a valid answer. More generally, the question may be answered using any suitable format corresponding to the form of the received query. For example, in a similar context, yes / no answers may be provided to the various types of queries just referenced, such as when the query is posed as "Is the first bar the highest bar?"

[0112] Structural questions related to discerning the structure of a particular bar graph may also be considered. For example, such structural questions may include "how many bars are there?" or "how many bars does each label include?" (e.g., in an example training dataset where multiple bars may be included for each label, as shown and described below with respect to Figure 8 described).

[0113] As mentioned above Figure 1 And those that are referenced and illustrated can answer specific value-based questions, such as "What is the value of the highest bar?" Figure 1 159 as a value line, which shows that the value of the second bar 124 is equal to 3. In various implementations, the bar graph used for training may include integers for a predetermined range of values, so that training for identifying such values ​​can be implemented using a fixed number of answer classes in conjunction with a softmax classifier model trained by the model trainer 128 to give a probability for each label / value.

[0114] In additional or alternative implementations, and in many real-world applications, the value range of a bar graph may include a continuous range of values ​​rather than a set of discrete predetermined values. In such a scenario, a model may be trained using a model trainer 128 that regresses directly on the actual values ​​of a specific bar rather than relying on a set of predetermined value answers.

[0115] Figure 7 An example of an attention-weighted feature vector for an input data visualization image is shown. Figure 7In the simplified example of , a bar chart is shown with bar 702, bar 704, and bar 706. As can be observed, the second bar 704 is the largest of the 3. Therefore, for the question "Is the second bar greater than the third bar?", the illustrated attention weighted features will result in regions 708, 710, and 712 being identified. As can be observed, the identified answer positions 708, 712 are relevant because they identify the absence of bar values ​​for bars 702, 706 at corresponding value ranges, while region 710 is relevant for identifying that bar 704 does include a value within a value range that is relevant to answering the received question.

[0116] although Figure 7 A non-limiting example of an attention map is shown, such as when possible answer locations are highlighted. However, it should be appreciated that other techniques may be used to predict answer locations within a data visualization. For example, the answer locations may be derived directly, e.g., as xy coordinates, without explicitly generating an attention map. In other examples, a set of rectangles or other appropriate shapes may be used to predict and illustrate answer locations.

[0117] Figure 8 Another example bar chart that can be used in visualizing the training data set 134 and whose questions can be answered using the DV query handler 130 is provided. Figure 8 As shown in and above about Figure 6 The referenced Figure 8 An example of includes grouping bars 802, 804, 806, and 808 such that each group of bars corresponds to a single label. For example, each bar of each group may represent one of three countries, companies, or other types of entities, and each label may identify a particular characteristic of each of the three entities. In this manner, for example, three companies may be considered relative to each other with respect to labels such as number of employees, annual gross revenue, or other characteristics.

[0118] All of the above descriptions are Figure 8 An example bar chart of Figure 8 It is further shown that many different types of data visualizations can be utilized and can be easily generated to be included in a visualization training dataset. For example, it can be observed that for Figure 8 For bar charts of this type, you can include questions such as “How many groups are there for each label?” even though such questions might be meaningless or useless with other types of data visualization.

[0119] Fig. 9 Shows Figure 1Example implementations of the source data generator 160 and associated source data query 164 and source data 166. As has been described with respect to this, Figure 1 The system 100 can be configured to use pre-constructed source data queries 164 to analyze the data visualization 901 and recover its underlying data (in Fig. 9 902). Specifically, as shown, the illustrated data visualization 901 includes a bar chart in which bars 904, 906, 908 are associated with labels 910, 912, and 914, respectively. In response to a request for source data 902, a query of source data query 164 may be applied. For example, the source data query may include a structural query (916), such as how many bars are there? Of course, as described above, such a structural question may be determined based on its irrelevance to the type of data visualization 901 being considered. For example, a similar question in the context of a pie chart may include "How many slices are included?"

[0120] Subsequently, various appropriate value and label questions may be applied. For example, having determined that three bars exist as a result of operating the structured related query of structure query 916, a pair of value and label questions may be applied to each of the identified bars. For example, value question 918 "What is the value of the first bar?" may be applied to the first bar, while label question 920 "What is the label of the first bar?" may be applied. Similar value question 922 and label question 924 may be applied to the second bar, while a final value question 926 and label question 928 may be applied to the third bar. Thus, it may be observed that source data generator 160 may be configured to utilize information determined from one or more structure queries 916 to then proceed to capture the necessary value and label content.

[0121] As a result, in Fig. 9 , the source data 902 may be constructed, including constructing a data table in which the various labels 910, 912, 914 are included in corresponding rows of the name column 930. Similarly, the values ​​of the bars 904, 906, 908 are included in appropriate rows in the same row within the value column 932. Thus, as described above, the resulting source data 902 may be used for any desired or suitable purpose associated with the formatted data table, including, for example, performing data analysis, text searching, and the generation of any specific data visualizations.

[0122] therefore, Figures 1 to 9 The systems and methods need not be used to identify and classify visual elements using output answers; for example, they need not be used to classify the top / best "K" answers within a set of predetermined answers. Instead, Figures 1 to 9The systems and methods are configured to address the problem of labels that do not have a static meaning, but rather have a local meaning that may only be valid in the context of a particular data visualization. Using the described techniques, the location of a label or other answer can be found and ultimately interpreted within the context of the entire particular data visualization being examined. These and other features are acquired even if the type of data visualization being considered contains sparse image data that cannot typically be examined using known types of natural image processing that rely on statistically meaningful patterns that occur naturally in photographs (e.g., the likelihood that the sky will contain the sun or the moon, or that an image of a car is likely to include an image of a road but is unlikely to include an image of a whale).

[0123] The implementation of the various techniques described herein may be implemented in digital electronic circuits, or in computer hardware, firmware, software, or a combination thereof. The implementation may be implemented as a computer program product, i.e., a computer program tangibly embodied in an information carrier, for example, in a machine-readable storage device, the computer program product being used to be executed by a data processing device or to control the operation of the data processing device, the data processing device being, for example, a programmable processor, a computer, or multiple computers. Computer programs such as the above-mentioned (multiple) computer programs may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. The computer program may be deployed to be executed on one computer, or to be executed on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.

[0124] The method steps may be performed by one or more programmable processors executing a computer program to perform functions by operating on input data and generating output. The method steps may also be performed by, and the apparatus may be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0125] For example, processors suitable for the execution of computer programs include general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The elements of a computer may include at least one processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer may also include one or more mass storage devices for storing data, or be operably coupled to receive data from or transfer data to one or more mass storage devices for storing data, or both, such as magnetic disks, magneto-optical disks, or optical disks. Information carriers suitable for embodying computer program instructions and data include all forms of non-volatile memory, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into a dedicated logic circuit.

[0126] To provide interaction with a user, implementations may be implemented on a computer having a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, voice, or tactile input.

[0127] The implementation may be implemented in a computing system that includes a back-end component (e.g., as a data server) or includes a middleware component (e.g., an application server) or includes a front-end component (e.g., a client computer) having a graphical user interface or a web browser through which a user can interact with the implementation, or any combination of such back-end, middleware, or front-end components. The components may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0128] Although certain features of the described implementations have been illustrated as described herein, many modifications, substitutions, changes and equivalents will now occur to those skilled in the art. It should therefore be understood that the appended claims are intended to cover all such modifications and changes that fall within the scope of the embodiments.

Claims

1. A computer program product tangibly embodied on a non-transitory computer-readable storage medium and comprising instructions that, when executed by at least one computing device, are configured to cause the at least one computing device to: Identify data visualization (DV); Generating a data visualization feature map representing the data visualization, including maintaining correspondence of spatial relationships between features of a map of the data visualization within the data visualization feature map and corresponding features of the data visualization, wherein each position of the data visualization feature map represents a data visualization feature vector, and the data visualization feature vector represents a corresponding position of the data visualization; identifying queries for answers included in the data visualization; encoding the query into a query feature vector; generating an attention map using an attention weight distribution of attention weights distributed over the data visualization feature map, the attention weights having weight values ​​set based on a relevance of the query feature vector to each attention weight and a corresponding feature vector of the data visualization feature map; generating an attention-weighted feature map comprising an array of weighted data visualization feature vectors, comprising multiplying each attention weight value of the attention weight distribution by a corresponding data visualization feature vector of the data visualization feature map, the corresponding data visualization feature vector comprising the data visualization feature vector; generating a prediction of at least one answer location within the data visualization based on the attention-weighted feature map; as well as The answer is determined from the at least one answer position.

2. The computer program product of claim 1 , wherein the instructions, when executed, are further configured to cause the at least one computing device to: The data visualization included within an image file is identified during a computer-based search of at least one document having at least one embedded image file that includes the image file. 3 . The computer program product of claim 1 , wherein the data visualization comprises at least one visual element visually arranged to represent and generated from source data and associated values ​​and labels.

4. The computer program product of claim 1 , wherein the instructions, when executed to generate the data visualization feature map, are further configured to cause the at least one computing device to: An image file containing the data visualization is input to a convolutional neural network trained using a dataset of data visualizations and associated query / answer pairs.

5. The computer program product of claim 1 , wherein the instructions, when executed to encode the query, are further configured to cause the at least one computing device to: The query is input to a long short-term memory model trained using a dataset of query / answer pairs.

6. The computer program product of claim 1 , wherein the instructions, when executed to generate the attention graph, are further configured to cause the at least one computing device to: generating a composite data visualization feature vector from said array of weighted data visualization feature vectors; concatenating the composite data visualization feature vector with the query feature vector to obtain a joint query / answer feature vector; and The answer location is determined from the joint query / answer feature vector.

7. The computer program product of claim 1 , wherein the instructions, when executed to determine the answer, are further configured to cause the at least one computing device to: Optical character recognition (OCR) is performed within the at least one answer location to provide the answer as answer text.

8. The computer program product of claim 1 , wherein the instructions, when executed to determine the answer, are further configured to cause the at least one computing device to: applying a set of source data queries to the data visualization feature map, the set of source data queries including the at least one query; and Source data from which the data visualization is created is generated based on answers obtained from the application to queries of the source data.

9. The computer program product of claim 8, wherein the instructions, when executed to determine the answer, are further configured to cause the at least one computing device to: A new data visualization is generated based on the generated source data.

10. The computer program product of claim 1, wherein the instructions, when executed to determine the answer, are further configured to cause the at least one computing device to: An overlay image of the answer is plotted within the plot of the data visualization and positioned to visually identify the answer location.

11. A computer-implemented method, the method include: Receive queries for data visualization (DV); encoding the query into a query feature vector; generating a data visualization feature map representing at least one spatial region of the data visualization, the data visualization feature map comprising an array of data visualization feature vectors corresponding to each of a plurality of positions within the at least one spatial region, wherein each position of the data visualization feature map represents a data visualization feature vector representing a corresponding position of the data visualization; generating an attention map, comprising applying a weighted distribution of attention weights to the data visualization feature vectors, wherein for each data visualization feature vector, a corresponding position in the plurality of positions is associated with a weight indicating a relative likelihood of including the answer therein; generating an attention-weighted feature map comprising an array of weighted data visualization feature vectors, comprising multiplying each attention weight value of the attention weight distribution by a corresponding data visualization feature vector of the data visualization feature map, the corresponding data visualization feature vector comprising the data visualization feature vector; generating a prediction of at least one answer location within the at least one spatial region based on the attention-weighted feature map; as well as An answer to the query is determined from the at least one answer location.

12. The method according to claim 11, include: inputting an image file containing the data visualization into a convolutional neural network trained using a dataset of data visualizations and associated query / answer pairs; as well as The query is input to a long short-term memory model trained using a dataset of query / answer pairs.

13. The method according to claim 11, include: Optical character recognition (OCR) is performed within the at least one answer location to provide the answer as answer text.

14. A computer system, include: at least one memory including instructions; as well as at least one processor operably coupled to the at least one memory and arranged and configured to execute instructions which, when executed, cause the at least one processor to: generating a visualization training dataset, the visualization training dataset comprising a plurality of training data visualizations and visualization parameters, and query / answer pairs for the plurality of training data visualizations; training a feature map generator to generate a feature map for each of the training data visualizations; training a query feature vector generator to generate a query feature vector for each of the queries in the query / answer pairs; training an answer location generator to generate an answer location within each of the training data visualizations for an answer to a corresponding query in the query / answer pair based on the trained output of the feature map generator and the trained query feature vector; Inputting a new data visualization and a new query into the trained feature map generator and the trained query feature vector to obtain a new feature map and a new query feature vector; as well as generating a new answer position within the new data visualization for the new query based on the new feature map and the new query feature vector, wherein the answer location generator comprises an attention map that assigns a plurality of attention weights to each feature map of each of the training data visualizations, each attention weight being assigned to a spatial location of a corresponding feature map and indicating a relative likelihood that each spatial location includes the answer location, Further wherein the answer position generator is configured to generate an attention-weighted feature map using the attention map and the plurality of attention weights, comprising multiplying each attention weight by a corresponding feature vector of the feature map, the attention-weighted feature map comprising an array of weighted feature vectors having weights set by the plurality of attention weights.

15. The system of claim 14, wherein the system is further configured to: The visualization training dataset is synthetically generated.

16. The system of claim 14, wherein the system is further configured to: applying a set of source data queries against the new feature map; and Source data from which the data visualization is created is generated based on answers obtained from the application to queries of the source data.

Citation Information

Patent Citations

  • Systems and methods for visual question answering

    CN106649542A