Information processing device, information processing method, and program
The information processing device evaluates data visualization candidates using user context to ensure alignment with user insights, addressing the inconsistency in existing technologies by providing relevant and accurate visualization results.
Patent Information
- Application Number
- JP2023546584
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-07
- Publication Date
- 2025-09-17
- Estimated Expiration
- 2041-09-07
AI Technical Summary
Existing data visualization technologies fail to consistently provide insights desired by users due to a lack of user context consideration in template data, leading to suboptimal visualization results.
An information processing device and method that acquires an evaluation dataset and context data to evaluate multiple insight subjects, using an evaluation model to determine the relevance between the data and user context, thereby identifying visualization candidates that align with user insights.
Enables evaluation of whether candidate visualizations provide the insights desired by users, ensuring relevance and accuracy in data visualization outcomes.
Smart Images

Figure 0007740343000003 
Figure 0007740343000004 
Figure 0007740343000005
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] In data analysis work, it is common to go through a cycle of "hypothesis setting, analysis / visualization, hypothesis verification," but this work requires a great deal of time and effort. Automatic insight discovery technology is a technology that automatically discovers visualization candidates that humans consider useful based on data characteristics. This can significantly reduce the workload in data analysis work. For example, Patent Document 1 listed below describes a method for generating instance data that visualizes data to be visualized based on template data containing keywords that express a method for visualizing the results of data analysis, and regenerating the instance data based on the evaluation value of the instance metadata. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2018 / 173251 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the visualization results of data desired by users vary depending on the content of the data, the needs of the user, etc., and are not uniformly determined. The technology described in Patent Document 1 has a problem in that if the template data does not capture the user context, the presented visualization candidates are not necessarily the visualization results desired by the user.
[0005] One aspect of the present invention has been made in consideration of the above-mentioned problems, and one of its objectives is to provide a technology that enables evaluation of whether candidate data visualizations provide the insights a user is looking for. [Means for solving the problem]
[0006] An information processing device according to one aspect of the present invention includes an acquisition means for acquiring an evaluation dataset and context data, and an evaluation means for evaluating a plurality of insight subjects generated by referring to at least the evaluation dataset in accordance with the context data.
[0007] An information processing method according to one aspect of the present invention includes at least one processor acquiring an evaluation dataset and context data, and evaluating a plurality of insight subjects generated by referring to at least the evaluation dataset according to the context data.
[0008] A program according to one aspect of the present invention causes a computer to execute a process of acquiring an evaluation dataset and context data, and a process of evaluating a plurality of insight subjects generated by referring to at least the evaluation dataset in accordance with the context data. [Effects of the Invention]
[0009] According to one aspect of the present invention, it is possible to evaluate whether candidate visualizations of data provide the insights a user is looking for. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a block diagram showing a configuration of an information processing device according to a first exemplary embodiment of the present invention. [Figure 2] 1 is a flowchart showing the flow of an information processing method according to a first exemplary embodiment of the present invention. [Figure 3]FIG. 2 is a diagram showing examples of insight subjects and evaluation results according to the first exemplary embodiment of the present invention. [Figure 4] FIG. 10 is a block diagram showing the configuration of an information processing device according to a second exemplary embodiment of the present invention. [Figure 5] FIG. 10 is a flowchart showing the flow of an information processing method according to a second exemplary embodiment of the present invention. [Figure 6] FIG. 10 is a diagram showing an example of input data according to the second exemplary embodiment of the present invention. [Figure 7] FIG. 10 is a diagram illustrating an example of context and visualization information according to the second exemplary embodiment of the present invention. [Figure 8] FIG. 10 is a diagram illustrating an example of generating a feature vector according to the second exemplary embodiment of the present invention. [Figure 9] 10A and 10B are diagrams illustrating examples of aggregated data and statistics according to the second exemplary embodiment of the present invention. [Figure 10] FIG. 10 is a diagram illustrating an example of an evaluation model according to the second exemplary embodiment of the present invention. [Figure 11] FIG. 10 is a diagram showing an example of displaying an insight subject together with an evaluation result according to the second exemplary embodiment of the present invention. [Figure 12] FIG. 10 is a diagram showing an example of displaying visualization information together with evaluation results according to the second exemplary embodiment of the present invention. [Figure 13] FIG. 10 is a diagram showing an example of displaying an insight subject together with an evaluation result according to the second exemplary embodiment of the present invention. [Figure 14] FIG. 10 is a block diagram showing the configuration of an information processing device according to a third exemplary embodiment of the present invention. [Figure 15] FIG. 10 is a diagram illustrating an example of a computer that executes instructions of a program that is software that realizes each function of the information processing device. DETAILED DESCRIPTION OF THE INVENTION
[0011] Exemplary Embodiment 1 A first exemplary embodiment of the present invention will be described in detail with reference to the drawings. This exemplary embodiment is a basic form of the exemplary embodiments described below.
[0012] <Configuration of information processing device> The configuration of an information processing device 1 according to this exemplary embodiment will be described with reference to FIG. 1. FIG. 1 is a block diagram showing the configuration of the information processing device 1. The information processing device 1 is a device that evaluates whether data visualization candidates provide insights desired by a user. As shown in the figure, the information processing device 1 includes an acquisition unit 11 and an evaluation unit 12. The acquisition unit 11 acquires an evaluation dataset and context data. The evaluation unit 12 evaluates multiple insight subjects generated by referring to at least the evaluation dataset according to the context data.
[0013] (Evaluation dataset) The evaluation dataset is data used by the information processing device 1 to evaluate visualization candidates for data. The evaluation dataset includes at least one of evaluation data, which is data to be visualized, and associated data related to the evaluation data. However, the data included in the evaluation dataset is not limited to the above-mentioned examples, and the evaluation dataset may include other information.
[0014] (Evaluation data) The evaluation data is data to be visualized, and is, for example, multidimensional data including multiple records. Examples of the evaluation data include data indicating the monthly sales record of a certain store, data indicating the size and area of the store, data indicating the product code, product name, and unit price of products sold in the store, and / or data indicating the customer's gender, age, place of residence, occupation, etc. However, the evaluation data is not limited to this and may be other data. As an example, the evaluation data is visualized as a chart (pie chart, bar graph, line graph, etc.) that represents the contents of the evaluation data.
[0015] (Related data) The related data is data related to the evaluation data. For example, the related data includes aggregated data indicating the aggregation results of the evaluation data, statistics of the aggregated data, and / or related information, which is a collection of various information used to visualize the evaluation data. For example, the related information includes some or all of the name, data type, type of aggregation method, and type of chart design of the data used to visualize the evaluation data. Note that the data included in the related data is not limited to the above-described examples, and the related data may include other data.
[0016] (context data) The context data is data that indicates what kind of insight the user is seeking. As an example, the context data includes at least one of a context, which is data related to the insight the user is seeking, and a feature vector that represents the context in a vector space. Note that the data included in the context data is not limited to the above-described examples, and the context data may include other data.
[0017] (context) The context is data related to the insights desired by the user, and is, for example, linguistic information extracted from a user query or metadata. Specifically, for example, the context is the words "product A" and "customer" extracted from a user query such as "about customers of product A." As another example, the context is the words "sales" and "transition" extracted from a user query such as "about sales trends." Furthermore, the context is, for example, the words "product A" and "customer" extracted from metadata whose "search history" is "customers of product A." Furthermore, the context is, for example, the words "sales" and "transition" extracted from metadata whose "search history" is "sales trends." However, the context is not limited to linguistic information and may be other information. The context may be, for example, location information indicating the user's location, information indicating the degree of association between words, or information indicating a site browsing history.
[0018] (Insight Subject) The insight subject is data generated with reference to at least the evaluation dataset. The insight subject, for example, includes at least one of data representing the visualization result of the evaluation data and data used to visualize the evaluation data. The visualization result of the evaluation data is, for example, a chart (pie chart, bar graph, line graph, etc.) representing the contents of the evaluation data. Furthermore, for example, the insight subject may be part of the related data described above, for example, related information included in the related data. In other words, the insight subject may be part of the evaluation dataset. However, the insight subject is not limited to the above examples and may be other data.
[0019] (Insight) In this specification, an insight refers to a visualization result that a person perceives as useful and data representing such a visualization result. In other words, an insight refers to an insight subject that a person perceives as useful.
[0020] There is no particular limitation on the method by which the acquisition unit 11 acquires the evaluation dataset and the context data. For example, the acquisition unit 11 may acquire the evaluation dataset and the context data by reading them from an external storage device or an internal storage device, or may acquire the evaluation dataset and the context data via a communication IF or an input / output IF.
[0021] Furthermore, the method by which the evaluation unit 12 evaluates multiple insight subjects according to the context data is not particularly limited. As an example, the evaluation unit 12 calculates an evaluation value for each of the multiple insight subjects, which is an evaluation result of whether the insight provides the insight desired by the user. Hereinafter, this evaluation value is also referred to as an insight score. The insight score is a great help in finding an insight subject that provides the insight desired by the user even when output as is. Furthermore, by using the insight score, it is possible to automatically detect insight subjects that have a high insight score, i.e., that are likely to provide the insight desired by the user.
[0022] As an example, the evaluation unit 12 evaluates multiple insight subjects using an evaluation model that receives input of related data and context data and outputs an evaluation value. The evaluation model may be a predefined score function or may be a trained model constructed by machine learning. When using a score function, as an example, the evaluation unit 12 evaluates multiple insight subjects using a score function that outputs a higher evaluation value the higher the relevance between the related data and the context data. However, the evaluation method used by the evaluation unit 12 is not limited to these, and other methods may be used.
[0023] The visualization results of the evaluation data vary depending on the content of the related information used for the visualization. Each of the multiple visualization results obtained by visualizing the evaluation data in multiple different patterns is hereinafter also referred to as a "visualization candidate." The visual features that the multiple visualization candidates of the evaluation data present to the user are different for each of the multiple visualization candidates.
[0024] An insight subject corresponds one-to-one with a visualization candidate in the evaluation data. Therefore, the evaluation unit 12 evaluates a plurality of insight subjects according to the context data, thereby evaluating a plurality of visualization candidates according to the context data.
[0025] <Flow of information processing method> The flow of the information processing method S1 according to this exemplary embodiment will be described with reference to Fig. 2. Fig. 2 is a flow diagram showing the flow of the information processing method S1.
[0026] In step S11, at least one processor acquires an evaluation dataset and context data. Then, in step S12, at least one processor evaluates multiple insight subjects generated by referring to at least the evaluation dataset according to the context data. This completes the information processing method S1 in FIG. 2.
[0027] Note that the processes of S11 to S12 may be executed by one processor, or the processes of S11 and S12 may be executed by different processors. In the latter case, the respective processors may be included in one information processing device, or may be included in different information processing devices. Furthermore, at least one processor that executes the processes of S11 to S12 may be included in the information processing device 1.
[0028] FIG. 3 is a diagram showing examples of insight subjects and evaluation results. In the example of FIG. 3, insight subjects V1 to V8 are data representing visualization candidates of the evaluation data. The evaluation results are the results of the evaluation unit 12 calculating the insight scores for each of the insight subjects V1 to V8. In the example of FIG. 3, the insight score for insight subject V1 is "0.2," and the insight score for insight subject V2 is "0.1." Similarly, the insight scores for insight subjects V3 to V8 are "0.8," "0.6," "0.3," "0.5," "0.9," and "0.7," respectively.
[0029] The information processing device 1 according to this exemplary embodiment is configured to include an acquisition unit 11 that acquires an evaluation dataset and context data, and an evaluation unit 12 that performs evaluations according to the context data for multiple insight subjects generated with reference to at least the evaluation dataset. Therefore, the information processing device 1 according to this exemplary embodiment has the effect of making it possible to evaluate whether visualization candidates for data provide insights desired by a user.
[0030] The functions of the information processing device 1 described above can also be realized by a program. The program according to this exemplary embodiment causes a computer to execute a process of acquiring an evaluation dataset and context data, and a process of evaluating multiple insight subjects generated by referring to at least the evaluation dataset according to the context data. Therefore, the program according to this exemplary embodiment has the effect of making it possible to evaluate whether data visualization candidates provide the insights desired by the user.
[0031] Furthermore, the information processing method S1 according to this exemplary embodiment employs a configuration in which at least one processor acquires an evaluation dataset and context data, and evaluates a plurality of insight subjects generated by referring to at least the evaluation dataset according to the context data. Therefore, the information processing method S1 according to this exemplary embodiment has the effect of enabling evaluation of visualization candidates as to whether they provide the insight desired by the user.
[0032] Exemplary Embodiment 2 A second exemplary embodiment of the present invention will be described in detail with reference to the drawings. Note that components having the same functions as those described in the first exemplary embodiment are given the same reference numerals, and their description will not be repeated.
[0033] <Configuration of information processing device> 4 is a block diagram showing the configuration of information processing device 1A. Information processing device 1A includes control unit 10A that controls all components of information processing device 1A and storage unit 17 that stores various data used by information processing device 1A. Information processing device 1A also includes communication unit 18 that enables information processing device 1A to communicate with other devices, display unit 19 that enables information processing device 1A to display and output data, and input unit 20 that accepts input to information processing device 1A. While the following describes an example in which display unit 19 displays and outputs data, information processing device 1A may also output data in the form of, for example, printout or audio output. Display unit 19 and input unit 20 may also be external devices attached to information processing device 1A.
[0034] The control unit 10A includes an acquisition unit 11, an evaluation unit 12, a first generation unit 13, and a second generation unit 14. The storage unit 17 stores an evaluation dataset DS, context data CD, evaluation model parameters EMP, evaluation results ER, and display data DD.
[0035] (Evaluation dataset DS) The evaluation dataset DS includes evaluation data and related data VD related to the evaluation data. The evaluation data is data to be visualized, and examples thereof include data showing the monthly sales record of a certain store, data showing the size and area of the store, data showing the product codes, product names and unit prices of products sold in the store, and / or data showing the gender, age, place of residence, occupation, etc. of customers.
[0036] (Related Data VD) The related data VD is data related to the evaluation data. Related information related to the evaluation data V Feature vector d representing related information V in vector space V Aggregated data s obtained by aggregating data included in the evaluation data and corresponding to the related information V V , and Aggregated dataV Statistics of t V At least one of the following is included.
[0037] (Related Information V) The related information V is, for example, a collection of various pieces of information used to visualize the evaluation data, and includes, for example, the following information: Attribute information for each piece of data included in the evaluation data - Information about the aggregation method (filter, aggregation function, column name used as the aggregation key, etc.) (information about the filter to be applied to the evaluation data, etc.) Information about the chart design (x-axis, y-axis, chart type, plot type, etc.) (information about the relationship between each axis of the chart and items, etc.)
[0038] (Feature vector d V ) Related information feature vector d V is a vector space representation of the related information V. Any vectorization method can be used, for example, distributed representation of words.
[0039] (Aggregated data V ) Aggregated Data V is the data obtained by aggregating the numerical values corresponding to the related information V from the evaluation data. V are plotted on a chart as a visualization result of the relevant information V.
[0040] (statistics t V ) Aggregated Data V Statistics of t V is the aggregate data V The statistics used are arbitrary, but for example, the following are statistics t V It is available as. Maximum, minimum, median · Mean, standard deviation, variance Cardinality Percentage of zero values, percent of missing values Kurtosis, skewness Entropy Gini coefficient
[0041] (Context Data CD) The context data CD includes: Context C, and Feature vector d representing the context in vector space C At least one of the following is included.
[0042] (Context C) Context C is data related to the insights desired by the user. As an example, context C is data expressing the insights desired by the user in natural language, and includes data related to the quality and quantity of the insights desired by the user. Context C may be extracted from a user query Q and / or metadata M, which will be described later. As an example, context C includes the words "product A" and "customer."
[0043] (Feature vector d C ) Feature vector d of context C C is a vector space representation of the context C. Any vectorization method can be used, but as an example, distributed representation of words may be used.
[0044] (User query Q) The user query Q is a query about the insight the user desires, and is given by the user in natural language. The user query Q includes, for example, the following information: Information about the data to be analyzed (e.g., "Product A," "Sales") Hypothesis for insight (e.g., "~ is increasing" or "~ is prominent") -Characteristics of the chart you are planning to use (e.g., regional aggregation, pie chart)
[0045] (Metadata M) The metadata M is information from which the insight desired by the user can be estimated. As an example, the metadata M is automatically collected by a predetermined system. The metadata M includes, for example, the following information: User search history (e.g., searching for "Product A, Customer") User analysis history (e.g., customer analysis of product A conducted in the past) User rating history (e.g., a chart about a customer of product A was highly rated) User behavior history (e.g., spent xx minutes on the website or store for product A)
[0046] (Evaluation model parameter EMP) The evaluation model parameter EMP is a parameter that defines the evaluation model f. The evaluation model f is a model that receives related data VD and context data CD as input and quantitatively evaluates the insight subject corresponding to the input related data VD. Any model that can be used to estimate the evaluation result of the insight subject can be used as the evaluation model f. For example, a rule-based model as described below, or a model constructed by machine learning, can be used as the evaluation model f. The output of the evaluation model f is, for example, a score or label probability that represents the evaluation result. The evaluation model f will be described later.
[0047] (Evaluation result ER) The evaluation result ER is data indicating the evaluation result of the insight subject by the evaluation unit 12. As an example, the evaluation result ER is an insight score y^ representing the evaluation result for each of the multiple insight subjects.
[0048] (Insight score y^) The insight score y^ is a quantitative index of the quality of visualization calculated based on the output value of the evaluation model f. The insight score y^ may be, for example, the output value of the evaluation model f, or may be a value obtained by applying processing such as normalization and / or weighting to the output value of the evaluation model f. A specific example of a method for calculating the insight score y^ will be described later.
[0049] (Display data DD) The display data DD is data for presenting to the user the evaluation result of the insight subject by the information processing device 1A, that is, data related to the evaluation result of the insight subject as to whether the insight desired by the user is provided.
[0050] (Acquisition part 11) The acquisition unit 11 acquires the evaluation dataset DS and the context data CD. As an example, the acquisition unit 11 acquires the evaluation dataset DS and the context data CD by reading them out from the storage unit 17. However, the method of acquiring the evaluation dataset DS and the context data CD is not particularly limited. For example, the acquisition unit 11 may acquire the evaluation dataset DS and the context data CD input by a user of the information processing device 1A via the input unit 20. Furthermore, for example, the acquisition unit 11 may acquire the evaluation dataset DS and the context data CD from an external device by communication via the communication unit 18.
[0051] (Evaluation Section 12) The evaluation unit 12 performs evaluation on the multiple insight subjects generated by at least referring to the evaluation dataset DS according to the context data CD. As an example, the evaluation unit 12 calculates an insight score y^ for each of the multiple insight subjects, generates an evaluation result ER indicating the calculation result, and stores it in the storage unit 17.
[0052] (First generation unit 13 and second generation unit 14) The first generation unit 13 generates a plurality of insight subjects by referring to the evaluation dataset DS. The first generation unit 13 also generates display data DD related to the evaluation results of the evaluation unit 12. The second generation unit 14 generates at least a portion of the context data CD and at least a portion of the related data VD.
[0053] <Flow of information processing method> The flow of the information processing method according to this exemplary embodiment will be described with reference to the drawings. Fig. 5 is a flow diagram showing the flow of the information processing method. In the following, a case will be described in which the related information V is visualization information used to visualize evaluation data. In the following, visualization information that is an example of the related information V will also be referred to as "visualization information V."
[0054] (Step S101) In step S101, the acquisition unit 11 acquires input data D and context generation data. The input data D is an example of evaluation data according to the present specification. The input data D may include data to be plotted on a chart, and any format may be used for the input data D. For example, the acquisition unit 11 acquires the input data D via the input unit 20 or the communication unit 18.
[0055] FIG. 6 is a diagram showing an example of input data D. In the example of FIG. 6, input data D includes sales data, store data, product data, and customer data. The sales data, store data, product data, and customer data are all multidimensional data sets including multiple records. The sales data is multidimensional data including the data items of "date," "product code," "customer code," "store code," and "sales." The store data is multidimensional data including the data items of "store code," "store name," "area," and "size." The product data is multidimensional data including the data items of "product code," "product name," "category," and "unit price." The customer data is multidimensional data including the data items of "customer code," "age," "gender," "place of residence," "occupation," and "income."
[0056] (Data for generating context) The context generation data is data for generating a context C, and includes, for example, one or both of a user query Q and metadata M. The context generation data may include multiple user queries, and may also include multiple metadata. However, the context generation data is not limited to user queries and metadata, and may be other data. Furthermore, the context generation data may be data that can be used as context C as is. For example, the acquisition unit 11 may acquire the context generation data via the input unit 20 or the communication unit 18, or may acquire the context generation data by reading it from the storage unit 17.
[0057] (Step S102) In step S102, the second generating unit 14 generates the evaluation data set DS and the context data CD. Specific examples of generating the evaluation data set DS and the context data CD will be described below.
[0058] (Generation of evaluation dataset DS) The second generation unit 14 first acquires visualization information V. The second generation unit 14 may acquire the visualization information V by reading it from a predetermined storage area of the storage unit 17, or may acquire the visualization information V via the input unit 20 or the communication unit 18. In this case, the second generation unit 14 acquires multiple pieces of visualization information V. The visualization information V includes, for example, attribute information of each piece of data included in the input data D, information regarding the relationship between each axis of the chart and each item, a filter to be applied to the input data D, a chart type, an aggregation method, and the like.
[0059] The second generation unit 14 generates a feature vector d that expresses the acquired visualization information V in a vector space using an arbitrary language model. V Generate a feature vector d V is generated for each of the plurality of pieces of visualization information V. The second generation unit 14 also generates aggregate data s V , and aggregated data sV Statistics t, which is a collection of various statistics about V Generate.
[0060] The second generation unit 14 generates the obtained visualization information V and the generated feature vector d V , aggregated data V , statistic t V The evaluation data set DS is generated by including the associated data VD containing the plurality of visualization information V and the plurality of feature vectors d V and a pair of visualization information V and feature vector d V may be included.
[0061] (Generating context data CD) Furthermore, the second generation unit 14 performs any natural language processing on the data for context generation acquired by the acquisition unit 11 in step S101 to generate a context C. Note that the second generation unit 14 may use the data for context generation as the context C as it is.
[0062] As one example, the second generation unit 14 performs natural language processing on a user query "about customers of product A" to generate a context C of "product A" and "customer." As another example, the second generation unit 14 performs natural language processing on a user query "about sales trends" to generate a context C of "sales" and "trend." As another example, the second generation unit 14 performs natural language processing on metadata whose "search history" is "customers of product A" to generate a context C of "product A" and "customer." As another example, the second generation unit 14 performs natural language processing on metadata whose "search history" is "sales trends" to generate a context C of "sales" and "trend."
[0063] The second generation unit 14 generates a feature vector d that expresses the generated context C in a vector space using an arbitrary language model. C Generate a feature vector dC and the context C.
[0064] FIG. 7 is a diagram showing an example of the context C and the visualization information V. FIG. 8 is a diagram showing an example of the feature vector d C and feature vector d V 7 shows an example of generating a feature vector d from the visualization information V. In the example of FIG. 7, the context C includes the words "product A" and "customer." The visualization information V includes attribute information of each data included in the input data D, information on the relationship between each axis of the chart and the item, a filter to be applied to the input data D, a chart type, a counting method, and other information. Also, as shown in FIG. 8, a feature vector d V is generated, and the feature vector d is generated from the context C. C is generated.
[0065] FIG. 9 shows the aggregate data s generated by the second generation unit 14. V and statistics t V In the example of FIG. 9, the aggregated data s V is data obtained by aggregating data included in the input data D and corresponding to the visualization information V. V is the aggregate data V This is data that represents the statistics of
[0066] (Step S103) 5, the first generation unit 13 generates a plurality of insight subjects by referring to the evaluation dataset DS. When the insight subject is data indicating a visualization candidate, the first generation unit 13 generates a plurality of insight subjects by referring to the evaluation data and the related data VD, for example. In this case, the first generation unit 13 generates, for example, the aggregated data included in the related data VD. s VThe first generation unit 13 generates an insight subject representing a visualization result obtained by plotting the above data on a chart in a display format represented by the visualization information V. At this time, the first generation unit 13 generates an insight subject for each of the multiple pieces of visualization information V, thereby generating multiple insight subjects. Furthermore, since one insight subject is generated for one piece of visualization information V, there is a one-to-one correspondence between the visualization information V and the insight subject. Note that the insight subject is not limited to data representing a visualization candidate; for example, the visualization information V may be treated as an insight subject as it is.
[0067] (Step S104) In step S104, the evaluation unit 12 evaluates each of the multiple insight subjects with reference to the context data CD. At this time, the evaluation unit 12 gives a higher evaluation to an insight subject that has a higher relevance to the context data CD, for example.
[0068] More specifically, the evaluation unit 12 performs evaluation for each of the multiple insight subjects by referring to the related data VD and the context data CD. At this time, since the multiple insight subjects correspond one-to-one to the related information V, the evaluation unit 12 performs evaluation for each piece of visualization information V. In other words, the evaluation unit 12 performs evaluation for each piece of related information V included in the related data VD for each of the multiple insight subjects.
[0069] As specific examples of evaluation performed by the evaluation unit 12, rule-based evaluation and learning-based evaluation will be described.
[0070] (Rule-based evaluation) In the rule-based case, the evaluation unit 12 calculates the score y0^ using the related data VD, and then calculates the insight score y^ using the score y0^. In this case, the evaluation unit 12 may use the score y0^ as the insight score y^ as is, or may calculate the insight score y^ by applying processing such as normalization or weighting to the score y0^.
[0071] The method for calculating the score y0^ is not limited, but the evaluation unit 12 may, for example, use a rule-based score function defined for each type of insight, or may calculate the score y0^ using a model that learns the features of the chart that provides the insight.
[0072] When a score function is used, the score function is, for example, a function that outputs a higher evaluation value the stronger the relevance between the associated data VD and the context data CD. In other words, the evaluation unit 12 performs evaluations on multiple insight subjects using a score function that is a predefined score function that outputs a higher evaluation value the stronger the relevance between the associated data VD and the context data CD.
[0073] (Example 1 of rule-based evaluation) For example, the evaluation unit 12 sets the insight score y^ for related data VD that has low relevance to the context data CD to zero or a negative value, thereby lowering the evaluation result. There is no limitation on the method for calculating the degree of relevance (similarity) between the context data CD and the related data VD, but the evaluation unit 12 uses, for example, set similarity (Jaccard, Dice, Simpson, etc.), string similarity (Hamming distance, Levenshtein distance, Jaro-Winkler distance, etc.), or similarity of distributed representations (word2vec, fastText, BERT, etc.).
[0074] (Example 2 of rule-based evaluation) Furthermore, the evaluation unit 12 may calculate the insight score y^ using a score weighted by the similarity between the context data CD and the related data VD. More specifically, for example, the score y^ calculated using the related data VD and the similarity sim(CD, VD ) may be used as the insight score y^.
[0075] (Learning-based assessment) In the learning-based case, the evaluation unit 12 performs evaluations on multiple insight subjects using an evaluation model f, which is a pre-trained evaluation model that receives related data VD and context data CD and outputs evaluation values. The machine learning method of the evaluation model f is not limited, and, for example, a decision tree-based, linear regression, or neural network method may be used, or one or more of these methods may be used. Examples of decision tree-based methods include LightGBM (Light Gradient Boosting Machine) and XGBoost. Examples of linear regression methods include support vector regression, Ridge regression, Lasso regression, and ElasticNet. Examples of neural networks include deep learning.
[0076] Any training data that is considered to have insights can be used in training the evaluation model f. For example, charts created by data analysts in the past may be considered to contain features that provide insights, and the visualization information V of such charts may be used as positive samples for training. Also, visualization information V of charts that are considered to have no insights may be used as negative samples for training.
[0077] 10 is a diagram showing an example of the evaluation model f. In the example of FIG. 10, the input of the evaluation model f is a feature vector d V , feature vector d C , aggregated data s V , and the statistic t V The output of the evaluation model f is the evaluation result, for example, a label probability indicating whether the model provides the insight desired by the user.
[0078] (Example 1 of a learning-based evaluation model) When a teacher label y for an insight in visualization information V is given, an evaluation model can be trained as a classification model. For example, when y∈{0,1} is given as a label indicating that there is insight when it is 1 and that there is no insight when it is 0, a machine learning model can be trained to minimize the loss function E(θ) given by the following equation (1) as a two-class classification task. In equation (1), N is the number of training data.
number
[0079] The output of the machine learning model that minimizes the loss function above is p(y=1|VD i ,CD i ), that is, the probability that an insight is determined to exist, and this can be used as the insight score y^.
[0080] (Example 2 of a learning-based evaluation model) When a score or ranking representing the quality of visualization for each piece of visualization information V is provided as training data, an evaluation model can be trained as a regression model. For example, if y is the score provided by the training data, a machine learning model can be trained to minimize the loss function E(θ) given by the following equation (2). In equation (2), N is the number of training data.
number
[0081] The output of the machine learning model that minimizes the above loss function is a score that represents the quality of the visualization, similar to the score of the training data, and this may be used as the insight score y^.
[0082] (Step S105) 5, the evaluation unit 12 outputs information related to the insight subject to the display unit 19, and the display unit 19 displays the information related to the insight subject. Specifically, for example, the display unit 19 displays at least one of the multiple insight subjects together with the evaluation result by the evaluation unit 12 or in a display mode according to the evaluation result by the evaluation unit 12. The display mode according to the evaluation result includes, for example, the display order or the display size.
[0083] Display examples of evaluation results will be described with reference to FIGS. 11 to 13. FIG. 11 is a diagram showing an example of displaying insight subjects together with evaluation results. In the example of FIG. 11, insight subjects V7, V3, V8, ... are charts showing visualization results of input data D, and the visual features of the insight subjects V7, V3, V8, ... are different from one another. The insight score y^ of each insight subject is displayed adjacent to each of the insight subjects V7, V3, V8, .... Furthermore, multiple insight subjects V7, V3, V8, ... are displayed in descending order of the insight score y^.
[0084] According to the example of FIG. 11, multiple insight subjects are displayed in descending order of insight score y^, making it easy for the user to understand which insight subject has a high evaluation.
[0085] Fig. 12 is a diagram showing an example of displaying visualized information V together with the evaluation results. In the example of Fig. 12, the display unit 19 displays each piece of related information V included in the related data in association with the evaluation by the evaluation unit 12. Specifically, the display unit 19 displays visualized information V11 to V18 in association with the insight scores y^ corresponding to each piece of visualized information V11 to V18.
[0086] Fig. 13 is a diagram showing an example of displaying insight subjects together with evaluation results. In the example of Fig. 13, the display unit 19 displays a chart (bar graph) that is a visualization result of the input data D, and also displays the insight score y^ corresponding to the displayed chart together with the chart.
[0087] As described above, the information processing device 1A according to this exemplary embodiment is configured such that the evaluation unit 12 gives a higher evaluation to an insight subject that has a higher relevance to context data. Therefore, according to the information processing device 1A according to this exemplary embodiment, in addition to the effects achieved by the information processing device 1 according to exemplary embodiment 1, an effect of being able to perform an evaluation that makes it easy to grasp the degree of relevance between context data and an insight subject can be obtained.
[0088] Exemplary Embodiment 3 A third exemplary embodiment of the present invention will be described in detail with reference to the drawings. Note that components having the same functions as those described in the first exemplary embodiment are denoted by the same reference numerals, and their description will not be repeated.
[0089] Fig. 14 is a block diagram showing the configuration of an information processing device 1B according to this exemplary embodiment. As shown in Fig. 14, the information processing device 1B includes a control unit 10B instead of the control unit 10A of the information processing device 1A according to exemplary embodiment 2. The control unit 10B includes a learning unit 15 in addition to an acquisition unit 11, an evaluation unit 12, a first generation unit 13, and a second generation unit 14.
[0090] In this exemplary embodiment, the input unit 20 receives feedback from the user regarding the evaluation result of the evaluation unit 12. Furthermore, the learning unit 15 re-learns the evaluation model f by referring to the feedback from the user.
[0091] For example, the learning unit 15 records, as feedback from the user, in the storage unit 17 or the like, the user's operation history regarding the information related to the insight subject (the insight score y^, the visualized information V, the chart, etc.) displayed by the display unit 19. The user's operation history includes, for example, the display time of the information related to the insight subject, pressing of the rating button for the information related to the insight subject, etc.
[0092] The learning unit 15 re-learns the evaluation model f by reflecting feedback from users. For example, the learning unit 15 re-learns the evaluation model f by using visualization information V with high evaluations as positive samples and visualization information with low evaluations as negative samples.
[0093] The information processing device 1B according to this exemplary embodiment is configured such that the input unit 20 receives feedback from the user regarding the evaluation result, and the learning unit 15 re-learns the evaluation model by referring to the feedback from the user. Therefore, the information processing device 1B according to this exemplary embodiment has the effect of further improving the evaluation accuracy of the evaluation model in addition to the effect achieved by the information processing device 1 according to the exemplary embodiment 1.
[0094] [Modification] In the above-described exemplary embodiment 1, the processing performed by one information processing device 1 may be shared among multiple information processing devices. In other words, part of the processing performed by the information processing device 1 may be executed by at least one other information processing device. In other words, when each of the above-described processes is performed by at least one processor, the at least one processor may be included in one information processing device 1, or may be included in different information processing devices. This also applies to the information processing device 1A in the above-described exemplary embodiment 2 and the information processing device 1B in the above-described exemplary embodiment 3.
[0095] [Software implementation example] Some or all of the functions of the information processing devices 1, 1A, and 1B may be realized by hardware such as an integrated circuit (IC chip), or by software.
[0096] In the latter case, the information processing devices 1, 1A, and 1B are realized, for example, by a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in FIG. 15. The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program P for operating the computer C as the information processing devices 1, 1A, and 1B. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing each function of the information processing devices 1, 1A, and 1B.
[0097] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.
[0098] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, mouse, display, and printer.
[0099] Furthermore, the program P can be recorded on a non-transitory tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.
[0100] [Appendix 1] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means disclosed in the above-described embodiments are also included in the technical scope of the present invention.
[0101] [Appendix 2] Some or all of the above-described embodiments can also be described as follows: However, the present invention is not limited to the following described aspects.
[0102] (Appendix 1) an acquisition means for acquiring an evaluation dataset and context data; evaluation means for performing evaluation according to the context data on a plurality of insight subjects generated by referring to at least the evaluation dataset; An information processing device comprising:
[0103] According to the above configuration, it is possible to evaluate whether a candidate for visualization of data provides the insight desired by the user.
[0104] (Appendix 2) The evaluation means 2. The information processing device of claim 1, wherein a higher rating is given to an insight subject that is more relevant to the context data.
[0105] According to the above configuration, it is possible to perform an evaluation that makes it easy to grasp the degree of relevance between the context data and the insight subject.
[0106] (Appendix 3) further comprising a first generation means for generating the plurality of insight subjects by referring to the evaluation dataset; 3. The information processing device according to claim 1, wherein the evaluation means performs an evaluation on each of the plurality of insight subjects by referring to the context data.
[0107] According to the above configuration, it is possible to evaluate whether each of the multiple insight subjects generated by referring to the evaluation dataset provides the insight desired by the user.
[0108] (Appendix 4) the evaluation data set includes evaluation data and associated data related to the evaluation data, the first generation means generates the plurality of insight subjects by referring to the evaluation data and the related data; The information processing device according to claim 3, wherein the evaluation means performs an evaluation for each of the plurality of insight subjects by referring to the related data and the context data.
[0109] According to the above configuration, each of the multiple insight subjects generated by referring to the evaluation dataset and related data related to the evaluation dataset can be evaluated to determine whether it provides the insight desired by the user.
[0110] (Appendix 5) 5. The information processing device according to claim 4, wherein the evaluation means performs an evaluation for each of the plurality of insight subjects on a piece of related information included in the related data.
[0111] According to the above configuration, it is possible to evaluate the insight subject for each piece of related information.
[0112] (Appendix 6) The information processing device according to claim 4 or 5, further comprising a second generating means for generating at least a portion of the context data and at least a portion of the associated data.
[0113] According to the above configuration, it is possible to evaluate whether each of a plurality of insight subjects generated by referring to the evaluation data set and related data provides the insight desired by the user.
[0114] (Appendix 7) The context data includes: Context, and Context feature vector 7. The information processing device according to any one of appendices 4 to 6, comprising at least one of the following:
[0115] According to the above configuration, it is possible to evaluate whether each of a plurality of insight subjects generated by referring to the evaluation data set and related data provides the insight desired by the user.
[0116] (Appendix 8) The related data includes: Related information related to the evaluation data; a feature vector of the related information; aggregated data obtained by aggregating data included in the evaluation data and corresponding to the related information; and Statistics of the aggregated data 8. The information processing device according to any one of appendices 4 to 7, comprising at least one of the following:
[0117] According to the above configuration, it is possible to evaluate whether each of a plurality of insight subjects generated by referring to the evaluation data set and related data provides the insight desired by the user.
[0118] (Appendix 9) The evaluation means 9. The information processing device according to any one of appendices 4 to 8, wherein the plurality of insight subjects are evaluated using a predefined score function that outputs a higher evaluation value the higher the relevance between the associated data and the context data.
[0119] According to the above configuration, it is possible to perform evaluation using a score function for each of a plurality of insight subjects generated by referring to the evaluation dataset and related data.
[0120] (Appendix 10) The evaluation means 9. The information processing device according to any one of appendices 4 to 8, wherein an evaluation is performed on the plurality of insight subjects using a pre-trained evaluation model that receives the related data and the context data and outputs an evaluation value.
[0121] According to the above configuration, it is possible to perform evaluation using an evaluation model for each of a plurality of insight subjects generated by referring to the evaluation dataset and related data.
[0122] (Appendix 11) further comprising a receiving means for receiving feedback from a user on the evaluation result of the evaluation means, 11. The information processing device according to claim 10, wherein the evaluation means re-learns the evaluation model by referring to feedback from the user.
[0123] According to the above configuration, it is possible to further improve the evaluation accuracy of the evaluation model that evaluates the insight subject.
[0124] (Appendix 12) 12. The information processing device according to any one of appendices 4 to 11, further comprising a display means for displaying information related to the insight subject.
[0125] According to the above configuration, the user can understand the evaluation of the insight subject from the information displayed by the display means.
[0126] (Appendix 13) The display means 13. The information processing device according to claim 12, wherein at least one of the plurality of insight subjects is displayed together with the evaluation result by the evaluation means or in a display mode according to the evaluation result by the evaluation means.
[0127] According to the above configuration, the insight subject displayed by the display means makes it easier for the user to understand the evaluation of the insight subject.
[0128] (Appendix 14) The display means 13. The information processing device according to claim 12, wherein each piece of related information included in the related data is displayed in association with the evaluation by the evaluation means.
[0129] According to the above configuration, the user can understand the evaluation of each of the multiple insight subjects from the information displayed by the display means.
[0130] (Appendix 15) At least one processor Obtaining an evaluation dataset and context data; and performing an evaluation according to the context data on a plurality of insight subjects generated by referring to at least the evaluation dataset; An information processing method including:
[0131] (Appendix 16) On the computer, A process of obtaining an evaluation dataset and context data; A process of evaluating a plurality of insight subjects generated by referring to at least the evaluation dataset according to the context data; A program that executes the following.
[0132] [Appendix 3] Some or all of the above-described embodiments can also be expressed as follows.
[0133] An information processing device comprising at least one processor, the processor executing an acquisition process for acquiring an evaluation dataset and context data, and an evaluation process for evaluating a plurality of insight subjects generated by referring to at least the evaluation dataset according to the context data.
[0134] The information processing device may further include a memory that stores a program for causing the processor to execute the acquisition process and the evaluation process. The program may be recorded on a computer-readable, non-transitory, tangible recording medium. [Explanation of symbols]
[0135] 1, 1A, 1B Information processing equipment 10A, 10B Control section 11 Acquisition unit (acquisition means) 12 Evaluation section (evaluation means) 13 First generation unit (first generation means) 14 Second generation unit (second generation means) 15 Learning section (assessment means) 17 Memory section 18 Communications Department 19 Display section 20 Input unit (reception means)
Claims
1. An evaluation dataset including evaluation data to be visualized and related data related to the evaluation data; and an acquisition means for acquiring context data representing an insight, which is a visualization result desired by a user; evaluation means for evaluating, by using a score function or an evaluation model, each of a plurality of insight subjects generated by referring to at least the evaluation data and the related data, in accordance with the context data; the score function outputs an evaluation value indicating a degree of relevance between the related data corresponding to each of the plurality of insight subjects and the context data; the evaluation model receives the associated data and the context data corresponding to each of the plurality of insight subjects, and outputs an evaluation value indicating a probability that each of the plurality of insight subjects is an insight desired by the user; Information processing device.
2. The information processing apparatus according to claim 1 , further comprising a first generating means for generating the plurality of insight subjects by referring to the evaluation data and the related data.
3. The information processing apparatus according to claim 1 , wherein the evaluation means performs evaluation for each of the plurality of insight subjects for each piece of related information included in the related data.
4. 4. The information processing device according to claim 1, further comprising a second generating means for generating at least a part of the context data and at least a part of the associated data.
5. The context data includes: Context, and Context feature vector 5. The information processing device according to claim 1, wherein at least one of the above is included.
6. The related data includes: Related information related to the evaluation data; a feature vector of the related information; aggregated data obtained by aggregating data included in the evaluation data and corresponding to the related information; and Statistics of the aggregated data 6. The information processing device according to claim 1, wherein at least one of the above is included.
7. At least one processor Obtaining an evaluation dataset including evaluation data of a visualization target and associated data related to the evaluation data, and context data representing an insight, which is a visualization result desired by a user; and Using a score function or an evaluation model, an evaluation is performed according to the context data for each of a plurality of insight subjects generated by referring to at least the evaluation data and the related data. Including, the score function outputs an evaluation value indicating a degree of relevance between the related data corresponding to each of the plurality of insight subjects and the context data; the evaluation model receives the associated data and the context data corresponding to each of the plurality of insight subjects, and outputs an evaluation value indicating a probability that each of the plurality of insight subjects is an insight desired by the user; Information processing methods.
8. On the computer, A process of acquiring an evaluation dataset including evaluation data of a visualization target and related data related to the evaluation data, and context data representing an insight, which is a visualization result desired by a user; a process of evaluating each of a plurality of insight subjects generated by referring to at least the evaluation data and the related data, according to the context data, using a score function or an evaluation model; Execute the score function outputs an evaluation value indicating a degree of relevance between the related data corresponding to each of the plurality of insight subjects and the context data; the evaluation model receives the associated data and the context data corresponding to each of the plurality of insight subjects, and outputs an evaluation value indicating a probability that each of the plurality of insight subjects is an insight desired by the user; program.
Citation Information
Patent Citations
Vehicle inspection and diagnosis system
JP2004322862A
Analysis support device, analysis support method and analysis support program
JP2016153981A
Sales support server, sales support terminal, and sales support system
JP2016224873A
Methods and systems for creating new presentations using existing presentations
US20180095945A1
Customer service appraisal device, customer service appraisal system, and customer service appraisal method
WO2015194115A1