Method and device for assisting users in exploring data sets and data tables

By outputting the first data analysis information and recommending the second data analysis information based on the machine learning model, it assists users in exploring data sets and data tables, solves the problem of high threshold for using existing tools, and realizes a low-threshold and efficient data analysis process.

CN110874644BActive Publication Date: 2025-09-12THE FOURTH PARADIGM BEIJING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201911104860.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-11-13
Publication Date
2025-09-12
Estimated Expiration
2039-11-13

AI Technical Summary

Technical Problem

Existing data analysis tools have a high threshold for use, requiring users to spend a lot of time learning and require extensive business experience. They are difficult to meet the needs of business personnel who have data analysis demands but lack data analysis capabilities.

Method used

By responding to the user's field selection operation, the first data analysis information is output, and the second data analysis information is recommended based on the machine learning model, statistical correlation or business rules, to assist users in exploring data sets and data tables, use machine learning models to automatically train and predict data analysis information, and provide chart display and explanatory information.

Benefits of technology

It lowers the threshold for using data analysis, helps users quickly understand and explore data sets and data tables, and improves the efficiency and accuracy of data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110874644B_ABST
    Figure CN110874644B_ABST
Patent Text Reader

Abstract

Disclosed are a method and apparatus for assisting users in exploring data sets and data tables. The data set includes multiple pieces of data, each piece of data including the values ​​of one or more fields. In response to a user's field selection operation, first data analysis information is output, and the first data analysis information is used to characterize the data statistics of the field values ​​corresponding to the field selected by the user; and one or more second data analysis information are recommended to the user, and the second data analysis information is data analysis information predicted based on the field selected by the user. Thus, by outputting the first data analysis information, the user can understand the data statistics of the selected field. The second data analysis information is data analysis information that the user may be interested in, which is predicted based on the field selected by the user. By recommending the second data analysis information to the user, the user's data exploration cost can be reduced, allowing the user to explore data conveniently and quickly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the field of data exploration, and more particularly, to a method and apparatus for assisting users in exploring data sets and data tables. Background Art

[0002] With the advancement of modern science and technology, the rapid development and application of information technology has led to a comprehensive increase in the level of informatization across all industries. Data is growing at an unprecedented rate across society, characterized by large quantities, diverse types, and rapid updates. It has gradually become a key production factor in all industries. This abundant data contains a wealth of valuable information, but users need statistical analysis to extract meaningful results from this data.

[0003] Agile BI (Business Intelligence) tools currently available on the market can provide data analysis services, but these tools have a high barrier to entry and require a significant learning curve and extensive business experience. For business professionals who desire data analysis but lack the necessary analytical skills, using existing tools can be a significant challenge.

[0004] Therefore, a solution is needed that can facilitate users to explore data. Summary of the Invention

[0005] The exemplary embodiments of the present invention are intended to overcome the drawback of the high threshold for using data analysis tools in the prior art.

[0006] According to a first aspect of the present invention, a method for assisting users in exploring a data set is proposed. The data set includes multiple data items, each of which includes values ​​of one or more fields. The method includes: outputting first data analysis information in response to a user's field selection operation, the first data analysis information being used to characterize the data statistics of the field values ​​corresponding to the field selected by the user; and recommending one or more second data analysis information to the user, the second data analysis information being data analysis information predicted based on the field selected by the user.

[0007] Optionally, the second data analysis information is used to characterize the data statistics of the field values ​​corresponding to other fields obtained through prediction, and / or the second data analysis information is used to characterize the data statistics of the field value combination corresponding to the field combination composed of the field selected by the user and other fields obtained through prediction.

[0008] Optionally, the step of recommending one or more second data analysis information to the user includes: predicting data analysis information suitable for recommendation to the user based on a machine learning model; and / or predicting data analysis information suitable for recommendation to the user based on statistical correlation; and / or predicting data analysis information suitable for recommendation to the user based on business rules.

[0009] Optionally, the step of predicting data analysis information suitable for recommendation to the user based on the machine learning model includes: based on the relevant information of the field currently selected by the user, using the machine learning model to predict the probability that the user will select each other field later, wherein the machine learning model is trained in the following manner: taking the relevant information of the previously selected field and the relevant information of other fields as input, and taking the probability of other fields being selected as output; analyzing the data statistics of the field values ​​corresponding to the fields whose probability values ​​are greater than the first predetermined threshold to obtain second data analysis information, or analyzing the data statistics of the field values ​​corresponding to the fields whose probability values ​​are greater than the first predetermined threshold on the basis of the first data analysis information of the field selected by the user to obtain second data analysis information.

[0010] Optionally, the method also includes: obtaining user acceptance feedback on one or more second data analysis information recommended to the user, and updating the machine learning model based on the obtained acceptance feedback.

[0011] Optionally, the step of predicting data analysis information suitable for recommendation to the user based on statistical correlation includes: according to the statistical correlation between different fields, obtaining fields whose statistical correlation with the field selected by the user is higher than a second predetermined threshold; analyzing the data statistics of the field values ​​corresponding to the fields whose statistical correlation is higher than the second predetermined threshold to obtain second data analysis information, or analyzing the data statistics of the field values ​​corresponding to the fields whose statistical correlation is higher than the second predetermined threshold based on the first data analysis information of the field selected by the user to obtain second data analysis information.

[0012] Optionally, the first data analysis information is presented in the first interface area, and the second data analysis information is presented in the second interface area. The method further includes: presenting the second data analysis information in the second interface area in the first interface area in response to a user operation.

[0013] Optionally, the method also includes: in response to a user's connection operation on two data analysis information in the first interface area, connecting one of the two data analysis information to another data analysis information; and updating the other data analysis information to represent the data statistics of the field value represented by one data analysis information under the field dimension corresponding to the other data analysis information.

[0014] Optionally, the method also includes: in response to a user's connection operation on two data analysis information in the first interface area, connecting one of the two data analysis information to another data analysis information; in response to a user's selection operation on a field value represented by one of the two data analysis information in a connected state, the data statistics of the selected field value in the data analysis information are highlighted relative to the data statistics of the unselected field values, and / or the other data analysis information in the two data analysis information in a connected state is updated to be used to represent the data statistics of the selected field value under the field dimension corresponding to the other data analysis information.

[0015] Optionally, the method also includes: displaying an interface for exploring a data set, wherein the left area of ​​the interface is used to display name icons corresponding to each field of the data set, the middle area of ​​the interface is used to display a chart of the first data analysis information, and the right area of ​​the interface is used to display a chart of the second data analysis information. When importing a data set, the name icons corresponding to each field included in each data in the data set are displayed in the left area of ​​the interface, wherein, in response to the user's field selection operation, outputting the first data analysis information includes: in response to the user selecting the name icon of a specific field with the cursor, dragging it to the middle area of ​​the interface and releasing it, displaying a chart of the first data analysis information in the middle area of ​​the interface, and recommending one or more second data analysis information to the user includes: displaying a chart of the first data analysis information in the middle area of ​​the interface while displaying one or more charts of the second data analysis information in the right area of ​​the interface.

[0016] Optionally, the method also includes: when importing a data set, pre-calculating data statistics of field values ​​corresponding to each field, and caching the calculated data statistics in memory, wherein displaying a chart of the first data analysis information in the middle area of ​​the interface includes: displaying a chart of the first data analysis information generated based on the corresponding data statistics extracted from the memory in the middle area of ​​the interface.

[0017] Optionally, the method further includes: in response to a user selecting a field value in a chart of the first data analysis information, one or more corresponding charts of the second data analysis information are automatically updated to data statistics of the field value under corresponding other field dimensions.

[0018] Optionally, the method also includes: when one or more frequently selected field values ​​in the chart of the first data analysis information are pre-calculated and selected, the corresponding one or more charts of the second data analysis information are automatically updated to the data statistics of the frequently selected field values ​​under the corresponding other field dimensions, and the calculated data statistics are cached in the memory, wherein the corresponding one or more charts of the second data analysis information displayed are automatically updated to the data statistics of a certain field value under the corresponding other field dimensions, including: the corresponding one or more charts of the second data analysis information displayed are automatically updated to the corresponding data statistics cached in the memory.

[0019] Optionally, the method also includes: in response to the user selecting with the cursor at least one chart of the second data analysis information displayed in the right area of ​​the interface, dragging it to the middle area of ​​the interface and releasing it, displaying at least one chart of the second data analysis information in the middle area of ​​the interface.

[0020] Optionally, the method also includes: in response to a user's connection operation on two charts in the middle area of ​​the interface, connecting one of the two charts to the other chart; in response to a user's selection operation on a field value represented in the first chart in a connected state, highlighting the selected field value relative to the unselected field value, and / or updating the second chart in a connected state to represent the data statistics of the selected field value under the field dimension corresponding to the second chart.

[0021] Optionally, the method also includes: in response to a user's connection operation on two charts in the middle area of ​​the interface, connecting one of the two charts to another chart; and updating the other chart to represent the data statistics of the field values ​​represented by one chart under the field dimension corresponding to the other chart.

[0022] Optionally, the method also includes: responding to a user's prediction request for a target field, automatically training a machine learning model using field values ​​corresponding to at least part of other fields in a single piece of data as input and field values ​​corresponding to the target field in a single piece of data as output; and displaying first explanatory information for characterizing the importance of one or more fields in at least part of other fields to the machine learning model in predicting the target field.

[0023] Optionally, the method also includes: when the user selects the name icon of the target field in the left area of ​​the interface or selects the chart of the first data analysis information of the target field in the middle area of ​​the interface, in response to the operation of starting automatic training of the machine learning model performed in the right area of ​​the interface, the machine learning model is automatically trained with the field values ​​corresponding to at least part of other fields in the single piece of data as input and the field value corresponding to the target field in the single piece of data as output; and first explanatory information for characterizing the importance of one or more fields in at least part of other fields to the machine learning model in predicting the target field is displayed in the right area of ​​the interface.

[0024] Optionally, automatically training the machine learning model includes: based on hyperparameters preset based on experience, automatically searching in as small a hyperparameter space as possible to train the machine learning model.

[0025] Optionally, displaying the first explanatory information for characterizing the importance of one or more fields in at least some of the other fields to the prediction of the target field by the machine learning model includes: displaying the first explanatory information in a gradient form during the process of calculating the importance.

[0026] Optionally, the method also includes: assigning the score obtained by the machine learning model for predicting a single piece of data to each field in at least some other fields according to the distribution method of Shapley values, so as to obtain the score of the field under the single piece of data; determining the importance of the field to the machine learning model for predicting the target field based on the sum of the scores of each field in at least some other fields under multiple pieces of data, wherein the importance is positively correlated with the total score.

[0027] Optionally, the method also includes: according to the distribution method of Shapley values, assigning the score obtained by the machine learning model for the prediction of a single piece of data to each field in at least some other fields to obtain the score of the field under the single piece of data; for a single field, presenting the score of the field under multiple pieces of data in a two-dimensional coordinate system, where one coordinate axis in the two-dimensional coordinate system is used to represent the field value corresponding to the field, and the other coordinate axis is used to represent the score of the field.

[0028] Optionally, the method further includes: using a display characteristic of a coordinate point in a two-dimensional coordinate system to represent a field value corresponding to another field in the data corresponding to the coordinate point.

[0029] Optionally, the method also includes: in response to a user's selection operation on a certain piece of data, outputting a prediction result of the machine learning model for the data; and outputting second explanatory information on the importance of one or more fields in at least part of the other fields in the data to the prediction result.

[0030] Optionally, the method also includes: providing the user with a control for inputting a certain piece of data in the right area of ​​the interface, and receiving the data input by the user; displaying the prediction result of the machine learning model for the data in the right area of ​​the interface, and displaying second explanatory information on the importance of one or more fields in at least some of the other fields in the data to the prediction result.

[0031] Optionally, the method also includes: according to the distribution method of Shapley value, assigning the score obtained by the machine learning model for the prediction of the data to each field in at least some other fields, so as to obtain the score of each field under the data, and the importance is positively correlated with the score.

[0032] Optionally, the method also includes: presenting the prediction results obtained by the machine learning model for multiple data in a two-dimensional coordinate system, there are multiple coordinate points in the two-dimensional coordinate system, each coordinate point corresponds to a piece of data, the display characteristics of the coordinate points are used to characterize the prediction results of the data, and the distance between the two coordinate points in the two-dimensional space is positively correlated with the distance between the two data corresponding to the two coordinate points in the multi-dimensional space.

[0033] Optionally, each piece of data has multiple dimensions, and each dimension corresponds to a field in at least some other fields. The method also includes: according to the distribution method of Shapley values, assigning the score obtained by the machine learning model for predicting a single piece of data to each field in at least some other fields to obtain the score of the field under the single piece of data. The score of the field is the value of the data under the corresponding dimension, wherein the position of each piece of data in the multidimensional space is determined based on the values ​​of the multiple dimensions of the data.

[0034] Optionally, the method also includes: in response to a user's selection operation on one or more coordinate points in a two-dimensional coordinate system, selecting a predetermined number of coordinate points with the same prediction results as the selected coordinate point from the vicinity of the selected coordinate point to obtain a plurality of clustered coordinate points; based on the order of size of the field scores, extracting one or more key fields from the plurality of data corresponding to the plurality of clustered coordinate points to obtain a key field group; and outputting the key field group.

[0035] Optionally, the method also includes: in response to a user's adjustment operation on the field values ​​of one or more fields in a piece of data, outputting a prediction result of the machine learning model for the adjusted piece of data; and outputting second explanatory information on the importance of one or more fields in at least part of the other fields in the adjusted piece of data to the prediction result.

[0036] Optionally, the method further includes: based on the user's expected prediction result for the target field of a certain piece of data, using a machine learning model to output changes in field values ​​of at least some other fields in the data.

[0037] According to the second aspect of the present invention, a method for assisting users in exploring data tables is also proposed, comprising: in response to a user opening a data table in an application, running a plug-in for implementing the method described in the first aspect of the present invention for a data set in the data table.

[0038] According to the third aspect of the present invention, a method for assisting users in exploring data tables is also proposed, including: in response to a user opening a data table in an application, running a plug-in in the application to perform the following steps: displaying an exploration area in a predetermined area of ​​the data table; in response to a user's prediction request for a target field in the data table, automatically training a machine learning model using field values ​​corresponding to at least part of other fields in a single piece of data as input and field values ​​corresponding to the target field in the single piece of data as output; and outputting first explanatory information in the exploration area for characterizing the importance of at least part of other fields to the target field.

[0039] Optionally, the predetermined area is at least one of the left side, the right side, the top and the bottom of the data table.

[0040] Optionally, the data table is an Excel table.

[0041] Optionally, automatically training the machine learning model includes: based on hyperparameters preset based on experience, automatically searching in as small a hyperparameter space as possible to train the machine learning model.

[0042] Optionally, outputting first explanation information for characterizing the importance of at least part of other fields to the target field in the exploration area includes: displaying the first explanation information in a gradual manner during calculation of the importance.

[0043] Optionally, the method also includes: assigning the score obtained by the machine learning model for predicting a single piece of data to each field in at least some other fields according to the distribution method of Shapley values, so as to obtain the score of the field under the single piece of data; determining the importance of the field to the machine learning model for predicting the target field based on the sum of the scores of each field in at least some other fields under multiple pieces of data, wherein the importance is positively correlated with the total score.

[0044] Optionally, the method also includes: outputting a two-dimensional coordinate graph in the exploration area for representing the impact of a single field on a target field, wherein one coordinate axis in the two-dimensional coordinate graph is used to represent the field value corresponding to the field, and the other coordinate axis is used to represent the score of the field. The two-dimensional coordinate graph includes multiple coordinate points, each coordinate point corresponds to a piece of data, and the score of the field is obtained by assigning the score obtained by the machine learning model for the prediction of a single piece of data to each field in at least some other fields based on the distribution method of the Shapley value.

[0045] Optionally, the method further includes: using display characteristics of a coordinate point in the two-dimensional coordinate graph to represent a field value corresponding to another field in the data corresponding to the coordinate point.

[0046] Optionally, the method also includes: in response to a user's prediction request for a piece of data in a data table, outputting a prediction result of the machine learning model for the piece of data in the exploration area; and outputting second explanatory information in the exploration area on the importance of one or more fields in at least part of the other fields in the piece of data to the prediction result.

[0047] Optionally, the method also includes: according to the distribution method of Shapley values, allocating the score obtained by the machine learning model for the prediction of the data to each field in at least some other fields, so as to obtain the score of each field under the data, and the importance of the field to the prediction result is positively correlated with the score of the field.

[0048] Optionally, the method also includes: in response to a user's adjustment operation on the field values ​​of one or more fields in a piece of data in a data table, outputting a prediction result of the machine learning model for the adjusted piece of data; and outputting second explanatory information in the exploration area on the importance of one or more fields in at least some of the other fields in the adjusted piece of data to the prediction result.

[0049] Optionally, the method further includes: based on the user's expected prediction result for the target field of a certain data in the data table, using the machine learning model to output the changes in the field values ​​of at least some other fields in the data.

[0050] Optionally, the method also includes: in response to a user's selection operation on one or more data columns in a data table, outputting first data analysis information in the exploration area, the first data analysis information being used to characterize the data statistics of the field values ​​corresponding to the data columns selected by the user; and recommending one or more second data analysis information to the user, the second data analysis information being data analysis information predicted based on the data columns selected by the user.

[0051] Optionally, the second data analysis information is used to characterize the data statistics of the field values ​​corresponding to other data columns obtained through prediction, and / or the second data analysis information is used to characterize the data statistics of the field value combination corresponding to the data column combination consisting of the data column selected by the user and the other data columns obtained through prediction.

[0052] Optionally, the step of recommending one or more second data analysis information to the user includes: predicting data analysis information suitable for recommendation to the user based on a machine learning model; and / or predicting data analysis information suitable for recommendation to the user based on statistical correlation; and / or predicting data analysis information suitable for recommendation to the user based on business rules.

[0053] Optionally, the step of predicting data analysis information suitable for recommendation to the user based on the machine learning model includes: based on the relevant information of the data column currently selected by the user, using the machine learning model to predict the probability that the user will select each other data column later, wherein the machine learning model is trained in the following manner: taking the relevant information of the previously selected data column and the relevant information of other data columns as input, and taking the probability of other data columns being selected as output; analyzing the data statistics of the field values ​​corresponding to the data columns whose probability values ​​are greater than the first predetermined threshold to obtain second data analysis information, or analyzing the data statistics of the field values ​​corresponding to the data columns whose probability values ​​are greater than the first predetermined threshold on the basis of the first data analysis information of the data column selected by the user to obtain second data analysis information.

[0054] Optionally, the method also includes: obtaining user acceptance feedback on one or more second data analysis information recommended to the user, and updating the machine learning model based on the obtained acceptance feedback.

[0055] Optionally, the step of predicting data analysis information suitable for recommendation to the user based on statistical correlation includes: obtaining, according to the statistical correlation between different data columns, a data column whose statistical correlation with the data column selected by the user is higher than a second predetermined threshold; analyzing the data statistics of the field values ​​corresponding to the data column whose statistical correlation is higher than the second predetermined threshold to obtain second data analysis information, or analyzing the data statistics of the field values ​​corresponding to the data column whose statistical correlation is higher than the second predetermined threshold on the basis of the first data analysis information of the data column selected by the user to obtain second data analysis information.

[0056] Optionally, the exploration area includes a first interface area and a second interface area. In response to a user's selection operation on one or more data columns in a data table, outputting first data analysis information in the exploration area includes: in response to the user selecting a specific data column with a cursor, displaying a chart of the first data analysis information in the first interface area, and recommending one or more second data analysis information to the user includes: displaying one or more charts of the second data analysis information in the second interface area while displaying a chart of the first data analysis information in the first interface area.

[0057] Optionally, the method also includes: when opening a data table in an application, pre-calculating data statistics of field values ​​corresponding to each data column, and caching the calculated data statistics in memory, wherein displaying a chart of the first data analysis information in the first interface area includes: displaying a chart of the first data analysis information generated based on the corresponding data statistics extracted from the memory in the first interface area.

[0058] Optionally, the method further includes: in response to a user selecting a field value in a chart of the first data analysis information, one or more corresponding charts of the second data analysis information are automatically updated to data statistics of the field value under corresponding other field dimensions.

[0059] Optionally, the method also includes: when one or more frequently selected field values ​​in the chart of the first data analysis information are pre-calculated and selected, the corresponding one or more charts of the second data analysis information are automatically updated to the data statistics of the frequently selected field values ​​under the corresponding other field dimensions, and the calculated data statistics are cached in the memory, wherein the corresponding one or more charts of the second data analysis information displayed are automatically updated to the data statistics of a certain field value under the corresponding other field dimensions, including: the corresponding one or more charts of the second data analysis information displayed are automatically updated to the corresponding data statistics cached in the memory.

[0060] Optionally, the method further includes: in response to the user selecting with a cursor at least one chart of the second data analysis information displayed in the second interface area, dragging it to the first interface area and releasing it, displaying at least one chart of the second data analysis information in the first interface area.

[0061] Optionally, the method also includes: in response to a user's connection operation on two charts in the first interface area, connecting one of the two charts to another chart; in response to a user's selection operation on a field value represented in the first chart in a connected state, highlighting the selected field value relative to the unselected field value, and / or updating the second chart in a connected state to represent the data statistics of the selected field value under the field dimension corresponding to the second chart.

[0062] Optionally, the method also includes: in response to a user's connection operation on two charts in the first interface area, connecting one of the two charts to another chart; and updating the other chart to represent the data statistics of the field values ​​represented by one chart under the field dimensions corresponding to the other chart.

[0063] According to the fourth aspect of the present invention, a device for assisting users in exploring a data set is also proposed. The data set includes multiple data pieces, each of which includes the values ​​of one or more fields. The device includes: an output module for outputting first data analysis information in response to the user's field selection operation, and the first data analysis information is used to characterize the data statistics of the field value corresponding to the field selected by the user; and a recommendation module for recommending one or more second data analysis information to the user, and the second data analysis information is data analysis information predicted based on the field selected by the user.

[0064] Optionally, the second data analysis information is used to characterize the data statistics of the field values ​​corresponding to other fields obtained through prediction, and / or the second data analysis information is used to characterize the data statistics of the field value combination corresponding to the field combination composed of the field selected by the user and other fields obtained through prediction.

[0065] Optionally, the recommendation module includes: a first recommendation module, used to predict data analysis information suitable for recommendation to users based on a machine learning model; and / or a second recommendation module, used to predict data analysis information suitable for recommendation to users based on statistical correlation; and / or a third recommendation module, used to predict data analysis information suitable for recommendation to users based on business rules.

[0066] Optionally, the first recommendation module: based on the relevant information of the field currently selected by the user, uses a machine learning model to predict the probability that the user will select each other field later, wherein the machine learning model is trained in the following manner: taking the relevant information of the previously selected field and the relevant information of other fields as input, and taking the probability of other fields being selected as output; analyzing the data statistics of the field values ​​corresponding to the fields whose probability values ​​are greater than the first predetermined threshold to obtain second data analysis information, or analyzing the data statistics of the field values ​​corresponding to the fields whose probability values ​​are greater than the first predetermined threshold based on the first data analysis information of the field selected by the user to obtain second data analysis information.

[0067] Optionally, the first recommendation module is also used to obtain the user's acceptance feedback on one or more second data analysis information recommended to the user, and update the machine learning model based on the obtained acceptance feedback.

[0068] Optionally, the second recommendation module: according to the statistical correlation between different fields, obtains the fields whose statistical correlation with the field selected by the user is higher than the second predetermined threshold; analyzes the data statistics of the field values ​​corresponding to the fields whose statistical correlation is higher than the second predetermined threshold to obtain second data analysis information, or analyzes the data statistics of the field values ​​corresponding to the fields whose statistical correlation is higher than the second predetermined threshold based on the first data analysis information of the field selected by the user to obtain second data analysis information.

[0069] Optionally, the device also includes: a display module for displaying a first interface area and a second interface area, wherein the first data analysis information is presented in the first interface area and the second data analysis information is presented in the second interface area, and in response to the user's operation, the display module presents the second data analysis information in the second interface area in the first interface area.

[0070] Optionally, in response to the user's connection operation on two data analysis information in the first interface area, the display module connects one of the two data analysis information to another data analysis information, and the other data analysis information is updated to represent the data statistics of the field value represented by one data analysis information under the field dimension corresponding to the other data analysis information.

[0071] Optionally, in response to the user's connection operation on two data analysis information in the first interface area, the display module connects one of the two data analysis information to the other data analysis information. In response to the user's selection operation on the field value represented by one of the two data analysis information in the connected state, the display module highlights the data statistics of the selected field value in the data analysis information relative to the data statistics of the unselected field values, and / or the display module updates the other data analysis information of the two data analysis information in the connected state to represent the data statistics of the selected field value under the field dimension corresponding to the other data analysis information.

[0072] Optionally, the device also includes: a display module for displaying an interface for exploring a data set, wherein the left area of ​​the interface is used to display name icons corresponding to each field of the data set, the middle area of ​​the interface is used to display a chart of the first data analysis information, and the right area of ​​the interface is used to display a chart of the second data analysis information. When importing a data set, the display module displays the name icons corresponding to each field included in each data in the data set in the left area of ​​the interface, and the output module displays a chart of the first data analysis information in the middle area of ​​the interface in response to the user selecting the name icon of a specific field with the cursor, dragging it to the middle area of ​​the interface and releasing it. In addition, while displaying the chart of the first data analysis information in the middle area of ​​the interface, the recommendation module displays one or more charts of the second data analysis information in the right area of ​​the interface.

[0073] Optionally, the device also includes: a first calculation module, which is used to pre-calculate the data statistics of the field values ​​corresponding to each field when importing a data set, and cache the calculated data statistics in the memory, and the output module displays a chart of the first data analysis information generated based on the corresponding data statistics extracted from the memory in the middle area of ​​the interface.

[0074] Optionally, in response to the user selecting a field value in a chart of the first data analysis information, the display module automatically updates one or more corresponding charts of the second data analysis information to data statistics of the field value under corresponding other field dimensions.

[0075] Optionally, the device also includes: a second calculation module, which is used to pre-calculate that when one or more frequently selected field values ​​in the chart of the first data analysis information are selected, the corresponding one or more charts of the second data analysis information are automatically updated to the data statistics of the frequently selected field values ​​under the corresponding other field dimensions, and cache the calculated data statistics in the memory, wherein the display module automatically updates the corresponding one or more charts of the second data analysis information to the corresponding data statistics cached in the memory.

[0076] Optionally, in response to the user selecting the at least one second data analysis information chart displayed in the right area of ​​the interface with the cursor, dragging it to the middle area of ​​the interface and releasing it, the display module displays the at least one second data analysis information chart in the middle area of ​​the interface.

[0077] Optionally, in response to a user's connection operation on two charts in the middle area of ​​the interface, the display module connects one of the two charts to the other chart, and in response to a user's selection operation on a field value represented in the first chart in a connected state, the display module highlights the selected field value relative to the unselected field value, and / or the display module updates the second chart in a connected state to represent the data statistics of the selected field value under the field dimension corresponding to the second chart.

[0078] Optionally, in response to the user's connection operation on two charts in the middle area of ​​the interface, the display module connects one of the two charts to another chart, and the display module updates the other chart to represent the data statistics of the field values ​​represented by one chart under the field dimension corresponding to the other chart.

[0079] Optionally, the device also includes: a training module for responding to a user's prediction request for a target field, using the field values ​​corresponding to at least part of other fields in a single piece of data as input and the field value corresponding to the target field in a single piece of data as output, to automatically train a machine learning model; a display module for displaying first explanatory information used to characterize the importance of one or more fields in at least part of other fields to the machine learning model in predicting the target field.

[0080] Optionally, the device also includes: a training module for automatically training the machine learning model in response to an operation of starting automatic training of the machine learning model performed in the right area of ​​the interface, using field values ​​corresponding to at least part of other fields in a single piece of data as input and field values ​​corresponding to the target field in a single piece of data as output, when the user selects the name icon of the target field in the left area of ​​the interface or selects a chart of the first data analysis information of the target field in the middle area of ​​the interface; a display module for displaying in the right area of ​​the interface first explanatory information used to characterize the importance of one or more fields in at least part of other fields to the machine learning model in predicting the target field.

[0081] Optionally, the training module automatically searches in as small a hyperparameter space as possible based on hyperparameters preset based on experience to train the machine learning model.

[0082] Optionally, the display module displays the first explanation information in a gradual manner during the process of calculating the importance.

[0083] Optionally, the device also includes: a third calculation module, used to assign the score obtained by the machine learning model for predicting a single piece of data to each field in at least some other fields according to the distribution method of Shapley values, so as to obtain the score of the field under the single piece of data; a determination module, used to determine the importance of the field to the prediction of the target field by the machine learning model based on the sum of the scores of each field in at least some other fields under multiple pieces of data, wherein the importance is positively correlated with the total score.

[0084] Optionally, the device also includes: a fourth calculation module, which is used to assign the score obtained by the machine learning model for predicting a single piece of data to each field in at least some other fields according to the distribution method of the Shapley value, so as to obtain the score of the field under the single piece of data. For a single field, the display module presents the score of the field under multiple pieces of data in a two-dimensional coordinate system, where one coordinate axis in the two-dimensional coordinate system is used to represent the field value corresponding to the field, and the other coordinate axis is used to represent the score of the field.

[0085] Optionally, the display module further uses the display characteristics of a coordinate point in the two-dimensional coordinate system to represent a field value corresponding to another field in the data corresponding to the coordinate point.

[0086] Optionally, in response to a user's selection operation on a certain piece of data, the display module also outputs the prediction result of the machine learning model for the data, and outputs second explanatory information on the importance of one or more fields in at least part of the other fields in the data to the prediction result.

[0087] Optionally, the display module is also used to provide the user with a control for inputting a certain piece of data in the right area of ​​the interface, and receive the data input by the user through the control. The display module is also used to display the prediction result of the machine learning model for the data in the right area of ​​the interface, and display second explanatory information on the importance of one or more fields in at least part of the other fields in the data to the prediction result.

[0088] Optionally, the device also includes: a fifth calculation module, which is used to assign the score obtained by the machine learning model for predicting the data to each field in at least some other fields according to the distribution method of the Shapley value, so as to obtain the score of each field under the data, and the importance is positively correlated with the score.

[0089] Optionally, the display module is also used to present the prediction results obtained by the machine learning model for multiple data in a two-dimensional coordinate system. There are multiple coordinate points in the two-dimensional coordinate system, and each coordinate point corresponds to a piece of data. The display characteristics of the coordinate points are used to characterize the prediction results of the data. The distance between two coordinate points in the two-dimensional space is positively correlated with the distance between the two data corresponding to the two coordinate points in the multidimensional space.

[0090] Optionally, each piece of data has multiple dimensions, and each dimension corresponds to a field in at least some of the other fields. The device also includes: a sixth calculation module, which is used to assign the score obtained by the machine learning model for predicting a single piece of data to each field in at least some of the other fields according to the distribution method of the Shapley value, so as to obtain the score of the field under the single piece of data. The score of the field is the value of the data under the corresponding dimension, wherein the position of each piece of data in the multidimensional space is determined based on the values ​​of the multiple dimensions of the data.

[0091] Optionally, the device also includes: a selection module, which is used to respond to the user's selection operation on one or more coordinate points in the two-dimensional coordinate system, and select a predetermined number of coordinate points with the same prediction results as the selected coordinate point from the vicinity of the selected coordinate point to obtain multiple clustered coordinate points; an extraction module, which is used to extract one or more key fields from the multiple data corresponding to the multiple clustered coordinate points based on the order of the size of the field scores to obtain a key field group, and the display module is also used to output the key field group.

[0092] Optionally, the display module is also used to output the prediction result of the machine learning model for the adjusted data in response to the user's adjustment operation on the field value of one or more fields in a piece of data, and output second explanatory information on the importance of one or more fields in at least part of the other fields in the adjusted data to the prediction result.

[0093] Optionally, the display module is also used to use a machine learning model to output changes in field values ​​of at least some other fields in a piece of data based on the user's expected prediction results for the target field of the data.

[0094] According to the fifth aspect of the present invention, a device for assisting users in exploring data tables is also proposed, including: a running module for running a plug-in for implementing the method described in the first aspect of the present invention for a data set in the data table in response to a user opening the data table in an application.

[0095] According to the sixth aspect of the present invention, a device for assisting users in exploring data tables is also proposed, including: a display module for displaying an exploration area in a predetermined area of ​​the data table in response to a user opening a data table in an application; a training module for responding to a user's prediction request for a target field in the data table, using field values ​​corresponding to at least part of other fields in a single piece of data as input and field values ​​corresponding to the target field in the single piece of data as output, to automatically train a machine learning model, and the display module is also used to output first explanatory information in the exploration area for characterizing the importance of at least part of other fields to the target field.

[0096] Optionally, the predetermined area is at least one of the left side, the right side, the top and the bottom of the data table.

[0097] Optionally, the data table is an Excel table.

[0098] Optionally, the training module automatically searches in as small a hyperparameter space as possible based on hyperparameters preset based on experience to train the machine learning model.

[0099] Optionally, the display module displays the first explanation information in a gradual manner during the process of calculating the importance.

[0100] Optionally, the device also includes: a first calculation module, used to assign the score obtained by the machine learning model for predicting a single piece of data to each field in at least some other fields according to the distribution method of Shapley values, so as to obtain the score of the field under the single piece of data; a determination module, used to determine the importance of the field to the prediction of the target field by the machine learning model based on the sum of the scores of each field in at least some other fields under multiple pieces of data, wherein the importance is positively correlated with the total score.

[0101] Optionally, the display module is also used to output a two-dimensional coordinate graph in the exploration area for representing the impact of a single field on a target field. One coordinate axis in the two-dimensional coordinate graph is used to represent the field value corresponding to the field, and the other coordinate axis is used to represent the score of the field. The two-dimensional coordinate graph includes multiple coordinate points, each coordinate point corresponds to a piece of data, and the score of the field is obtained by assigning the score obtained by the machine learning model for the prediction of a single piece of data to each field in at least some other fields based on the Shapley value distribution method.

[0102] Optionally, the display module is further configured to use the display characteristics of a coordinate point in the two-dimensional coordinate graph to represent a field value corresponding to another field in the data corresponding to the coordinate point.

[0103] Optionally, in response to a user's prediction request for a piece of data in a data table, the display module outputs the prediction result of the machine learning model for the piece of data in the exploration area, and outputs second explanatory information in the exploration area on the importance of one or more fields in at least some of the other fields in the piece of data to the prediction result.

[0104] Optionally, the device also includes: a second calculation module, which is used to assign the score obtained by the machine learning model for predicting the data to each field in at least some other fields according to the distribution method of Shapley values, so as to obtain the score of each field under the data, and the importance of the field to the prediction result is positively correlated with the score of the field.

[0105] Optionally, in response to a user's adjustment operation on the field values ​​of one or more fields in a piece of data in a data table, the display module also outputs the prediction result of the machine learning model for the adjusted piece of data, and outputs second explanatory information in the exploration area on the importance of one or more fields in at least some of the other fields in the adjusted piece of data to the prediction result.

[0106] Optionally, the display module also uses a machine learning model to output changes in field values ​​of at least some other fields in a certain data item in the data table based on the user's expected prediction results for the target field of the data item.

[0107] Optionally, the device also includes: an output module for outputting first data analysis information in the exploration area in response to a user's selection operation on one or more data columns in a data table, the first data analysis information being used to characterize the data statistics of the field values ​​corresponding to the data columns selected by the user; a recommendation module for recommending one or more second data analysis information to the user, the second data analysis information being data analysis information predicted based on the data columns selected by the user.

[0108] Optionally, the second data analysis information is used to characterize the data statistics of the field values ​​corresponding to other data columns obtained through prediction, and / or the second data analysis information is used to characterize the data statistics of the field value combination corresponding to the data column combination consisting of the data column selected by the user and the other data columns obtained through prediction.

[0109] Optionally, the recommendation module includes: a first recommendation module, used to predict data analysis information suitable for recommendation to users based on a machine learning model; and / or a second recommendation module, used to predict data analysis information suitable for recommendation to users based on statistical correlation; and / or a third recommendation module, used to predict data analysis information suitable for recommendation to users based on business rules.

[0110] Optionally, the first recommendation module: based on the relevant information of the data column currently selected by the user, uses a machine learning model to predict the probability that the user will select each other data column later, wherein the machine learning model is trained in the following manner: taking the relevant information of the previously selected data column and the relevant information of other data columns as input, and taking the probability of other data columns being selected as output; analyzing the data statistics of the field values ​​corresponding to the data columns whose probability values ​​are greater than the first predetermined threshold to obtain second data analysis information, or analyzing the data statistics of the field values ​​corresponding to the data columns whose probability values ​​are greater than the first predetermined threshold on the basis of the first data analysis information of the data column selected by the user to obtain second data analysis information.

[0111] Optionally, the first recommendation module also obtains the user's acceptance feedback on the one or more second data analysis information recommended to it, and updates the machine learning model based on the obtained acceptance feedback.

[0112] Optionally, the second recommendation module: according to the statistical correlation between different data columns, obtains the data column whose statistical correlation with the data column selected by the user is higher than the second predetermined threshold; analyzes the data statistics of the field values ​​corresponding to the data column whose statistical correlation is higher than the second predetermined threshold to obtain second data analysis information, or analyzes the data statistics of the field values ​​corresponding to the data column whose statistical correlation is higher than the second predetermined threshold based on the first data analysis information of the data column selected by the user to obtain second data analysis information.

[0113] Optionally, the exploration area includes a first interface area and a second interface area, and the output module displays a chart of the first data analysis information in the first interface area in response to the user selecting a specific data column with the cursor. In addition, while displaying the chart of the first data analysis information in the first interface area, the recommendation module displays one or more charts of the second data analysis information in the second interface area.

[0114] Optionally, the device also includes: a third calculation module, which is used to pre-calculate the data statistics of the field values ​​corresponding to each data column when opening the data table in the application, and cache the calculated data statistics in the memory, wherein the output module displays a chart of the first data analysis information generated based on the corresponding data statistics extracted from the memory in the first interface area.

[0115] Optionally, in response to the user selecting a field value in a chart of the first data analysis information, the recommendation module automatically updates one or more corresponding charts of the second data analysis information to data statistics of the field value under corresponding other field dimensions.

[0116] Optionally, the device also includes: a fourth calculation module, which is used to pre-calculate that when one or more frequently selected field values ​​in the chart of the first data analysis information are selected, the corresponding one or more charts of the second data analysis information are automatically updated to the data statistics of the frequently selected field values ​​under the corresponding other field dimensions, and cache the calculated data statistics in the memory, and the recommendation module automatically updates the corresponding displayed charts of the one or more second data analysis information to the corresponding data statistics cached in the memory.

[0117] Optionally, in response to the user selecting the at least one chart of the second data analysis information displayed in the second interface area with a cursor, dragging it to the first interface area and releasing it, the display module displays the at least one chart of the second data analysis information in the first interface area.

[0118] Optionally, in response to a user's connection operation on two charts in the first interface area, the display module connects one of the two charts to the other chart, and in response to a user's selection operation on a field value represented in the first chart in a connected state, the display module highlights the selected field value relative to the unselected field value, and / or the display module updates the second chart in a connected state to represent the data statistics of the selected field value under the field dimension corresponding to the second chart.

[0119] Optionally, in response to the user's connection operation on two charts in the first interface area, the display module connects one of the two charts to another chart, and the display module updates the other chart to represent the data statistics of the field values ​​represented by one chart under the field dimension corresponding to the other chart.

[0120] According to the seventh aspect of the present invention, a system is also proposed, comprising at least one computing device and at least one storage device storing instructions, wherein, when the instructions are executed by the at least one computing device, they prompt the at least one computing device to execute the method as described in any one of the first to third aspects of the present invention.

[0121] According to the eighth aspect of the present invention, a computer-readable storage medium storing instructions is also proposed, wherein, when the instructions are executed by at least one computing device, the at least one computing device is prompted to execute the method described in any one of the first to third aspects of the present invention.

[0122] In the method and apparatus for assisting users in exploring datasets and data tables according to exemplary embodiments of the present invention, first data analysis information is output, allowing users to understand the data statistics of the selected field. Second data analysis information is predicted based on the user-selected field and is likely to be of interest to the user. By recommending this second data analysis information to the user, the user's data exploration costs can be reduced, enabling users to explore datasets quickly and conveniently. BRIEF DESCRIPTION OF THE DRAWINGS

[0123] These and / or other aspects and advantages of the present invention will become more apparent and more readily understood from the following detailed description of embodiments of the present invention in conjunction with the accompanying drawings, in which:

[0124] Figure 1 A flowchart of a method for assisting a user in exploring a data set according to an exemplary embodiment of the present invention is shown;

[0125] Figures 2 to 9 A schematic diagram showing a data exploration interface according to an exemplary embodiment of the present invention is shown;

[0126] Figures 10 to 14 A schematic diagram showing an exploration area shown in a data table according to an exemplary embodiment of the present invention;

[0127] Figure 15 A structural block diagram of an apparatus for assisting a user in exploring a data set according to an exemplary embodiment of the present invention is shown;

[0128] Figure 16 A structural block diagram of an apparatus for assisting a user in exploring a data table according to an exemplary embodiment of the present invention is shown. DETAILED DESCRIPTION

[0129] In order to enable those skilled in the art to better understand the present invention, exemplary embodiments of the present invention are further described in detail below with reference to the accompanying drawings and specific embodiments.

[0130] Figure 1 A flow chart of a method for assisting a user in exploring a data set according to an exemplary embodiment of the present invention is shown. The data set includes multiple pieces of data, each piece of data includes values ​​of one or more fields. Figure 1 The method shown can be implemented entirely in software by a computer program or executed by a specially configured computing device. Figure 1The method shown.

[0131] See also Figure 1 In step S110, in response to the user's field selection operation, first data analysis information is output, where the first data analysis information is used to characterize the data statistics of the field value corresponding to the field selected by the user.

[0132] Outputting the first data analysis information may refer to outputting the first data analysis information in a visual manner. For example, the first data analysis information may be presented to the user in the form of, but not limited to, a chart. Specifically, the first data analysis information may be a chart generated by analyzing field values ​​corresponding to a field selected by the user in the dataset.

[0133] In step S120 , one or more second data analysis information is recommended to the user, where the second data analysis information is data analysis information predicted based on the field selected by the user.

[0134] Recommending one or more pieces of second data analysis information to a user may refer to presenting the one or more pieces of second data analysis information to the user in a visual manner. For example, the second data analysis information may be presented to the user in the form of, but not limited to, a chart, i.e., the second data analysis information may be a chart recommended to the user.

[0135] The second data analysis information can be used to characterize the data statistics of the field values ​​corresponding to other fields obtained through prediction, and / or the second data analysis information can also be used to characterize the data statistics of the field value combination corresponding to the field combination composed of the field selected by the user and other fields obtained through prediction.

[0136] Taking the employee information table as an example, when the user selects the "Department" field, the first data analysis information can be a chart used to represent the data statistics of the field values ​​corresponding to the "Department" field (such as Human Resources Department, R&D Department, Sales Department). The second data analysis information recommended to the user can be a chart based on the statistics obtained by analyzing the "Whether to work overtime" field grouped by "Department", or it can be a chart based on the statistics obtained by analyzing the field values ​​(yes, no) corresponding to the "Whether to resign" field.

[0137] In the present invention, one or more of a variety of methods such as, but not limited to, machine learning models, statistical correlations, business rules, etc. may be used to predict data analysis information suitable for recommendation to users.

[0138] Taking the prediction of data analysis information suitable for recommendation to users based on a machine learning model as an example, the probability that the user will select each other field in the future can be predicted using a machine learning model based on the relevant information of the field currently selected by the user (such as but not limited to field name, data type, data distribution, etc.), wherein the machine learning model is trained in the following manner: taking the relevant information of the previously selected field and the relevant information of other fields as input, and taking the probability of other fields being selected as output; analyzing the data statistics of the field values ​​corresponding to the fields whose probability values ​​are greater than the first predetermined threshold to obtain second data analysis information, or analyzing the data statistics of the field values ​​corresponding to the fields whose probability values ​​are greater than the first predetermined threshold on the basis of the first data analysis information of the field selected by the user to obtain second data analysis information.

[0139] Optionally, the user's acceptance feedback on one or more second data analysis information recommended to the user may be obtained, and the machine learning model may be updated based on the obtained acceptance feedback.

[0140] Taking data analysis information suitable for recommendation to users based on statistical correlation prediction as an example, according to the statistical correlation between different fields, fields whose statistical correlation with the field selected by the user is higher than a second predetermined threshold can be obtained; the data statistics of the field values ​​corresponding to the fields whose statistical correlation is higher than the second predetermined threshold are analyzed to obtain the second data analysis information, or the data statistics of the field values ​​corresponding to the fields whose statistical correlation is higher than the second predetermined threshold are analyzed on the basis of the first data analysis information of the field selected by the user to obtain the second data analysis information.

[0141] For example, predicting data analysis information suitable for user recommendations based on business rules can be used to summarize frequently used data analysis and visualization patterns in business scenarios based on business experience. When predicting the second data analysis information, this can be based on business experience in the same or similar scenarios. For example, if user interviews or actual business experience reveal that users often focus on "net profit" when analyzing "sales," then when the "sales" field is selected, data analysis information related to "net profit" can be recommended.

[0142] In the present invention, step S110 and step S120 can be performed simultaneously, in no particular order. That is, in response to the user's field selection operation, the first data analysis can be output and one or more second data analysis information can be recommended to the user at the same time.

[0143] By outputting the first data analysis information, users can understand the data statistics of the selected field. The second data analysis information is the data analysis information that the user may be interested in, which is predicted based on the user's selected field. Therefore, by recommending the second data analysis information to the user, the user's data exploration cost can be reduced, allowing users to explore the dataset conveniently and quickly.

[0144] As an example, when displaying the first data analysis information / the second data analysis information, the outliers and / or significant values ​​therein may be highlighted (e.g., highlighted), where an outlier refers to a value in the data that is significantly different from other values, and a significant value refers to a value or value segment that appears more frequently.

[0145] The present invention can display an interface for exploring a data set. The first data analysis information and the second data analysis information can be presented in different areas of the interface. For example, the first data analysis information can be presented in a first interface area, and the second data analysis information can be presented in a second interface area. The first interface area and the second interface area are two different interface areas within the same interface.

[0146] The first interface area can be considered an interactive area, where users can perform predetermined operations on the data analysis information displayed in the first interface area to achieve corresponding functions. The second interface area is a recommendation area, where users can drag the second recommended information that meets their needs to the first interface area to perform subsequent operations on it.

[0147] That is, in response to a user operation, the second data analysis information in the second interface area can also be presented in the first interface area. For example, in response to a user selecting the second data analysis information displayed in the second interface area with a cursor, dragging the cursor to the first interface area, and releasing the cursor, the second data analysis information presented in the second interface area can be presented in the first interface area. After the second data analysis information in the second interface area is dragged to the first interface area, the second interface area may no longer display the second data analysis information, or may continue to display the second data analysis information.

[0148] Thus, the first interface area can display one or more data analysis information, wherein the data analysis information displayed in the first interface area can be either the first data analysis information or the second data analysis information from the second interface area.

[0149] As an example of the present invention, in response to a user's operation to connect two pieces of data analysis information in the first interface area, one piece of data analysis information is connected to the other piece of data analysis information, and the other piece of data analysis information is updated to represent the data statistics of the field value represented by the first piece of data analysis information under the field dimension corresponding to the other piece of data analysis information. Thus, by connecting the two pieces of data analysis information, the user can obtain the data statistics under the dimensions represented by the two pieces of data analysis information.

[0150] For example, the first interface area may include Chart 1, which is a statistical analysis of the field values ​​corresponding to the field "Department," and Chart 2, which is a statistical analysis of the field values ​​corresponding to the field "Resignation Status." In response to a user's operation to connect Chart 1 and Chart 2, Chart 1 can be visually connected to Chart 2 (e.g., by a line), and Chart 2 can be automatically updated to display a chart showing the resignation status of different departments, obtained by grouping the resignations by department.

[0151] As another example of the present invention, in response to a user's connection operation on two data analysis information in the first interface area, one of the two data analysis information is connected to the other data analysis information; in response to a user's selection operation on a field value represented by one of the two data analysis information in the connected state, the data statistics of the selected field value in the data analysis information are highlighted relative to the data statistics of the unselected field values, and / or the other of the two data analysis information in the connected state is updated to represent the data statistics of the field dimension corresponding to the other data analysis information for the selected field value. Thus, in response to the user connecting the two data analysis information and filtering one of the data analysis information, the other data analysis information can be automatically updated to the filtered data statistics.

[0152] For example, the first interface area may include a chart 1 obtained by statistically analyzing the field values ​​corresponding to the field "Department" and a chart 2 obtained by statistically analyzing the field values ​​corresponding to the field "Whether to resign", wherein Chart 1 can be used to represent the headcount distribution of the "Human Resources Department", "R&D Department" and "Sales Department". In response to the user's connection operation on Chart 1 and Chart 2, Chart 1 can be visually connected to Chart 2 (for example, by a line), and in response to the user's selection operation on "R&D Department" in Chart 1, the data statistics of "R&D Department" in Chart 1 are highlighted relative to the data statistics of "Human Resources Department" and "Sales Department", and / or Chart 2 is updated to a chart used to represent the resignation statistics of "R&D Department".

[0153] Figures 2 to 9A schematic diagram of a data exploration interface according to an exemplary embodiment of the present invention is shown. Figures 2 to 9 The method of assisting a user in exploring a data set according to the present invention is further described.

[0154] See also Figure 2 , you can display an interface for exploring the data set, which is divided into three parts: the left area of ​​the interface, the middle area of ​​the interface, and the right area of ​​the interface (that is, the insights area shown in the figure).

[0155] The left area of ​​the interface is used to display the name icons corresponding to each field of the dataset. Figure 2 Taking the employee information table as an example, the name icons corresponding to some fields of the dataset (department, gender, whether overtime is worked, and whether the employee has resigned) are displayed. It should be noted that all or some of the fields of the dataset can be displayed in the left area of ​​the interface according to the actual situation. For example, if the number of fields in the dataset is small (such as below the threshold), all fields can be displayed in the left area of ​​the interface. If the number of fields in the dataset is large (such as above the threshold), some fields can be selectively displayed in the left area of ​​the interface.

[0156] The middle area of ​​the interface corresponds to the first interface area mentioned above, and is used to display a chart of the first data analysis information. The right area of ​​the interface corresponds to the second interface area mentioned above, and is used to display a chart of the second data analysis information.

[0157] When importing a data set, the present invention can display the name icons corresponding to the various fields included in each data in the data set in the left area of ​​the interface.

[0158] In response to a user selecting the name icon of a specific field with the cursor, dragging the cursor to the center area of ​​the interface, and releasing the cursor, a chart of the first data analysis information is displayed in the center area of ​​the interface. Simultaneously with the chart of the first data analysis information being displayed in the center area of ​​the interface, one or more charts of the second data analysis information are displayed in the right area of ​​the interface. The charts displayed in the right area of ​​the interface may be predicted based on any one of the prediction methods mentioned above or a combination of the prediction methods.

[0159] like Figure 3 As shown, in response to the user selecting the name icon of the "Department" field with the cursor, dragging it to the middle area of ​​the interface and releasing it, a chart representing the data statistics of the field values ​​corresponding to the "Department" field in the data set is displayed in the middle area of ​​the interface. At the same time, a chart obtained by counting the number of "Whether to work overtime" grouped by "Department" and a chart obtained by summing the "Monthly Income" grouped by "Department" are displayed in the right area of ​​the interface.

[0160] To ensure that data analysis information can be displayed to users in a timely manner, when importing a data set, the data statistics of the field values ​​corresponding to each field can be pre-calculated and cached in memory. Therefore, in response to the user's field selection operation, there is no need to further analyze the field values ​​corresponding to the user-selected field. Instead, a chart of the first data analysis information generated based on the corresponding data statistics extracted from memory can be displayed in the center area of ​​the interface.

[0161] In response to a user selecting a field value in a chart of the first data analysis information, one or more corresponding charts of the second data analysis information are automatically updated to data statistics of the field value under corresponding other field dimensions.

[0162] Still Figure 3 For example, in response to the user selecting "R&D Department" in the chart displayed in the middle area of ​​the interface, the chart obtained by counting the number of "whether to work overtime" grouped by "Department" and displayed in the right area of ​​the interface is automatically updated to a chart obtained by counting the number of "whether to work overtime" grouped by "R&D Department", and the chart obtained by summing up the "monthly income" grouped by "Department" and displayed in the right area of ​​the interface is automatically updated to a chart obtained by summing up the "monthly income" grouped by "R&D Department".

[0163] Accordingly, in order to enable data analysis information to be displayed to users in a timely manner, it can be pre-calculated that when one or more frequently selected field values ​​in the chart of the first data analysis information are selected, the corresponding one or more charts of the second data analysis information are automatically updated to the data statistics of the frequently selected field values ​​under the corresponding other field dimensions, and the calculated data statistics are cached in the memory, wherein the corresponding one or more charts of the second data analysis information displayed are automatically updated to the data statistics of a certain field value under the corresponding other field dimensions, including: the corresponding one or more charts of the second data analysis information displayed are automatically updated to the corresponding data statistics cached in the memory.

[0164] The user can move a chart of the second data analysis information displayed in the right area of ​​the interface to the central area of ​​the interface by performing a predetermined operation. For example, in response to the user selecting at least one chart of the second data analysis information displayed in the right area of ​​the interface with a cursor, dragging the cursor to the central area of ​​the interface, and releasing the cursor, the at least one chart of the second data analysis information is displayed in the central area of ​​the interface. After dragging a chart from the right area of ​​the interface to the central area of ​​the interface, the chart may no longer be displayed in the right area of ​​the interface, or the chart may continue to be displayed.

[0165] The user can continue to select new fields to display multiple charts of the first data analysis information in the middle area of ​​the interface. Figure 4As shown, after the user selects the name icon of the "Department" field, drags it to the middle area of ​​the interface and releases it, the user can also select the name icon of the "Whether to Resign" field, drag it to the middle area of ​​the interface and release it, so that a chart representing the data statistics of the field values ​​corresponding to the "Whether to Resign" field in the data set is displayed in the middle area of ​​the interface.

[0166] In response to the addition of a chart of the first data analysis information in the middle area of ​​the interface, the chart of the second data analysis information displayed in the right area of ​​the interface is also automatically updated. The chart of the second data analysis information displayed in the right area of ​​the interface can be automatically updated to a chart of the data analysis information predicted based on the field selected by the user last time. And / or the chart of the second data analysis information displayed in the right area of ​​the interface can also be automatically updated to a chart of the data analysis information predicted based on multiple fields previously selected by the user. For the prediction method, please refer to the relevant description above and will not be repeated here.

[0167] like Figure 4 As shown, the chart of the second data analysis information displayed in the right area of ​​the interface can be updated to a chart used to represent the high correlation between "whether to resign" and "whether to work overtime", and a chart used to represent the statistical number of "whether to resign" grouped by "department".

[0168] like Figure 5 As shown, in response to a user's connection operation on two charts in the middle area of ​​the interface, one of the two charts is connected to the other chart, and in response to a user's selection operation on a field value (e.g., sales department) represented in the first chart in a connected state (e.g., a chart corresponding to a department), the selected field value (e.g., sales department) is highlighted relative to the unselected field values ​​(e.g., human resources department, research and development department), and / or the second chart in a connected state (e.g., a chart corresponding to whether or not the employee has resigned) is updated to represent the data statistics of the selected field value (sales department) under the field dimension (whether or not the employee has resigned) corresponding to the second chart. Thus, by connecting two charts and filtering one of the charts, the other chart can be automatically updated to the filtered data statistics.

[0169] In response to a user's linking operation on two charts in the middle area of ​​the interface, connecting one chart to the other, the other chart can be updated to represent the data statistics of the field values ​​represented by the first chart under the field dimensions corresponding to the second chart. By linking the two data analysis information, the data statistics under the dimensions represented by the two data analysis information can be obtained.

[0170] like Figure 4As shown, the middle area of ​​the interface includes Chart 1, which is obtained by counting the field values ​​corresponding to the field "Department", and Chart 2, which is obtained by counting the field values ​​corresponding to the field "Resignation". Chart 1 can be used to represent the headcount distribution of the "Human Resources Department", "R&D Department", and "Sales Department". In response to the user's connection operation on Chart 1 and Chart 2, Chart 1 can be visually connected to Chart 2 (for example, by connecting a line). At this time, Chart 2 can be automatically updated to a chart of the resignation situation of different departments obtained by counting whether the employees have resigned according to department groups.

[0171] In response to a user's prediction request for a target field, the present invention can automatically train a machine learning model using the field values ​​corresponding to at least some other fields in a single piece of data as input and the field value corresponding to the target field in the single piece of data as output. When automatically training a machine learning model, an automatic search can be performed in a minimal hyperparameter space based on hyperparameters preset based on experience to train the machine learning model. This allows a machine learning model to be generated in a short period of time that is effective, concise, and highly interpretable, thereby assisting users in making business decisions.

[0172] After training the machine learning model, first explanatory information can be displayed to indicate the importance of one or more fields in at least some of the other fields to the target field predicted by the machine learning model. During the importance calculation process, the first explanatory information can be displayed in a gradient manner. The first explanatory information can be used to indicate the importance of all fields to the target field predicted by the machine learning model, or it can be used to indicate certain fields that are more important to the target field predicted by the machine learning model.

[0173] like Figures 3 to 5 As shown, the right area of ​​the interface can provide a control for starting the automatic training of the machine learning model. When the user selects the name icon of the target field in the left area of ​​the interface or selects the chart of the first data analysis information of the target field in the middle area of ​​the interface, in response to the operation of starting the automatic training of the machine learning model performed in the right area of ​​the interface (for example, clicking the control "AutoML-Automatic Training Parsing Model"), the field values ​​corresponding to at least part of other fields in the single data are used as input, and the field value corresponding to the target field in the single data is used as output to automatically train the machine learning model. In addition, the first explanatory information used to characterize the importance of one or more fields in at least part of other fields to the prediction of the target field by the machine learning model can be displayed in the right area of ​​the interface.

[0174] In the present invention, the calculation of the importance of a field can adopt the distribution method of the Shapley value in game theory, and convert the calculation of the importance of the field into a fair distribution problem of rights and interests (i.e., predicted values) in the case of multi-field collaboration. As an example, the present invention can distribute the score obtained by the machine learning model for predicting a single piece of data to each field in at least some other fields according to the distribution method of the Shapley value, so as to obtain the score of the field under the single piece of data, and then determine the importance of the field to the target field predicted by the machine learning model based on the sum of the scores of each field in at least some other fields under multiple pieces of data, and the importance is positively correlated with the sum of the scores.

[0175] The score of a field obtained by the Shapley value distribution method for a single piece of data can be positive or negative. A positive value indicates that the field improves the predicted value of the target field obtained by the machine learning model, that is, it has a positive effect on the target field. A negative value indicates that the field reduces the predicted value of the target field obtained by the machine learning model, that is, it has a negative effect on the target field.

[0176] When calculating the total score of each field under multiple data, the sum of the absolute values ​​of the scores of the field under multiple data can be used as the importance of the field to the machine learning model predicting the target field, or the average of the absolute values ​​of the scores of the field under multiple data can be used as the importance of the field to the machine learning model predicting the target field.

[0177] In the process of calculating the importance, the first explanatory information can be displayed in a gradient form. For example, since it takes a long time to calculate the importance, the first explanatory information can be roughly displayed in a relatively vague state first. As the calculation is completed more and more, the displayed first explanatory information gradually becomes clearer. Regarding the display form of the first explanatory information, as an example, different display characteristics can be given to the field according to the importance of the calculated field. For example, the greater the importance of the field, the darker the color. Optionally, different colors can be used to identify whether the field has a positive or negative effect on the target field. For example, red can be used to represent a positive effect, and blue can be used to represent a negative effect. That is, the greater the importance of the field to the target field and the positive effect it has, the redder its color is; the greater the importance of the field to the target field and the negative effect it has, the bluer its color is.

[0178] like Figure 6A As shown, the first explanation information (ie, the automatic analysis result of the model shown on the right) can be displayed in the right area of ​​the interface. Figure 6B Shown Figure 6A A magnified schematic diagram of the automatic analysis results of the model.

[0179] like Figure 6BAs shown, based on the calculated importance of the fields, business-oriented language can be used to describe to the user the columns of fields that have the greatest impact on the target field (i.e., the greatest importance): whether to work overtime, equity option level, and job role.

[0180] Figure 6B Each row in the chart represents a field, the horizontal axis is the SHAP value, and each point represents a piece of data. The SHAP value corresponding to each point in a field is the score of the field value corresponding to the data corresponding to each point under the data. It should be noted that Figure 6B It is a grayscale image without color shown for the purpose of application documents. Figure 6B The vertical line on the right side of the text indicates that the size of the field value can be represented by a color gradient, that is, the size of the field value can be represented by a display characteristic. As an example, the field value can be represented from small to large by the form of color A (for example, blue) to color B (for example, red). For example, the larger the field value, the redder the color, and the smaller the field value, the bluer the color.

[0181] Figure 6B The red values ​​in the "Whether Overtime" field are located in the right half, indicating that more overtime hours increase the probability of leaving. The blue values ​​in the "Whether Overtime" field are located in the left half, indicating that less overtime hours decrease the probability of leaving. Similarly, the red values ​​in the "Stock Option Level" field are located in the left half, indicating that a high stock option level decreases the probability of leaving. The blue values ​​in the "Stock Option Level" field are mostly located in the right half, indicating that a low stock option level increases the probability of leaving.

[0182] Therefore, by displaying the relationship between the size of the field values ​​of different fields in the dataset and the score of the field under different data in the dataset (i.e., SHAP value) in different colors, users can intuitively understand the impact of the field value size of different fields on the target field.

[0183] Figure 6B The following figure shows the influence of the field value size on the target field in the form of color. Figure 7 As shown, the present invention can also display the influence of a field value of a certain field on a target field in the form of a two-dimensional coordinate graph.

[0184] Specifically, the present invention can assign the score obtained by the machine learning model for predicting a single piece of data to each field in at least part of the other fields according to the distribution method of Shapley values, so as to obtain the score of the field under the single piece of data. The score can be positive or negative, and the details can be found in the relevant description above, which will not be repeated here. For a single field, the score of the field under multiple pieces of data in the data set can be presented in a two-dimensional coordinate system, where one coordinate axis in the two-dimensional coordinate system is used to represent the field value corresponding to the field, and the other coordinate axis is used to represent the score of the field (i.e., the SHAP value).

[0185] See also Figure 7 , taking the "Age" field as an example, the horizontal axis is the age value, the vertical axis is the SHAP value, and each point in the figure represents a piece of data. Figure 7 It can be seen that the younger the age, the larger the SHAP value, indicating that the younger the age, the greater the probability of leaving the job, and the SHAP value of employees between the ages of 30 and 45 is smaller, indicating that the probability of leaving the job between the ages of 30 and 45 is smaller.

[0186] Therefore, by representing the relationship between the field value size and the score (ie, SHAP value) of a field in a two-dimensional coordinate system, users can also intuitively understand the impact of the field value size of a certain field on the target field.

[0187] Optionally, the display characteristics of a coordinate point in a two-dimensional coordinate system may be used to represent a field value corresponding to another field in the data corresponding to the coordinate point. Figure 7 It is a grayscale image without color shown for the purpose of application documents. Figure 7 The vertical line shown on the right side of the figure means that the length of service can be represented by color. For example, the length of service can be represented from short to long by the different forms from color A (for example, blue) to color B (for example, red). For example, the longer the length of service, the redder the color, and the shorter the length of service, the bluer the color. Figure 7 The display characteristics of the coordinate point can be used to represent the length of service in the data corresponding to the coordinate point. For example, the redder the coordinate point, the longer the length of service at the coordinate point, and the bluer the coordinate point, the shorter the length of service at the coordinate point. Figure 7 Coordinate points with high SHAP values ​​between the ages of 20 and 30 are blue, indicating that employees with high SHAP values ​​between the ages of 20 and 30 tend to have shorter service years. Coordinate points with low SHAP values ​​between the ages of 30 and 45 are mostly blue, indicating that employees between the ages of 30 and 45 have lower SHAP values ​​for shorter service years. Therefore, based on the distribution of the display characteristics of coordinate points in the two-dimensional coordinate system, users can understand the impact of multi-field dimensions on the target field.

[0188] In response to a user's selection operation on a piece of data, the prediction result of the machine learning model for the piece of data may be output, and second explanatory information on the importance of one or more fields in at least some of the other fields in the piece of data to the prediction result may also be output. For example, a control for inputting a piece of data may be provided to the user in the right area of ​​the interface, and the data input by the user may be received; the prediction result of the machine learning model for the piece of data may be displayed in the right area of ​​the interface, and second explanatory information on the importance of one or more fields in at least some of the other fields in the piece of data to the prediction result may be displayed.

[0189] The second explanatory information is used to represent the importance of each field in the data piece to the prediction result of the target field obtained by the machine learning model for the data piece. The importance here can also be calculated using the Shapley value distribution method. That is, based on the Shapley value distribution method, the score obtained by the machine learning model for the data piece can be distributed to each field in at least some other fields to obtain a score for each field in the data piece. The importance is positively correlated with the score.

[0190] like Figure 8 As shown in the figure, the probability of resignation predicted by a machine learning model for a certain data point is 0.49, and the importance (i.e., influence) of several field values ​​that have a significant impact on the predicted probability is displayed. Field values ​​above the predicted value of 0.49 (such as Level 2, Age 41, Distance from Workplace to Home 1, and Job Satisfaction 4) have a negative impact on the predicted result, while field values ​​below the predicted value (such as Overtime, Supervisory Role, Work-Life Balance 1, Years of Employee 8, Stock Option Level 0, Relationship Satisfaction 1, Environment Satisfaction 2, and Employee Number 1) have a positive impact on the predicted result. The numerical value to the left of the corresponding field value indicates the degree of influence on the predicted result. Figure 8 It is a grayscale image without color shown for the purpose of application documents. Figure 8 Different colors can be used to indicate whether the field value has a positive or negative impact on the prediction result. For example, a predicted value above 0.49 can be represented by color A (e.g., blue), and a predicted value below 0.49 can be represented by color B (e.g., red).

[0191] As an example, in response to a user adjusting the field values ​​of one or more fields in a piece of data, a prediction result of the machine learning model for the adjusted piece of data may be output, and second explanatory information may be output indicating the importance of one or more of at least some of the other fields in the adjusted piece of data to the prediction result. For details about the second explanatory information and how it is displayed, please refer to the relevant description above and will not be repeated here.

[0192] As an example, a machine learning model can also be used to output changes in the field values ​​of at least some other fields in a piece of data based on the user's desired prediction result for a target field in that piece of data. For example, when using a machine learning model to predict the employee's turnover probability based on their information, if the expected turnover probability decreases by 20%, the model can output changes in field values ​​such as the required salary increase, the required overtime reduction, and the required stock option increase.

[0193] The present invention can also present the prediction results obtained by the machine learning model for multiple data in a two-dimensional coordinate system. There are multiple coordinate points in the two-dimensional coordinate system, and each coordinate point corresponds to a piece of data. The display characteristics of the coordinate points are used to characterize the prediction results of the data. The distance between two coordinate points in the two-dimensional space is positively correlated with the distance between the two data corresponding to the two coordinate points in the multidimensional space. Among them, each piece of data has multiple dimensions, and each dimension corresponds to a field in at least part of the other fields. Thus, through dimensionality reduction, multiple pieces of data can be displayed in a plane (that is, a two-dimensional coordinate system), and the distance relationship between different data is retained while reducing the dimensionality, that is, the distance between two pieces of data in the multidimensional space is close, and the distance on the two-dimensional plane after dimensionality reduction is also close. This can achieve clustering of data that are predicted to be a certain classification due to similar reasons, so that users can clearly understand the classification characteristics of the data set. The specific implementation method of dimensionality reduction will not be repeated here.

[0194] As an example, the score obtained by the machine learning model for predicting a single piece of data can be assigned to each field in at least some other fields according to the distribution method of Shapley values ​​to obtain the score of the field under the single piece of data. The score of the field is the value of the data under the corresponding dimension, wherein the position of each piece of data in the multidimensional space is determined based on the values ​​of multiple dimensions of the data.

[0195] like Figure 9 Taking the two-dimensional coordinate system shown in the lower half as an example, the coordinate system can include multiple coordinate points. Figure 9 It is a grayscale image without color shown for the purpose of application documents. Figure 9 Coordinate points can be colored, and the color of the coordinate point can be used to represent the prediction result obtained by the machine learning model for the corresponding data. For example, when predicting whether to resign, color A (for example, blue) can be used to indicate a no prediction result (i.e., no resignation), and color B (for example, red) can be used to indicate a yes prediction result (i.e., resignation). This allows users to intuitively understand the distribution of prediction results (e.g., resignation status) for multiple data points.

[0196] As an example, in response to a user's selection operation for one or more coordinate points in a two-dimensional coordinate system, a predetermined number of coordinate points with the same prediction result as the selected coordinate point are selected from the vicinity of the selected coordinate point to obtain a plurality of clustered coordinate points; based on the order of the scores of the fields, one or more key fields are extracted from the plurality of data corresponding to the plurality of clustered coordinate points to obtain a key field group; and the key field group is output. The score of the field mentioned here may refer to the sum of the scores of the field under the data corresponding to the plurality of clustered coordinate points. For the calculation method of the score of the field under a single piece of data, please refer to the relevant description above and will not be repeated here. The output key field group may include the extracted key fields and the statistical results of the field values ​​of the key fields under the corresponding data (such as the average value).

[0197] like Figure 9 As shown, in response to a user selecting a coordinate point in a two-dimensional coordinate system predicted to be a departure point, two groups of typical features can be output. Typical features are similar features shared by employees predicted to be departures. Group 1 can include fields such as stock option level, number of companies worked for, and work environment satisfaction, as well as statistical results of their field values. Group 2 can include fields such as whether overtime was worked, monthly income, and stock option level, as well as statistical results of their field values.

[0198] So far combined Figures 2 to 9 The present invention further describes a method for assisting users in exploring a dataset. This method can be implemented as a data analysis platform, allowing users to import and explore datasets using the platform. The data analysis platform can be designed as a webpage program that can be accessed through a browser, or as an application that can be installed and run on electronic devices such as mobile phones and iPads.

[0199] The method of assisting a user in exploring a data set of the present invention can also be implemented as a plug-in that can be run in an application program capable of opening a data table (e.g., spreadsheet software Excel). That is, in response to a user opening a data table in an application program, the plug-in for implementing the method of assisting a user in exploring a data set of the present invention can be run for the data set in the data table.

[0200] The following takes the present invention implemented as a plug-in running in an application as an example to further illustrate the process of the method for assisting users in exploring a data table.

[0201] The present invention also provides a method for assisting a user in exploring a data table, comprising: in response to a user opening a data table in an application, running a plug-in in the application to perform the following steps S1 to S3. The application refers to an application capable of opening a data table, such as, but not limited to, Excel or other spreadsheet software.

[0202] Step S1: Displaying an exploration area in a predetermined area of ​​a data table. The predetermined area may be at least one of the left side, right side, top, and bottom of the data table. The data table may be an Excel spreadsheet.

[0203] Step S2: In response to a user's prediction request for a target field in the data table, the machine learning model is automatically trained using the field values ​​corresponding to at least a portion of other fields in the single data entry as input and the field value corresponding to the target field in the single data entry as output. The automatic training of the machine learning model may include automatically searching within a minimal hyperparameter space based on empirically preset hyperparameters to train the machine learning model.

[0204] Step S3: Output first explanatory information in the exploration area that characterizes the importance of at least some of the other fields to the target field. The first explanatory information can be used to characterize the importance of all fields to the machine learning model's prediction of the target field, or it can be used to characterize some fields that are more important to the machine learning model's prediction of the target field. The calculation of field importance can be found in the relevant description above and will not be repeated here.

[0205] In the process of calculating the importance, the first explanatory information can be displayed in a gradient form. For example, since it takes a long time to calculate the importance, the first explanatory information can be roughly displayed in a relatively vague state first. As the calculation is completed more and more, the displayed first explanatory information gradually becomes clearer. Regarding the display form of the first explanatory information, as an example, different display characteristics can be given to the field according to the importance of the calculated field. For example, the greater the importance of the field, the darker the color. Optionally, different colors can be used to identify whether the field has a positive or negative effect on the target field. For example, red can be used to represent a positive effect, and blue can be used to represent a negative effect. That is, the greater the importance of the field to the target field and the positive effect it has, the redder its color is; the greater the importance of the field to the target field and the negative effect it has, the bluer its color is.

[0206] For example, the application is Excel and the data table is Excel. Figure 10As shown, the Excel toolbar can be provided with an "Automatic Modeling" tool. The user can first use the cursor to select the data column "Attrition" as the target field for desired prediction, and then click the cursor on the "Automatic Modeling" tool to start automatic training of the machine learning model.

[0207] like Figure 11 As shown in the figure, after the training is completed, the exploration area can be displayed on the right side of the Excel table, and the explanation information of the impact of the field value size of different fields on the target field can be output in the exploration area. Figure 11 For the explanation information on the right side of the output, please refer to the above combined Figure 6B The description is not repeated here.

[0208] like Figure 12 As shown, a two-dimensional coordinate graph can also be output in the exploration area to represent the impact of a single field on the target field. One coordinate axis in the two-dimensional coordinate graph is used to represent the field value corresponding to the field, and the other coordinate axis is used to represent the score of the field. The two-dimensional coordinate graph includes multiple coordinate points, each coordinate point corresponds to a piece of data, and the score of the field is obtained by assigning the score obtained by the machine learning model for a single piece of data to each field in at least part of the other fields based on the distribution method of the Shapley value. Optionally, the display characteristics of the coordinate point in the two-dimensional coordinate graph can also be used to represent the field value corresponding to another field in the data corresponding to the coordinate point. About Figure 12 The two-dimensional coordinate diagram output on the right side can be found in the above text. Figure 7 The description is not repeated here.

[0209] like Figure 13 As shown, in response to a user's prediction request for a piece of data in a data table, the prediction result of the machine learning model for the piece of data is output in the exploration area, and second explanatory information of the importance of one or more fields in at least part of the other fields in the piece of data to the prediction result is output in the exploration area.

[0210] The second explanatory information is used to represent the importance of each field in the data piece to the prediction result of the target field obtained by the machine learning model for the data piece. The importance here can also be calculated using the Shapley value distribution method. That is, based on the Shapley value distribution method, the score obtained by the machine learning model for the data piece can be distributed to each field in at least some other fields to obtain a score for each field in the data piece. The importance is positively correlated with the score.

[0211] about Figure 13 The second explanation information output on the right side can be found in the above combination Figure 8In response to a user's prediction request for a piece of data in a data table, specific display characteristics can be assigned to the field values ​​of different fields in the piece of data in the data table, so as to use the display characteristics of the field values ​​to represent the impact of the field values ​​on the prediction results. Figure 13 It is a grayscale image without color shown for the purpose of application documents. Figure 13 Color A (such as blue) can be used to represent the negative effect of field values ​​on the prediction results, and color B (such as red) can be used to represent the positive effect of field values ​​on the prediction results. The darker the color of the field value, the greater the impact on the prediction results.

[0212] For example, in response to a user adjusting the values ​​of one or more fields in a data item in a data table, the machine learning model outputs a prediction result for the adjusted data item; and second explanatory information is output in the exploration area, indicating the importance of one or more fields in at least some other fields in the adjusted data item to the prediction result. For details about the second explanatory information and how it is displayed, please refer to the relevant description above and will not be repeated here.

[0213] like Figure 14 As shown, a control for adjusting one or more field values ​​in a piece of data can be displayed to the user. The user can adjust one or more field values ​​through the control. After the adjustment is completed, the user can click the Predict button to output the prediction result and second explanation information of the machine learning model for the adjusted data.

[0214] As an example, a machine learning model can also be used to output changes in the field values ​​of at least some other fields in a data table based on the user's desired prediction result for a target field in that data item. For example, when using a machine learning model to predict the likelihood of an employee leaving based on a piece of employee information, if the expected likelihood of the employee leaving is reduced by 20%, the model can output changes in field values ​​such as the required salary increase, the required overtime hours reduction, and the required stock option increase.

[0215] As an example, in response to a user's selection operation on one or more data columns in a data table, first data analysis information can be output in the exploration area. The first data analysis information is used to characterize the data statistics of the field values ​​corresponding to the data columns selected by the user, and one or more second data analysis information can also be recommended to the user. The second data analysis information is data analysis information obtained by prediction based on the data columns selected by the user.

[0216] The second data analysis information is used to characterize the data statistics of the field values ​​corresponding to other data columns obtained through prediction, and / or the second data analysis information is used to characterize the data statistics of the field value combination corresponding to the data column combination composed of the data column selected by the user and the other data columns obtained through prediction.

[0217] In the present invention, one or more of a variety of methods such as, but not limited to, machine learning models, statistical correlations, business rules, etc. may be used to predict data analysis information suitable for recommendation to users.

[0218] Taking the prediction of data analysis information suitable for recommendation to users based on a machine learning model as an example, the probability that the user will select each other data column in the future can be predicted using a machine learning model based on the relevant information of the data column currently selected by the user, wherein the machine learning model is trained in the following manner: taking the relevant information of the previously selected data column and the relevant information of other data columns as input, and taking the probability of other data columns being selected as output; analyzing the data statistics of the field values ​​corresponding to the data columns whose probability values ​​are greater than a first predetermined threshold to obtain second data analysis information, or analyzing the data statistics of the field values ​​corresponding to the data columns whose probability values ​​are greater than the first predetermined threshold on the basis of the first data analysis information of the data column selected by the user to obtain second data analysis information.

[0219] Optionally, the user's acceptance feedback on one or more second data analysis information recommended to the user may be obtained, and the machine learning model may be updated based on the obtained acceptance feedback.

[0220] Taking data analysis information suitable for recommendation to users based on statistical correlation prediction as an example, data columns whose statistical correlation with the data column selected by the user is higher than a second predetermined threshold can be obtained according to the statistical correlation between different data columns; the data statistics of the field values ​​corresponding to the data columns whose statistical correlation is higher than the second predetermined threshold are analyzed to obtain the second data analysis information, or the data statistics of the field values ​​corresponding to the data columns whose statistical correlation is higher than the second predetermined threshold are analyzed on the basis of the first data analysis information of the data column selected by the user to obtain the second data analysis information.

[0221] As an example, the exploration area may include a first interface area and a second interface area. In response to a user's selection operation on one or more data columns in a data table, outputting first data analysis information in the exploration area includes: in response to the user selecting a specific data column with a cursor, displaying a chart of the first data analysis information in the first interface area, and recommending one or more second data analysis information to the user includes: while displaying a chart of the first data analysis information in the first interface area, displaying one or more charts of the second data analysis information in the second interface area.

[0222] Optionally, when a data table is opened in an application, data statistics of the field values ​​corresponding to each data column can be pre-calculated and the calculated data statistics can be cached in the memory, wherein the chart displaying the first data analysis information in the first interface area includes: displaying a chart of the first data analysis information generated based on the corresponding data statistics extracted from the memory in the first interface area.

[0223] As an example, in response to a user selecting a field value in a chart of the first data analysis information, one or more corresponding charts of the second data analysis information are automatically updated to the data statistics of the field value under the corresponding other field dimensions.

[0224] Optionally, in response to the user selecting the at least one chart of the second data analysis information displayed in the second interface area with a cursor, dragging it to the first interface area and releasing it, the at least one chart of the second data analysis information is displayed in the first interface area.

[0225] As an example, in response to a user's connection operation on two charts in the first interface area, one of the two charts is connected to the other chart; in response to a user's selection operation on a field value represented in the first chart in a connected state, the selected field value is highlighted relative to the unselected field value, and / or the second chart in a connected state is updated to represent the data statistics of the selected field value under the field dimension corresponding to the second chart.

[0226] As an example, in response to a user's connection operation on two charts in the first interface area, one of the two charts is connected to another chart, and the other chart is updated to represent the data statistics of the field values ​​represented by one chart under the field dimension corresponding to the other chart.

[0227] The method of assisting a user in exploring a data set of the present invention may also be implemented as a device for assisting a user in exploring a data set. Figure 15 The following is a block diagram of a device for assisting a user in exploring a data set according to an exemplary embodiment of the present invention. The functional units of the device for assisting a user in exploring a data set may be implemented by hardware, software, or a combination of hardware and software that implements the principles of the present invention. It will be understood by those skilled in the art that Figure 15 The functional units described can be combined or divided into sub-units to implement the principles of the above invention. Therefore, the description herein can support any possible combination, division, or further limitation of the functional units described herein.

[0228] The following briefly describes the functional units that a device for assisting users in exploring data sets may have and the operations that each functional unit may perform. For the details involved, please refer to the relevant description above and will not be repeated here.

[0229] See also Figure 15 The apparatus 200 for assisting a user in exploring a data set includes an output module 210 and a recommendation module. The data set includes multiple pieces of data, each piece of data including values ​​of one or more fields.

[0230] Output module 210 is configured to output first data analysis information in response to a user's field selection operation. The first data analysis information is configured to represent the statistical data of the field value corresponding to the user-selected field. Recommendation module 220 is configured to recommend one or more second data analysis information to the user. The second data analysis information is data analysis information predicted based on the user-selected field. For details about the second data analysis information, please refer to the relevant description above and will not be repeated here.

[0231] As an example, the recommendation module 220 may include, but is not limited to, one or any combination of a first recommendation module, a second recommendation module, and a third recommendation module. The first recommendation module is used to predict data analysis information suitable for recommendation to users based on a machine learning model, the second recommendation module is used to predict data analysis information suitable for recommendation to users based on statistical correlations, and the third recommendation module is used to predict data analysis information suitable for recommendation to users based on business rules. The prediction mechanisms of the first, second, and third recommendation modules can be found in the relevant description above and will not be repeated here.

[0232] The apparatus 200 for assisting a user in exploring a data set may further include a display module.

[0233] Example 1

[0234] In this embodiment, the display module is configured to display a first interface area and a second interface area. The first data analysis information is presented in the first interface area, and the second data analysis information is presented in the second interface area. In response to a user operation, the display module further presents the second data analysis information in the second interface area in the first interface area.

[0235] In response to the user's connection operation on two data analysis information in the first interface area, the display module can also connect one of the two data analysis information to another data analysis information, and the other data analysis information is updated to represent the data statistics of the field value represented by one data analysis information under the field dimension corresponding to the other data analysis information.

[0236] In response to the user's connection operation on two data analysis information in the first interface area, the display module connects one of the two data analysis information to the other data analysis information. In response to the user's selection operation on the field value represented by one of the two data analysis information in the connected state, the display module highlights the data statistics of the selected field value in the data analysis information relative to the data statistics of the unselected field values, and / or the display module updates the other data analysis information of the two data analysis information in the connected state to represent the data statistics of the selected field value under the field dimension corresponding to the other data analysis information.

[0237] Example 2

[0238] In this embodiment, the display module can be used to display an interface for exploring a data set, wherein the left area of ​​the interface is used to display the name icons corresponding to the various fields of the data set, the middle area of ​​the interface (equivalent to the first interface area) is used to display the charts of the first data analysis information, and the right area of ​​the interface (equivalent to the second interface area) is used to display the charts of the second data analysis information. When importing a data set, the display module displays the name icons corresponding to the various fields included in each data in the data set in the left area of ​​the interface, and the output module displays the charts of the first data analysis information in the middle area of ​​the interface in response to the user selecting the name icon of a specific field with the cursor, dragging it to the middle area of ​​the interface, and releasing it. At the same time as the charts of the first data analysis information are displayed in the middle area of ​​the interface, the recommendation module displays one or more charts of the second data analysis information in the right area of ​​the interface.

[0239] The device 200 for assisting users in exploring a data set may also include a first calculation module for pre-calculating the data statistics of the field values ​​corresponding to each field when importing a data set, and caching the calculated data statistics in the memory, and the output module displays a chart of the first data analysis information generated based on the corresponding data statistics extracted from the memory in the middle area of ​​the interface.

[0240] In response to the user selecting a field value in a chart of the first data analysis information, the display module automatically updates one or more corresponding charts of the second data analysis information to data statistics of the field value under corresponding other field dimensions.

[0241] The device 200 for assisting users in exploring a data set may also include a second calculation module, which is used to pre-calculate that when one or more frequently selected field values ​​in a chart of the first data analysis information are selected, the corresponding one or more charts of the second data analysis information are automatically updated to the data statistics of the frequently selected field values ​​under the corresponding other field dimensions, and cache the calculated data statistics in the memory, wherein the display module automatically updates the corresponding one or more charts of the second data analysis information to the corresponding data statistics cached in the memory.

[0242] In response to the user selecting the at least one second data analysis information chart displayed in the right area of ​​the interface with a cursor, dragging it to the middle area of ​​the interface and releasing it, the display module displays the at least one second data analysis information chart in the middle area of ​​the interface.

[0243] In response to the user's connection operation on two charts in the middle area of ​​the interface, the display module connects one of the two charts to the other chart. In response to the user's selection operation on the field value represented in the first chart in the connected state, the display module highlights the selected field value relative to the unselected field value, and / or the display module updates the second chart in the connected state to represent the data statistics of the selected field value under the field dimension corresponding to the second chart.

[0244] In response to the user's connection operation on the two charts in the middle area of ​​the interface, the display module connects one of the two charts to the other chart, and the display module updates the other chart to represent the data statistics of the field value represented by one chart under the field dimension corresponding to the other chart.

[0245] In Embodiment 1 and Embodiment 2, the apparatus 200 for assisting a user in exploring a data set may further include a training module and a presentation module.

[0246] In Example 1, the training module is configured to automatically train a machine learning model in response to a user's prediction request for a target field, using the field values ​​corresponding to at least a portion of other fields in a single piece of data as input and the field value corresponding to the target field in the single piece of data as output. The presentation module is configured to present first explanatory information representing the importance of one or more fields in at least a portion of the other fields to the machine learning model's prediction of the target field.

[0247] In Example 2, the training module is used to automatically train the machine learning model using the field values ​​corresponding to at least a portion of other fields in a single piece of data as input and the field value corresponding to the target field in the single piece of data as output, in response to an operation of starting automatic training of the machine learning model executed in the right area of ​​the interface when the user selects the name icon of the target field in the left area of ​​the interface or selects the chart of the first data analysis information of the target field in the middle area of ​​the interface. The display module is used to display in the right area of ​​the interface first explanatory information that characterizes the importance of one or more fields in at least a portion of other fields to the machine learning model in predicting the target field.

[0248] In Examples 1 and 2, the training module can automatically search in a minimal hyperparameter space based on hyperparameters preset based on experience to train the machine learning model. The display module can display the first explanation information in a gradual manner during the calculation of the importance.

[0249] In Examples 1 and 2, the apparatus 200 for assisting a user in exploring a data set may further include a third calculation module and a determination module. The third calculation module is configured to assign, based on a Shapley value distribution method, a score obtained by the machine learning model for predicting a single piece of data to each of the at least some of the other fields, to obtain a score for the field under the single piece of data. The determination module is configured to determine, based on the sum of the scores of each field in the at least some of the other fields across multiple pieces of data, the importance of the field to the prediction of the target field by the machine learning model, wherein the importance is positively correlated with the sum of the scores.

[0250] In Examples 1 and 2, the apparatus 200 for assisting users in exploring a data set may further include a fourth calculation module. The fourth calculation module is configured to assign the score predicted by the machine learning model for a single piece of data to each of at least some of the other fields based on the Shapley value distribution method, to obtain the score of the field under the single piece of data. For a single field, the display module presents the score of the field under multiple pieces of data in a two-dimensional coordinate system, where one coordinate axis in the two-dimensional coordinate system is used to represent the field value corresponding to the field, and the other coordinate axis is used to represent the score of the field. The display module may also utilize the display characteristics of a coordinate point in the two-dimensional coordinate system to represent the field value corresponding to another field in the data corresponding to the coordinate point.

[0251] In Example 1 and Example 2, the display module is also used to present the prediction results obtained by the machine learning model for multiple data in a two-dimensional coordinate system. There are multiple coordinate points in the two-dimensional coordinate system, and each coordinate point corresponds to a piece of data. The display characteristics of the coordinate points are used to characterize the prediction results of the data. The distance between two coordinate points in the two-dimensional space is positively correlated with the distance between the two data corresponding to the two coordinate points in the multidimensional space.

[0252] As an example, each piece of data has multiple dimensions, and each dimension corresponds to a field in the at least some other fields. The device 200 for assisting users in exploring a data set may also include a sixth calculation module for assigning the score obtained by the machine learning model for predicting a single piece of data to each field in the at least some other fields according to the distribution method of the Shapley value, so as to obtain the score of the field under the single piece of data. The score of the field is the value of the data under the corresponding dimension, wherein the position of each piece of data in the multidimensional space is determined based on the values ​​of the multiple dimensions of the data.

[0253] Optionally, the apparatus 200 for assisting a user in exploring a dataset may further include a selection module and an extraction module. The selection module is configured to, in response to a user selecting one or more coordinate points in a two-dimensional coordinate system, select a predetermined number of coordinate points near the selected coordinate point that have the same prediction result as the selected coordinate point, thereby obtaining a plurality of clustered coordinate points. The extraction module is configured to extract one or more key fields from the plurality of data items corresponding to the plurality of clustered coordinate points based on the order of the field scores, thereby obtaining a key field group. The presentation module is further configured to output the key field group.

[0254] In Example 1 and Example 2, the display module is also used to output the prediction result of the machine learning model for the adjusted data in response to the user's adjustment operation on the field value of one or more fields in a certain data, and output second explanatory information on the importance of one or more fields in at least part of the other fields in the adjusted data to the prediction result.

[0255] In Example 1 and Example 2, the display module is also used to use a machine learning model to output changes in field values ​​of at least some other fields in a piece of data based on the user's expected prediction results for the target field of the data.

[0256] It should be understood that the specific implementation of the apparatus 200 for assisting a user in exploring a data set according to an exemplary embodiment of the present invention can be referred to in conjunction with Figures 1 to 9 The relevant specific implementation methods described are implemented and will not be repeated here.

[0257] The present invention can also be implemented as a device for assisting users in exploring data tables. The device may include a running module for running a plug-in for implementing the method of assisting users in exploring data sets of the present invention for a data set in the data table in response to a user opening the data table in an application.

[0258] The method of assisting a user in exploring a data table of the present invention may also be implemented as a device for assisting a user in exploring a data table. Figure 16The block diagram of the structure of the device for assisting a user in exploring a data table according to an exemplary embodiment of the present invention is shown. The functional units of the device for assisting a user in exploring a data table can be implemented by hardware, software, or a combination of hardware and software that implements the principles of the present invention. It will be understood by those skilled in the art that Figure 16 The functional units described can be combined or divided into sub-units to implement the principles of the above invention. Therefore, the description herein can support any possible combination, division, or further limitation of the functional units described herein.

[0259] The following briefly describes the functional units that a device for assisting users in exploring a data table may have and the operations that each functional unit may perform. For the details involved, please refer to the relevant description above and will not be repeated here.

[0260] See also Figure 16 The apparatus 300 for assisting a user in exploring a data table includes a display module 310 and a training module 320. The apparatus 300 for assisting a user in exploring a data table can be implemented as a plug-in installed in an application. The apparatus 300 for assisting a user in exploring a data table can be executed in response to a user opening a data table in the application.

[0261] The display module 310 is configured to display an exploration area in a predetermined area of ​​the data table in response to a user opening the data table in an application. The predetermined area may be at least one of the left side, right side, top, and bottom of the data table. The data table may be an Excel spreadsheet.

[0262] Training module 320 is configured to automatically train a machine learning model in response to a user's prediction request for a target field in a data table, using the field values ​​corresponding to at least some other fields in a single data entry as input and the field value corresponding to the target field in the single data entry as output. Training module 620 can automatically search within a minimal hyperparameter space based on empirically preset hyperparameters to train the machine learning model.

[0263] The display module 310 is further configured to output, in the exploration area, first explanatory information representing the importance of at least some of the other fields to the target field. The first explanatory information can be used to represent the importance of all fields to the machine learning model predicting the target field, or it can be used to represent some fields that are more important to the machine learning model predicting the target field. The calculation of the importance of the fields can be found in the relevant description above and will not be repeated here.

[0264] The display module can display the first explanatory information in a gradual manner during the process of calculating the importance. For example, since it takes a long time to calculate the importance, the first explanatory information can be roughly displayed in a relatively vague state first. As the calculation is completed more and more, the displayed first explanatory information gradually becomes clearer. Regarding the display form of the first explanatory information, as an example, different display characteristics can be given to the fields according to the importance of the fields calculated. For example, the greater the importance of the field, the darker the color. Optionally, different colors can be used to identify whether the field has a positive or negative effect on the target field. For example, red can be used to represent a positive effect and blue can be used to represent a negative effect. That is, the greater the importance of the field to the target field and the positive effect it has, the redder its color is; the greater the importance of the field to the target field and the negative effect it has, the bluer its color is.

[0265] As an example, the apparatus 300 for assisting a user in exploring a data table may further include a first calculation module and a determination module. The first calculation module is configured to assign, based on a Shapley value distribution method, a score obtained by the machine learning model for predicting a single piece of data to each of at least some other fields, thereby obtaining a score for the field in the single piece of data. The determination module is configured to determine, based on the sum of the scores of each of at least some other fields across multiple pieces of data, the importance of the field to the prediction of the target field by the machine learning model, wherein the importance is positively correlated with the sum of the scores.

[0266] As an example, the display module can also be used to output a two-dimensional coordinate graph in the exploration area for representing the impact of a single field on the target field, one coordinate axis in the two-dimensional coordinate graph is used to represent the field value corresponding to the field, and the other coordinate axis is used to represent the score of the field. The two-dimensional coordinate graph includes multiple coordinate points, each coordinate point corresponds to a piece of data, and the score of the field is obtained by assigning the score obtained by the machine learning model for the single piece of data to each field in the at least part of the other fields according to the distribution method of the Shapley value. Optionally, the display module can also be used to use the display characteristics of the coordinate point in the two-dimensional coordinate graph to represent the field value corresponding to another field in the data corresponding to the coordinate point.

[0267] As an example, in response to a user's prediction request for a certain data in a data table, the display module outputs the prediction result of the machine learning model for the data in the exploration area, and outputs second explanatory information on the importance of one or more fields of at least some other fields in the data to the prediction result in the exploration area. The device 300 for assisting users in exploring a data table may also include a second calculation module for assigning the score obtained by the machine learning model for the data to each of the at least some other fields according to the distribution method of the Shapley value, so as to obtain the score of each field under the data, and the importance of the field to the prediction result is positively correlated with the score of the field. In response to the user's adjustment operation on the field value of one or more fields in a certain data in the data table, the display module may also output the prediction result of the machine learning model for the adjusted data, and output second explanatory information on the importance of one or more fields of at least some other fields in the data to the prediction result in the exploration area.

[0268] As an example, the display module can also use the machine learning model to output the changes in the field values ​​of at least some other fields in a certain data in the data table based on the user's expected prediction results for the target field of the data.

[0269] The apparatus 200 for assisting a user in exploring a data table may further include an output module 210 and a recommendation module 220. The output module 210 is configured to output first data analysis information in an exploration area in response to a user selecting one or more data columns in the data table. The first data analysis information is configured to represent data statistics of field values ​​corresponding to the data columns selected by the user.

[0270] The recommendation module 220 is used to recommend one or more second data analysis information to the user. The second data analysis information is data analysis information predicted based on the data column selected by the user. For the second data analysis information, please refer to the relevant description above and will not be repeated here.

[0271] Recommendation module 220 may include, but is not limited to, one or any combination of a first recommendation module, a second recommendation module, and a third recommendation module. The first recommendation module is used to predict data analysis information suitable for recommendation to users based on a machine learning model, the second recommendation module is used to predict data analysis information suitable for recommendation to users based on statistical correlations, and the third recommendation module is used to predict data analysis information suitable for recommendation to users based on business rules. The prediction mechanisms of the first, second, and third recommendation modules can be found in the relevant description above and will not be repeated here.

[0272] The exploration area may include a first interface area and a second interface area. In response to the user selecting a specific data column with a cursor, the output module displays a chart of the first data analysis information in the first interface area. While displaying the chart of the first data analysis information in the first interface area, the recommendation module displays one or more charts of the second data analysis information in the second interface area.

[0273] The device 200 for assisting users in exploring data tables may also include a third calculation module, which is used to pre-calculate data statistics of field values ​​corresponding to each data column when opening the data table in an application, and cache the calculated data statistics in the memory, wherein the output module displays a chart of the first data analysis information generated based on the corresponding data statistics extracted from the memory in the first interface area.

[0274] In response to a user selecting a certain field value in a chart of the first data analysis information, the recommendation module automatically updates the corresponding one or more charts of the second data analysis information to the data statistics of the certain field value under the corresponding other field dimensions. The device 300 for assisting users in exploring data tables may also include a fourth calculation module for pre-calculating that when one or more frequently selected field values ​​in the chart of the first data analysis information are selected, the corresponding one or more charts of the second data analysis information are automatically updated to the data statistics of the frequently selected field values ​​under the corresponding other field dimensions, and caching the calculated data statistics in the memory. The recommendation module automatically updates the corresponding one or more charts of the second data analysis information to the corresponding data statistics cached in the memory.

[0275] In response to the user selecting the at least one second data analysis information chart displayed in the second interface area with a cursor, dragging the cursor to the first interface area and releasing the cursor, the display module displays the at least one second data analysis information chart in the first interface area.

[0276] In response to a user's connection operation on two charts in the first interface area, the display module connects one of the two charts to the other chart. In response to a user's selection operation on a field value represented in the first chart in a connected state, the display module highlights the selected field value relative to the unselected field value, and / or the display module updates the second chart in a connected state to represent the data statistics of the selected field value under the field dimension corresponding to the second chart.

[0277] In response to the user's connection operation on two charts in the first interface area, the display module connects one of the two charts to another chart, and the display module updates the other chart to represent the data statistics of the field value represented by one chart under the field dimension corresponding to the other chart.

[0278] It should be understood that the specific implementation of the device 300 for assisting a user in exploring a data table according to an exemplary embodiment of the present invention can be referred to in conjunction with the above. Figures 10 to 14 This is achieved by referring to the relevant description of the method for assisting users in exploring the data table, which will not be repeated here.

[0279] Reference above Figures 1 to 16 A method and apparatus for assisting a user in exploring a data set or data table according to an exemplary embodiment of the present invention are described. It should be understood that the above method may be implemented by a program recorded on a computer-readable medium. For example, according to an exemplary embodiment of the present invention, a computer-readable storage medium storing instructions may be provided, wherein the computer-readable medium records instructions for executing the method for assisting a user in exploring a data set according to the present invention (e.g., Figure 1 as shown) or a computer program that assists a user in exploring a data table.

[0280] The computer program in the computer readable medium can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. It should be noted that the computer program can be used to execute other Figure 1 In addition to the steps shown, it can also be used to perform additional steps other than the above steps or perform more specific processing when performing the above steps. The contents of these additional steps and further processing have been referred to. Figure 1 To avoid repetition, it will not be described again here.

[0281] It should be noted that the device for assisting users in exploring data sets and the device for assisting users in exploring data tables according to exemplary embodiments of the present invention can completely rely on the operation of computer programs to realize corresponding functions, that is, each device corresponds to each step in the functional architecture of the computer program, so that the entire device is called through a special software package (for example, lib library) to realize the corresponding function.

[0282] on the other hand, Figure 15 、 Figure 16 The various devices shown may also be implemented by hardware, software, firmware, middleware, microcode, or any combination thereof. When implemented by software, firmware, middleware, or microcode, the program code or code segments for performing the corresponding operations may be stored in a computer-readable medium such as a storage medium, so that a processor can perform the corresponding operations by reading and running the corresponding program code or code segments.

[0283] For example, an exemplary embodiment of the present invention can also be implemented as a computing device, which includes a storage component and a processor, wherein a set of computer-executable instructions is stored in the storage component, and when the set of computer-executable instructions is executed by the processor, a method for assisting a user in exploring a data set or a method for assisting a user in exploring a data table is executed.

[0284] Specifically, the computing device may be deployed in a server or client, or may be deployed on a node device in a distributed network environment. In addition, the computing device may be a PC, tablet device, personal digital assistant, smartphone, web application, or other device capable of executing the above-mentioned instruction set.

[0285] Here, the computing device is not necessarily a single computing device, but may be any collection of devices or circuits capable of executing the above instructions (or instruction sets) individually or in combination. The computing device may also be part of an integrated control system or system manager, or may be configured as a portable electronic device that is interconnected with a local or remote (e.g., via wireless transmission) interface.

[0286] In the computing device, the processor may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0287] According to the exemplary embodiment of the present invention, some operations described in the method of assisting users in exploring a data set or the method of assisting users in exploring a data table can be implemented by software, some operations can be implemented by hardware, and in addition, these operations can also be implemented by a combination of software and hardware.

[0288] The processor may execute instructions or codes stored in one of the memory components, wherein the memory component may also store data. Instructions and data may also be sent and received over a network via a network interface device, wherein the network interface device may employ any known transmission protocol.

[0289] The storage component can be integrated with the processor, for example, by placing RAM or flash memory within an integrated circuit microprocessor or the like. Furthermore, the storage component can include a separate device, such as an external disk drive, a storage array, or any other storage device usable by a database system. The storage component and the processor can be operatively coupled or can communicate with each other, for example, via an I / O port, a network connection, or the like, such that the processor can access files stored in the storage component.

[0290] In addition, the computing device may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, mouse, touch input device, etc.) All components of the computing device may be connected to each other via a bus and / or a network.

[0291] The operations involved in the method for assisting a user in exploring a data set or a method for assisting a user in exploring a data table according to an exemplary embodiment of the present invention can be described as various interconnected or coupled functional blocks or functional diagrams. However, these functional blocks or functional diagrams can be equally integrated into a single logical device or operate according to non-precise boundaries.

[0292] For example, as described above, the device for assisting users in exploring a data set or the device for assisting users in exploring a data table according to an exemplary embodiment of the present invention may include a storage component and a processor, wherein a set of computer-executable instructions is stored in the storage component, and when the set of computer-executable instructions is executed by the processor, the method for assisting users in exploring a data set or the method for assisting users in exploring a data table mentioned above is executed.

[0293] While various exemplary embodiments of the present invention have been described above, it should be understood that the foregoing description is merely illustrative and not exhaustive, and the present invention is not limited to the disclosed exemplary embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for assisting a user in exploring a data set, wherein the data set includes multiple pieces of data, each piece of data includes values ​​of one or more fields, and the data set includes an employee information table, the method comprising: In response to a field selection operation by a user, outputting first data analysis information, where the first data analysis information is used to represent data statistics of a field value corresponding to the field selected by the user; as well as Recommending one or more second data analysis information to the user, where the second data analysis information is data analysis information obtained by prediction based on the field selected by the user, and the second data analysis information is used to represent data statistics of field values ​​corresponding to other fields obtained through prediction, and / or the second data analysis information is used to represent data statistics of field value combinations corresponding to field combinations consisting of the field selected by the user and other fields obtained through prediction; In response to a user's prediction request for a target field, the machine learning model is automatically trained using the field values ​​corresponding to at least some other fields in the single data piece as input and the field value corresponding to the target field in the single data piece as output; Displaying first explanation information for characterizing the importance of one or more fields in the at least part of the other fields to the prediction of the target field by the machine learning model; or, An interface for exploring a data set is displayed, wherein the left area of ​​the interface is used to display name icons corresponding to various fields of the data set, the middle area of ​​the interface is used to display a chart of first data analysis information, and the right area of ​​the interface is used to display a chart of second data analysis information; when a user selects the name icon of a target field in the left area of ​​the interface or selects a chart of the first data analysis information of a target field in the middle area of ​​the interface, in response to an operation of starting automatic training of a machine learning model performed in the right area of ​​the interface, the machine learning model is automatically trained using field values ​​corresponding to at least part of other fields in a single piece of data as input and field values ​​corresponding to the target field in the single piece of data as output; first explanatory information for characterizing the importance of one or more fields of the at least part of other fields to the prediction of the target field by the machine learning model is displayed in the right area of ​​the interface; Among them, according to the distribution method of Shapley value, the score obtained by the machine learning model for predicting a single piece of data is distributed to each field in the at least part of the other fields to obtain the score of the field under the single piece of data; according to the sum of the scores of each field in the at least part of the other fields under multiple pieces of data, the importance of the field to the prediction of the target field by the machine learning model is determined, wherein the importance is positively correlated with the sum of the scores.

2. The method according to claim 1, wherein The step of recommending one or more second data analysis information to the user includes: Predicting data analysis information suitable for recommendations to users based on machine learning models; and / or Predicting data analysis information suitable for recommendation to users based on statistical relevance; and / or Predict data analysis information suitable for recommendation to users based on business rules.

3. The method according to claim 2, wherein: The steps for predicting data analysis information suitable for recommendation to users based on machine learning models include: Based on the information about the field currently selected by the user, using the machine learning model to predict the probability that the user will subsequently select each other field, wherein the machine learning model is trained in the following manner: taking the information about the previously selected field and the information about the other fields as input and outputting the probability of the other fields being selected; Analyze the data statistics of the field values ​​corresponding to the fields whose probability values ​​are greater than the first predetermined threshold to obtain the second data analysis information, or analyze the data statistics of the field values ​​corresponding to the fields whose probability values ​​are greater than the first predetermined threshold based on the first data analysis information of the field selected by the user to obtain the second data analysis information.

4. The method according to claim 3, further comprising: Obtain user acceptance feedback on one or more second data analysis information recommended to the user, and update the machine learning model based on the obtained acceptance feedback.

5. The method according to claim 2, wherein: The steps of predicting data analysis information suitable for recommendation to users based on statistical correlation include: According to the statistical correlations between different fields, obtaining a field whose statistical correlation with the field selected by the user is higher than a second predetermined threshold; Analyze the data statistics of the field values ​​corresponding to the fields whose statistical correlation is higher than the second predetermined threshold to obtain the second data analysis information, or analyze the data statistics of the field values ​​corresponding to the fields whose statistical correlation is higher than the second predetermined threshold based on the first data analysis information of the field selected by the user to obtain the second data analysis information.

6. The method according to claim 1, wherein The first data analysis information is presented in the first interface area, and the second data analysis information is presented in the second interface area. The method further includes: presenting the second data analysis information in the second interface area in the first interface area in response to a user operation.

7. The method according to claim 6, further comprising: In response to a user's connection operation on two pieces of data analysis information in the first interface area, connecting one piece of the two pieces of data analysis information to the other piece of data analysis information; The other data analysis information is updated to represent the data statistics of the field value represented by the one data analysis information under the field dimension corresponding to the other data analysis information.

8. The method according to claim 6, further comprising: In response to a user's connection operation on two pieces of data analysis information in the first interface area, connecting one piece of the two pieces of data analysis information to another piece of data analysis information; In response to a user's selection operation on a field value represented by one of the two data analysis information in a connected state, the data statistics of the selected field value in the data analysis information are highlighted relative to the data statistics of the unselected field values, and / or the other of the two data analysis information in a connected state is updated to represent the data statistics of the selected field value under the field dimension corresponding to the other data analysis information.

9. The method according to claim 1, further comprising: When importing a data set, the name icons corresponding to each field in each data set are displayed in the left area of ​​the interface. The outputting of the first data analysis information in response to the user's field selection operation includes: displaying a chart of the first data analysis information in the middle area of ​​the interface in response to the user selecting the name icon of a specific field with a cursor, dragging the cursor to the middle area of ​​the interface, and releasing the cursor; and The recommending one or more second data analysis information to the user includes: displaying a chart of the first data analysis information in the middle area of ​​the interface, and displaying one or more charts of the second data analysis information in the right area of ​​the interface.

10. The method according to claim 9, further comprising: When importing a data set, pre-calculate the data statistics of the field values ​​corresponding to each field and cache the calculated data statistics in memory. The displaying of the chart of the first data analysis information in the middle area of ​​the interface includes: displaying a chart of the first data analysis information generated based on corresponding data statistics extracted from the memory in the middle area of ​​the interface.

11. The method according to claim 9, further comprising: In response to a user selecting a field value in a chart of the first data analysis information, one or more corresponding charts of the second data analysis information are automatically updated to data statistics of the field value under corresponding other field dimensions.

12. The method according to claim 11, further comprising: When one or more frequently selected field values ​​pre-calculated in a chart of the first data analysis information are selected, the corresponding one or more charts of the second data analysis information are automatically updated to the data statistics of the frequently selected field values ​​under the corresponding other field dimensions, and the calculated data statistics are cached in the memory. Among them, the charts of the corresponding one or more second data analysis information are automatically updated to the data statistics of the certain field value under the corresponding other field dimensions, including: the charts of the corresponding one or more second data analysis information are automatically updated to the corresponding data statistics cached in the memory.

13. The method according to claim 9, further comprising: In response to the user selecting the at least one second data analysis information chart displayed in the right area of ​​the interface with a cursor, dragging it to the middle area of ​​the interface and releasing it, the at least one second data analysis information chart is displayed in the middle area of ​​the interface.

14. The method according to claim 9, further comprising: In response to a user's connection operation on two charts in the middle area of ​​the interface, connecting one of the two charts to the other chart; In response to a user's selection operation on a field value represented in a first connected chart, the selected field value is highlighted relative to unselected field values, and / or the second connected chart is updated to represent the data statistics of the selected field value under the field dimension corresponding to the second chart.

15. The method according to claim 9, further comprising: In response to a user's connection operation on two charts in the middle area of ​​the interface, connecting one of the two charts to the other chart; The other chart is updated to represent data statistics of the field value represented by the one chart under the field dimension corresponding to the other chart.

16. The method according to claim 1 or 9, The automatic training of the machine learning model includes: Automatically search the hyperparameter space to train machine learning models based on empirically preset hyperparameters.

17. The method according to claim 1 or 9, wherein Displaying the first explanatory information used to characterize the importance of one or more fields in at least part of the other fields to the prediction of the target field by the machine learning model includes: displaying the first explanatory information in a gradient form during the process of calculating the importance.

18. The method according to claim 1 or 9, further comprising: Allocating, according to a Shapley value allocation method, a score obtained by the machine learning model for predicting a single piece of data to each field in the at least some of the other fields, to obtain a score for the field under the single piece of data; For a single field, the scores of the field under multiple data are presented in a two-dimensional coordinate system. One coordinate axis in the two-dimensional coordinate system is used to represent the field value corresponding to the field, and the other coordinate axis is used to represent the score of the field.

19. The method according to claim 18, further comprising: The display characteristics of the coordinate point in the two-dimensional coordinate system are used to represent the field value corresponding to another field in the data corresponding to the coordinate point.

20. The method of claim 1, further comprising: In response to a user's selection operation on a piece of data, outputting a prediction result of the machine learning model on the piece of data; as well as Output second explanation information of the importance of one or more fields in at least part of the other fields in the data to the prediction result.

21. The method according to claim 9, further comprising: Provide users with a control to input a piece of data in the right area of ​​the interface, and receive the data input by the user; The right area of ​​the interface displays the prediction result of the machine learning model for the data, and displays second explanatory information on the importance of one or more fields in at least part of the other fields in the data to the prediction result.

22. The method according to claim 20 or 21, further comprising: According to the distribution method of Shapley values, the score obtained by the machine learning model for predicting the data is distributed to each field in the at least part of the other fields to obtain the score of each field under the data, and the importance is positively correlated with the score.

23. The method of claim 1, further comprising: The prediction results obtained by the machine learning model for multiple data are presented in a two-dimensional coordinate system. There are multiple coordinate points in the two-dimensional coordinate system, and each coordinate point corresponds to a piece of data. The display characteristics of the coordinate points are used to characterize the prediction results of the data. The distance between two coordinate points in the two-dimensional space is positively correlated with the distance between the two data corresponding to the two coordinate points in the multidimensional space.

24. The method according to claim 23, wherein Each piece of data has multiple dimensions, each dimension corresponds to one of the at least some of the other fields, and the method further includes: According to the distribution method of Shapley values, the score obtained by the machine learning model for predicting a single piece of data is distributed to each field in the at least part of the other fields to obtain the score of the field under the single piece of data. The score of the field is the value of the data under the corresponding dimension, wherein the position of each piece of data in the multidimensional space is determined based on the values ​​of multiple dimensions of the data.

25. The method according to claim 24, further comprising: In response to a user's selection operation on one or more coordinate points in the two-dimensional coordinate system, a predetermined number of coordinate points having the same prediction result as the selected coordinate point are selected from the vicinity of the selected coordinate point to obtain a plurality of clustered coordinate points; Extracting one or more key fields from the plurality of data corresponding to the plurality of cluster coordinate points based on the order of the scores of the fields to obtain a key field group; Output the key field group.

26. The method according to claim 20 or 21, further comprising: In response to a user's adjustment operation on the field values ​​of one or more fields in a piece of data, outputting a prediction result of the machine learning model for the adjusted piece of data; as well as Output second explanation information of the importance of one or more fields in at least part of the other fields in the adjusted data to the prediction result.

27. The method according to claim 20 or 21, further comprising: According to the user's expected prediction result for the target field of a certain data, the machine learning model is used to output the changes in the field values ​​of at least some other fields in the data.

28. A method for assisting a user in exploring a data table, comprising: In response to a user opening a data table in an application, a plug-in for implementing the method according to any one of claims 1 to 27 is executed for a data set in the data table.

29. A method for assisting a user in exploring a data table, comprising: In response to a user opening a data table in the application, the plug-in in the application is run to perform the following steps: displaying an exploration area in a predetermined area of ​​a data table, the data table including an employee information table; In response to a user's prediction request for a target field in a data table, the machine learning model is automatically trained using the field values ​​corresponding to at least some other fields in the single data entry as input and the field value corresponding to the target field in the single data entry as output; outputting, in the exploration area, first explanation information for characterizing the importance of at least part of the other fields to the target field; Allocating, according to a Shapley value allocation method, a score obtained by the machine learning model for predicting a single piece of data to each field in the at least some of the other fields, to obtain a score for the field under the single piece of data; According to the sum of the scores of each field in the at least part of the other fields under multiple data, the importance of the field to the prediction of the target field by the machine learning model is determined, wherein the importance is positively correlated with the sum of the scores.

30. The method according to claim 29, wherein The predetermined area is at least one of a left side, a right side, an upper side, and a lower side of the data table.

31. The method according to claim 29, wherein The data table is an Excel table.

32. The method according to claim 29, The automatic training of the machine learning model includes: Automatically search the hyperparameter space to train machine learning models based on empirically preset hyperparameters.

33. The method of claim 29, wherein: Outputting first explanation information for representing the importance of at least part of the other fields to the target field in the exploration area includes: displaying the first explanation information in a gradual manner during calculation of the importance.

34. The method of claim 29, further comprising: A two-dimensional coordinate graph is output in the exploration area to represent the impact of a single field on the target field. One coordinate axis in the two-dimensional coordinate graph is used to represent the field value corresponding to the field, and the other coordinate axis is used to represent the score of the field. The two-dimensional coordinate graph includes multiple coordinate points, each coordinate point corresponds to a piece of data, and the score of the field is obtained by assigning the score obtained by the machine learning model for the single piece of data to each field in the at least part of the other fields according to the distribution method of the Shapley value.

35. The method of claim 34, further comprising: The display characteristics of the coordinate point in the two-dimensional coordinate graph are used to represent the field value corresponding to another field in the data corresponding to the coordinate point.

36. The method of claim 29, further comprising: In response to a user's prediction request for a piece of data in the data table, outputting a prediction result of the machine learning model for the piece of data in the exploration area; as well as Second explanation information of the importance of one or more fields of at least part of the other fields in the data to the prediction result is output in the exploration area.

37. The method of claim 36, further comprising: According to the distribution method of Shapley values, the score obtained by the machine learning model for predicting the data is distributed to each field in the at least part of the other fields to obtain the score of each field under the data, and the importance of the field to the prediction result is positively correlated with the score of the field.

38. The method of claim 29, further comprising: In response to a user's adjustment operation on the field values ​​of one or more fields in a piece of data in a data table, outputting a prediction result of the machine learning model for the adjusted piece of data; as well as Outputting in the exploration area second explanation information of the importance of one or more fields of at least a portion of the other fields in the adjusted data to the prediction result.

39. The method of claim 29, further comprising: According to the user's expected prediction result for the target field of a certain data in the data table, the machine learning model is used to output the changes in the field values ​​of at least part of the other fields in the data.

40. The method of claim 29, further comprising: In response to a user selecting one or more data columns in a data table, outputting first data analysis information in an exploration area, the first data analysis information being used to represent data statistics of field values ​​corresponding to the data columns selected by the user; recommending one or more second data analysis information to the user, where the second data analysis information is data analysis information predicted based on the data column selected by the user; The second data analysis information is used to characterize the data statistics of the field values ​​corresponding to other data columns obtained through prediction, and / or The second data analysis information is used to represent data statistics of a combination of field values ​​corresponding to a combination of data columns formed by a data column selected by a user and other data columns obtained through prediction.

41. The method according to claim 40, wherein The step of recommending one or more second data analysis information to the user includes: Predicting data analysis information suitable for recommendations to users based on machine learning models; and / or Predicting data analysis information suitable for recommendation to users based on statistical relevance; and / or Predict data analysis information suitable for recommendation to users based on business rules.

42. The method according to claim 41, wherein The steps for predicting data analysis information suitable for recommendation to users based on machine learning models include: Based on the information about the data column currently selected by the user, using the machine learning model to predict the probability that the user will subsequently select each other data column, wherein the machine learning model is trained in the following manner: taking the information about the previously selected data column and the information about the other data columns as input and outputting the probability of the other data columns being selected; Analyze the data statistics of the field values ​​corresponding to the data column whose probability value is greater than the first predetermined threshold to obtain the second data analysis information, or analyze the data statistics of the field values ​​corresponding to the data column whose probability value is greater than the first predetermined threshold based on the first data analysis information of the data column selected by the user to obtain the second data analysis information.

43. The method of claim 42, further comprising: Obtain user acceptance feedback on one or more second data analysis information recommended to the user, and update the machine learning model based on the obtained acceptance feedback.

44. The method of claim 41, wherein The steps of predicting data analysis information suitable for recommendation to users based on statistical correlation include: According to the statistical correlations between different data columns, obtaining a data column whose statistical correlation with the data column selected by the user is higher than a second predetermined threshold; Analyze the data statistics of the field values ​​corresponding to the data columns whose statistical correlation is higher than the second predetermined threshold to obtain the second data analysis information, or analyze the data statistics of the field values ​​corresponding to the data columns whose statistical correlation is higher than the second predetermined threshold based on the first data analysis information of the data columns selected by the user to obtain the second data analysis information.

45. The method of claim 40, wherein The exploration area includes a first interface area and a second interface area, Outputting the first data analysis information in the exploration area in response to the user selecting one or more data columns in the data table includes: displaying a chart of the first data analysis information in the first interface area in response to the user selecting a specific data column with a cursor, and The recommending one or more second data analysis information to the user includes: simultaneously displaying a chart of the first data analysis information in the first interface area, and displaying one or more charts of the second data analysis information in the second interface area.

46. ​​The method of claim 45, further comprising: When a data table is opened in an application, the data statistics of the field values ​​corresponding to each data column are pre-calculated and the calculated data statistics are cached in memory. The displaying of the chart of the first data analysis information in the first interface area includes: displaying a chart of the first data analysis information generated based on corresponding data statistics extracted from the memory in the first interface area.

47. The method of claim 45, further comprising: In response to a user selecting a field value in a chart of the first data analysis information, one or more corresponding charts of the second data analysis information are automatically updated to data statistics of the field value under corresponding other field dimensions.

48. The method of claim 47, further comprising: When one or more frequently selected field values ​​pre-calculated in a chart of the first data analysis information are selected, the corresponding one or more charts of the second data analysis information are automatically updated to the data statistics of the frequently selected field values ​​under the corresponding other field dimensions, and the calculated data statistics are cached in the memory. Among them, the charts of the corresponding one or more second data analysis information are automatically updated to the data statistics of the certain field value under the corresponding other field dimensions, including: the charts of the corresponding one or more second data analysis information are automatically updated to the corresponding data statistics cached in the memory.

49. The method of claim 45, further comprising: In response to the user selecting the at least one second data analysis information chart displayed in the second interface area with a cursor, dragging the cursor to the first interface area and releasing the cursor, the at least one second data analysis information chart is displayed in the first interface area.

50. The method of claim 45, further comprising: In response to a user's connection operation on two charts in the first interface area, connecting one of the two charts to the other chart; In response to a user's selection operation on a field value represented in a first connected chart, the selected field value is highlighted relative to unselected field values, and / or the second connected chart is updated to represent the data statistics of the selected field value under the field dimension corresponding to the second chart.

51. The method of claim 45, further comprising: In response to a user's connection operation on two charts in the first interface area, connecting one of the two charts to the other chart; The other chart is updated to represent data statistics of the field value represented by the one chart under the field dimension corresponding to the other chart.

52. A device for assisting a user in exploring a data set, the data set comprising a plurality of data items, each data item comprising values ​​of one or more fields, the data set comprising an employee information table, the device comprising: an output module, configured to output first data analysis information in response to a field selection operation by a user, wherein the first data analysis information is used to represent data statistics of a field value corresponding to the field selected by the user; as well as a recommendation module, configured to recommend one or more second data analysis information to the user, wherein the second data analysis information is data analysis information obtained by prediction based on a field selected by the user, and the second data analysis information is used to represent data statistics of field values ​​corresponding to other fields obtained through prediction, and / or the second data analysis information is used to represent data statistics of field value combinations corresponding to field combinations consisting of the field selected by the user and other fields obtained through prediction; A training module for automatically training a machine learning model in response to a user's prediction request for a target field, using field values ​​corresponding to at least a portion of other fields in a single piece of data as input and field values ​​corresponding to the target field in the single piece of data as output; a display module for displaying first explanatory information representing the importance of one or more fields in the at least a portion of other fields to the prediction of the target field by the machine learning model; or, a display module for displaying an interface for exploring a data set, wherein the left area of ​​the interface is used to display name icons corresponding to various fields of the data set, the middle area of ​​the interface is used to display charts of first data analysis information, and the right area of ​​the interface is used to display charts of second data analysis information; a training module for automatically training a machine learning model by using field values ​​corresponding to at least part of other fields in a single piece of data as input and field values ​​corresponding to the target field in a single piece of data as output in response to an operation of starting automatic training of a machine learning model executed in the right area of ​​the interface when a user selects the name icon of a target field in the left area of ​​the interface or selects a chart of the first data analysis information of a target field in the middle area of ​​the interface; a display module for displaying first explanatory information in the right area of ​​the interface for characterizing the importance of one or more fields of the at least part of other fields to the machine learning model in predicting the target field; Among them, according to the distribution method of Shapley value, the score obtained by the machine learning model for predicting a single piece of data is distributed to each field in the at least part of the other fields to obtain the score of the field under the single piece of data; according to the sum of the scores of each field in the at least part of the other fields under multiple pieces of data, the importance of the field to the prediction of the target field by the machine learning model is determined, wherein the importance is positively correlated with the sum of the scores.

53. The apparatus of claim 52, wherein: The recommendation module includes: A first recommendation module, configured to predict data analysis information suitable for recommendation to users based on a machine learning model; and / or A second recommendation module is configured to predict data analysis information suitable for recommendation to users based on statistical correlation; and / or The third recommendation module is used to predict data analysis information suitable for recommendation to users based on business rules.

54. The apparatus of claim 53, wherein: The first recommendation module: Based on the information about the field currently selected by the user, using the machine learning model to predict the probability that the user will subsequently select each other field, wherein the machine learning model is trained in the following manner: taking the information about the previously selected field and the information about the other fields as input and outputting the probability of the other fields being selected; Analyze the data statistics of the field values ​​corresponding to the fields whose probability values ​​are greater than the first predetermined threshold to obtain the second data analysis information, or analyze the data statistics of the field values ​​corresponding to the fields whose probability values ​​are greater than the first predetermined threshold based on the first data analysis information of the field selected by the user to obtain the second data analysis information.

55. The apparatus of claim 54, wherein: The first recommendation module is also used to obtain user acceptance feedback on one or more second data analysis information recommended to the user, and update the machine learning model based on the obtained acceptance feedback.

56. The apparatus of claim 53, wherein: The second recommended module: According to the statistical correlations between different fields, obtaining a field whose statistical correlation with the field selected by the user is higher than a second predetermined threshold; Analyze the data statistics of the field values ​​corresponding to the fields whose statistical correlation is higher than the second predetermined threshold to obtain the second data analysis information, or analyze the data statistics of the field values ​​corresponding to the fields whose statistical correlation is higher than the second predetermined threshold based on the first data analysis information of the field selected by the user to obtain the second data analysis information.

57. The apparatus of claim 52, further comprising: A display module is used to display a first interface area and a second interface area, wherein the first data analysis information is presented in the first interface area and the second data analysis information is presented in the second interface area. In response to a user operation, the display module presents the second data analysis information in the second interface area in the first interface area.

58. The apparatus of claim 57, wherein In response to the user's connection operation on two data analysis information in the first interface area, the display module connects one of the two data analysis information to another data analysis information, and the other data analysis information is updated to represent the data statistics of the field value represented by the one data analysis information under the field dimension corresponding to the other data analysis information.

59. The apparatus of claim 57, wherein In response to a user's operation of connecting two pieces of data analysis information in the first interface area, the display module connects one piece of the two pieces of data analysis information to the other piece of data analysis information. In response to a user's selection operation on a field value represented by one of the two data analysis information in a connected state, the display module highlights the data statistics of the selected field value in the data analysis information relative to the data statistics of the unselected field values, and / or the display module updates the other of the two data analysis information in a connected state to represent the data statistics of the selected field value under the field dimension corresponding to the other data analysis information.

60. The apparatus of claim 52, further comprising: When importing a data set, the display module displays the name icons corresponding to each field included in each data set in the left area of ​​the interface. The output module displays a chart of the first data analysis information in the middle area of ​​the interface in response to the user selecting the name icon of a specific field with a cursor, dragging the cursor to the middle area of ​​the interface and releasing the cursor, and While the chart of the first data analysis information is displayed in the middle area of ​​the interface, the recommendation module displays one or more charts of the second data analysis information in the right area of ​​the interface.

61. The apparatus of claim 60, further comprising: The first calculation module is used to pre-calculate the data statistics of the field values ​​corresponding to each field when importing the data set, and cache the calculated data statistics in the memory. The output module displays a chart of first data analysis information generated based on corresponding data statistics extracted from the memory in the middle area of ​​the interface.

62. The apparatus of claim 60, wherein: In response to a user selecting a field value in a chart of the first data analysis information, the display module automatically updates one or more corresponding charts of the second data analysis information to data statistics of the field value under corresponding other field dimensions.

63. The apparatus of claim 62, further comprising: The second calculation module is used to pre-calculate that when one or more frequently selected field values ​​in the chart of the first data analysis information are selected, the corresponding one or more charts of the second data analysis information are automatically updated to the data statistics of the frequently selected field values ​​under the corresponding other field dimensions, and cache the calculated data statistics in the memory. The display module automatically updates the corresponding charts of the one or more second data analysis information to corresponding data statistics cached in the memory.

64. The apparatus of claim 60, wherein: In response to the user selecting the at least one second data analysis information chart displayed in the right area of ​​the interface with a cursor, dragging it to the middle area of ​​the interface and releasing it, the display module displays the at least one second data analysis information chart in the middle area of ​​the interface.

65. The apparatus of claim 60, wherein In response to a user's connection operation on two charts in the middle area of ​​the interface, the display module connects one of the two charts to the other chart. In response to a user's selection operation on a field value represented in a first chart in a connected state, the display module highlights the selected field value relative to unselected field values, and / or the display module updates a second chart in a connected state to represent the data statistics of the selected field value under the field dimension corresponding to the second chart.

66. The apparatus of claim 60, wherein In response to a user's connection operation on two charts in the middle area of ​​the interface, the display module connects one of the two charts to the other chart. The display module updates the other chart to represent data statistics of the field value represented by the one chart under the field dimension corresponding to the other chart.

67. The apparatus of claim 52 or 60, wherein: The training module automatically searches in the hyperparameter space based on hyperparameters preset according to experience to train the machine learning model.

68. The apparatus of claim 52 or 60, wherein: The display module displays the first explanation information in a gradual manner during the process of calculating the importance.

69. The apparatus of claim 52 or 60, further comprising: a fourth calculation module, configured to assign, based on a Shapley value allocation method, a score obtained by the machine learning model for predicting a single piece of data to each of the at least some of the other fields, to obtain a score for the field under the single piece of data; For a single field, the display module presents the scores of the field under multiple data in a two-dimensional coordinate system, where one coordinate axis is used to represent the field value corresponding to the field, and the other coordinate axis is used to represent the score of the field.

70. The apparatus of claim 69, wherein The display module further uses the display characteristics of the coordinate point in the two-dimensional coordinate system to represent the field value corresponding to another field in the data corresponding to the coordinate point.

71. The apparatus of claim 52, wherein: In response to the user's selection operation on a certain piece of data, the display module also outputs the prediction result of the machine learning model for the data, and outputs second explanatory information on the importance of one or more fields of at least part of the other fields in the data to the prediction result.

72. The apparatus of claim 60, wherein The display module is also used to provide a control for the user to input a certain piece of data in the right area of ​​the interface, and receive the data input by the user through the control. The display module is also used to display the prediction results of the machine learning model for the data in the right area of ​​the interface, and to display second explanatory information on the importance of one or more fields of at least part of the other fields in the data to the prediction results.

73. The apparatus according to claim 71 or 72, further comprising: The fifth calculation module is used to assign the score obtained by the machine learning model for predicting the data to each field in the at least part of the other fields according to the distribution method of the Shapley value, so as to obtain the score of each field under the data, and the importance is positively correlated with the score.

74. The apparatus of claim 52, wherein: The display module is also used to present the prediction results obtained by the machine learning model for multiple data in a two-dimensional coordinate system. There are multiple coordinate points in the two-dimensional coordinate system, and each coordinate point corresponds to a piece of data. The display characteristics of the coordinate points are used to characterize the prediction results of the data. The distance between two coordinate points in the two-dimensional space is positively correlated with the distance between the two data corresponding to the two coordinate points in the multidimensional space.

75. The apparatus of claim 74, wherein: Each piece of data has multiple dimensions, each dimension corresponds to one of the at least part of the other fields, and the apparatus further includes: The sixth calculation module is used to assign the score obtained by the machine learning model for predicting a single piece of data to each field of the at least part of the other fields according to the distribution method of the Shapley value, so as to obtain the score of the field under the single piece of data. The score of the field is the value of the data under the corresponding dimension, wherein the position of each piece of data in the multidimensional space is determined based on the values ​​of multiple dimensions of the data.

76. The apparatus of claim 75, further comprising: a selection module, configured to select, in response to a user's selection operation on one or more coordinate points in the two-dimensional coordinate system, a predetermined number of coordinate points having the same prediction result as the selected coordinate point from the vicinity of the selected coordinate point, so as to obtain a plurality of clustered coordinate points; An extraction module is used to extract one or more key fields from the multiple data corresponding to the multiple cluster coordinate points based on the order of the field scores to obtain a key field group. The display module is further configured to output the key field group.

77. The apparatus of claim 52 or 60, wherein: The display module is also used to output the prediction result of the machine learning model for the adjusted data in response to the user's adjustment operation on the field value of one or more fields in a certain data, and output second explanatory information on the importance of one or more fields in at least part of the other fields in the adjusted data to the prediction result.

78. The apparatus of claim 52 or 60, wherein: The display module is also used to use the machine learning model to output the changes in the field values ​​of at least some other fields in a certain piece of data based on the user's expected prediction results for the target field of the data.

79. A device for assisting a user in exploring a data table, comprising: The running module is configured to run a plug-in for implementing the method according to any one of claims 1 to 27 for a data set in the data table in response to a user opening the data table in the application.

80. A device for assisting a user in exploring a data table, comprising: a display module for displaying an exploration area in a predetermined area of ​​a data table in response to a user opening the data table in an application, the data table including an employee information table; The training module is used to respond to the user's prediction request for the target field in the data table, use the field values ​​corresponding to at least part of other fields in the single data as input, and use the field value corresponding to the target field in the single data as output to automatically train the machine learning model. The display module is further configured to output, in the exploration area, first explanation information representing the importance of at least part of the other fields to the target field; a first calculation module, configured to assign, based on a Shapley value assignment method, a score obtained by the machine learning model for predicting a single piece of data to each field in the at least some of the other fields, so as to obtain a score for the field in the single piece of data; A determination module is used to determine the importance of each field to the prediction of the target field by the machine learning model based on the sum of the scores of each field in the at least part of the other fields under multiple data, wherein the importance is positively correlated with the sum of the scores.

81. The apparatus of claim 80, wherein The predetermined area is at least one of a left side, a right side, an upper side, and a lower side of the data table.

82. The apparatus of claim 80, wherein The data table is an Excel table.

83. The apparatus of claim 80, wherein The training module automatically searches in the hyperparameter space based on hyperparameters preset according to experience to train the machine learning model.

84. The apparatus of claim 80, wherein The display module displays the first explanation information in a gradual manner during the process of calculating the importance.

85. The apparatus of claim 80, wherein The display module is also used to output a two-dimensional coordinate graph in the exploration area for representing the impact of a single field on the target field. One coordinate axis in the two-dimensional coordinate graph is used to represent the field value corresponding to the field, and the other coordinate axis is used to represent the score of the field. The two-dimensional coordinate graph includes multiple coordinate points, each coordinate point corresponds to a piece of data, and the score of the field is obtained by assigning the score obtained by the machine learning model for the prediction of the single piece of data to each field of the at least part of the other fields according to the distribution method of the Shapley value.

86. The apparatus of claim 85, wherein The display module is further configured to utilize the display characteristics of a coordinate point in the two-dimensional coordinate graph to represent a field value corresponding to another field in the data corresponding to the coordinate point.

87. The apparatus of claim 80, wherein In response to a user's prediction request for a piece of data in a data table, the display module outputs the prediction result of the machine learning model for the piece of data in the exploration area, and outputs second explanatory information in the exploration area on the importance of one or more fields of at least part of the other fields in the piece of data to the prediction result.

88. The apparatus of claim 87, further comprising: The second calculation module is used to assign the score obtained by the machine learning model for predicting the data to each field in the at least part of the other fields according to the distribution method of Shapley value, so as to obtain the score of each field under the data, and the importance of the field to the prediction result is positively correlated with the score of the field.

89. The apparatus of claim 88, wherein In response to a user's adjustment operation on the field values ​​of one or more fields in a piece of data in a data table, the display module also outputs the prediction result of the machine learning model for the adjusted piece of data, and outputs second explanatory information in the exploration area on the importance of one or more fields of at least part of the other fields in the adjusted piece of data to the prediction result.

90. The apparatus of claim 80, wherein The display module also uses the machine learning model to output changes in the field values ​​of at least some other fields in a certain data item in the data table based on the user's expected prediction results for the target field of the data item.

91. The apparatus of claim 80, further comprising: an output module, configured to output first data analysis information in the exploration area in response to a user's selection operation on one or more data columns in the data table, wherein the first data analysis information is used to represent data statistics of field values ​​corresponding to the data columns selected by the user; A recommendation module, configured to recommend one or more second data analysis information to the user, where the second data analysis information is data analysis information predicted based on the data column selected by the user; The second data analysis information is used to characterize the data statistics of the field values ​​corresponding to other data columns obtained through prediction, and / or The second data analysis information is used to represent data statistics of a combination of field values ​​corresponding to a combination of data columns formed by a data column selected by a user and other data columns obtained through prediction.

92. The apparatus of claim 91, wherein The second data analysis information is used to characterize the data statistics of the field values ​​corresponding to other data columns obtained through prediction, and / or The second data analysis information is used to represent data statistics of a combination of field values ​​corresponding to a combination of data columns formed by a data column selected by a user and other data columns obtained through prediction.

93. The apparatus of claim 91, wherein The recommendation module includes: A first recommendation module, configured to predict data analysis information suitable for recommendation to users based on a machine learning model; and / or A second recommendation module is configured to predict data analysis information suitable for recommendation to users based on statistical correlation; and / or The third recommendation module is used to predict data analysis information suitable for recommendation to users based on business rules.

94. The apparatus of claim 93, wherein: The first recommendation module: Based on the information about the data column currently selected by the user, using the machine learning model to predict the probability that the user will subsequently select each other data column, wherein the machine learning model is trained in the following manner: taking the information about the previously selected data column and the information about the other data columns as input and outputting the probability of the other data columns being selected; Analyze the data statistics of the field values ​​corresponding to the data column whose probability value is greater than the first predetermined threshold to obtain the second data analysis information, or analyze the data statistics of the field values ​​corresponding to the data column whose probability value is greater than the first predetermined threshold based on the first data analysis information of the data column selected by the user to obtain the second data analysis information.

95. The apparatus of claim 94, wherein The first recommendation module also obtains the user's acceptance feedback on the one or more second data analysis information recommended to the user, and updates the machine learning model based on the obtained acceptance feedback.

96. The apparatus of claim 93, wherein The second recommended module: According to the statistical correlations between different data columns, obtaining a data column whose statistical correlation with the data column selected by the user is higher than a second predetermined threshold; Analyze the data statistics of the field values ​​corresponding to the data columns whose statistical correlation is higher than the second predetermined threshold to obtain the second data analysis information, or analyze the data statistics of the field values ​​corresponding to the data columns whose statistical correlation is higher than the second predetermined threshold based on the first data analysis information of the data columns selected by the user to obtain the second data analysis information.

97. The apparatus of claim 91, wherein The exploration area includes a first interface area and a second interface area, The output module displays a chart of the first data analysis information in the first interface area in response to the user selecting a specific data column with a cursor, and While displaying a chart of the first data analysis information in the first interface area, the recommendation module displays one or more charts of the second data analysis information in the second interface area.

98. The apparatus of claim 97, further comprising: The third calculation module is used to pre-calculate the data statistics of the field values ​​corresponding to each data column when the data table is opened in the application, and cache the calculated data statistics in the memory. The output module displays a chart of first data analysis information generated based on corresponding data statistics extracted from the memory in the first interface area.

99. The apparatus of claim 97, wherein In response to a user selecting a field value in a chart of the first data analysis information, the recommendation module automatically updates one or more corresponding charts of the second data analysis information to data statistics of the field value under corresponding other field dimensions.

100. The apparatus of claim 99, further comprising: The fourth calculation module is used to pre-calculate that when one or more frequently selected field values ​​in the chart of the first data analysis information are selected, the corresponding one or more charts of the second data analysis information are automatically updated to the data statistics of the frequently selected field values ​​under the corresponding other field dimensions, and cache the calculated data statistics in the memory. The recommendation module automatically updates the charts of the corresponding one or more second data analysis information to the corresponding data statistics cached in the memory.

101. The apparatus of claim 97, wherein In response to the user selecting the at least one second data analysis information chart displayed in the second interface area with a cursor, dragging the cursor to the first interface area and releasing the cursor, the display module displays the at least one second data analysis information chart in the first interface area.

102. The apparatus of claim 97, wherein: In response to a user's connection operation on two charts in the first interface area, the display module connects one of the two charts to the other chart. In response to a user's selection operation on a field value represented in a first chart in a connected state, the display module highlights the selected field value relative to unselected field values, and / or the display module updates a second chart in a connected state to represent the data statistics of the selected field value under the field dimension corresponding to the second chart.

103. The apparatus of claim 97, wherein: In response to a user's connection operation on two charts in the first interface area, the display module connects one of the two charts to the other chart. The display module updates the other chart to represent data statistics of the field value represented by the one chart under the field dimension corresponding to the other chart.

104. A system comprising at least one computing device and at least one storage device storing instructions, wherein: When the instructions are executed by the at least one computing device, the instructions cause the at least one computing device to perform the method of any one of claims 1 to 51.

105. A computer-readable storage medium storing instructions, wherein: When the instructions are executed by at least one computing device, the at least one computing device is prompted to perform the method according to any one of claims 1 to 51.

Citation Information

Patent Citations

  • Rapid analysis tool used for spreadsheet application

    CN102982016A

  • Recommended content display method and device

    CN105808764A