Adding Machine Understanding to Spreadsheet Data
By introducing machine understanding into spreadsheet data, using knowledge graphs to identify entity types and determine semantic meanings, the problem that spreadsheet software in the prior art is difficult to automatically understand table data, and a more efficient and accurate user experience is achieved.
Patent Information
- Application Number
- CN202110047733.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-20
- Filing Date
- 2021-01-14
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-01-14
AI Technical Summary
Existing spreadsheet software is difficult to automatically identify and understand the semantic meaning of table data, resulting in poor user experience and waste of computing resources.
By introducing machine understanding into spreadsheet data, using knowledge graphs to identify entity types in cells and determine the semantic meaning of columns based on that entity type, thus providing more valuable features and functions.
It realizes automatic and real-time determination of the semantic meaning of columns, improves the efficiency and accuracy of spreadsheet software, and reduces the user's manual operation needs and error rates.
Smart Images

Figure CN112800773B_ABST
Abstract
Description
Technical Field
[0001] Aspects and embodiments of the present disclosure relate to electronic documents, and more particularly to adding machine understanding to spreadsheet data. Background Art
[0002] Spreadsheet documents can be used to organize and analyze large amounts of information. The information in a spreadsheet document can be contained in cells arranged in rows and columns on one or more worksheets. For example, spreadsheets can be used to manage and manipulate financial information, engineering information, or any organizational information. A user can use a spreadsheet software application to manipulate (e.g., create, edit, view, print, etc.) the spreadsheet. When editing a spreadsheet, the user can change the content of the spreadsheet by removing text, entering new text, formatting the spreadsheet layout, adding graphics or charts, or otherwise changing the content of the spreadsheet. Summary of the Invention
[0003] The following is a brief summary of the present disclosure to provide a basic understanding of some aspects of the present disclosure. This summary of the invention is not an extensive overview of the present disclosure. It is neither intended to identify key or critical elements of the present disclosure nor to delineate the scope of any particular embodiments of the present disclosure or the scope of any claims. Its sole purpose is to present some concepts of the present disclosure in a simplified form as a prelude to the more detailed description that is presented later.
[0004] A system and method for adding machine understanding to spreadsheet data are disclosed. A processing device can determine a data set, each piece of data including the content of a cell in a column of a spreadsheet presented to a user. The processing device can determine an entity type associated with the column based on the data set. The entity type can represent the semantic meaning of the data set in the column.
[0005] In some embodiments, to determine the entity type associated with the column, the processing device can identify one or more entities on a knowledge graph associated with each cell in the column. A knowledge graph can be described as a knowledge base having structured information about multiple semantic entities and the relationship connections between these semantic entities. The nodes of the knowledge graph can represent entities, and the relationship connections between the entities can be represented by edges. The knowledge graph can be a common knowledge graph or a privately maintained knowledge graph.
[0006] To identify one or more entities of each cell, the processing device may compare the data in the cell with the nodes on the knowledge graph. The processing device may then determine one or more commonly shared entity types of the column based on the number of cells in the column that share the entity type. To be considered a commonly shared entity type of the column, the number of cells sharing the entity type must meet a threshold condition. For example, the number of cells sharing the entity type must exceed 95% of the cells in the column to be considered a commonly shared entity type of the column. The method may then determine the entity type of the column by selecting the commonly shared entity type that is shared among the largest number of cells in the column.
[0007] In one embodiment, the processing device may then identify at least one graph related to the entity type associated with the column. The processing device may score various graph types based on a rule set that takes into account the semantic meaning of the column, or based on the output of a machine learning model. The rule set may guide which graphs should be considered for a particular semantic meaning associated with the column, and / or what score should be assigned to each of the graphs considered. The rules may be different for different users or different levels or categories of users, and the rules may be predetermined or configured based on user input. The machine learning model may be trained to assign scores to each graph type based on one or more entity types derived from the semantic meaning of the column, and optionally based on one or more characteristics of the user. The processing device may rank the graph types based on their scores and may provide the most relevant graph type (i.e., the graph type with the highest score) to be presented to the user. Additionally or alternatively, the processing device may provide more than one recommended graph type to the user based on the ranking of the graph types. For example, the processing device may provide the top N graph types to the user as recommendations based on the ranking, where N is an integer. As another example, the processing device may provide one or more graphs with scores that meet a threshold condition to the user as recommendations.
[0008] In one embodiment, a processing device may identify additional information related to the entity type of the column and may add one or more columns with the identified additional information to a spreadsheet. The processing device may identify entity types closely related to the entity type of the column based on a rule set that takes into account the semantic meaning of the column or based on the output of a machine learning model. The rule set may guide which entity types should be considered for a particular semantic meaning associated with the column and / or what score should be assigned to each of the considered entity types. As described above, this may be different for different users or different levels or classes of users and the rules may be predetermined or configured based on user input. A machine learning model may be trained to assign scores to each entity type based on the semantic meaning of the column and optionally based on one or more characteristics of the user. Entity types having scores that meet a threshold condition may be considered closely related. The processing device may then extract data from entities having closely related entity types to add to the spreadsheet for presentation to the user. Alternatively, the processing device may recommend adding additional information to the spreadsheet and, in response to receiving a user selection, may add one or more columns containing the additional information.
[0009] If the cells in a column do not meet the threshold condition such that the entity type of the column cannot be determined, the spreadsheet may operate as before without adding machine understanding of the data. However, the methods described herein may be practiced each time the spreadsheet is updated or at fixed intervals such that the entity type of each column may be determined subsequently.
[0010] The subject matter outlined above and described in more detail below is capable of automatically and in real-time determining the semantic meaning of columns. Such determined semantic meaning may automatically improve the existing functionality of spreadsheet software and provide other new functionality to users. These improvements in turn enhance the speed and efficiency of the computer running the spreadsheet software, reduce the error rate associated with the manual entity type determination process, and avoid unnecessary suggestions to the user.
[0011] The subject matter described herein is not limited to the advantages listed above by way of example. Other advantages are achievable and recognizable in view of the disclosure of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Aspects and embodiments of the present disclosure will be more fully understood from the detailed description given below and from the accompanying drawings of various aspects and embodiments of the present disclosure. However, the present disclosure should not be limited to the specific aspects or embodiments but is for explanatory and understanding purposes only.
[0013] Figure 1 Illustrated is an example system architecture for adding machine understanding to spreadsheet data for a spreadsheet software application in accordance with an embodiment of the present disclosure.
[0014] Figure 2 An example of a system architecture for adding machine understanding to spreadsheet data according to an embodiment of the present disclosure is depicted.
[0015] Figure 3 An example system according to an embodiment of the present disclosure is illustrated, which transfers data between entities in the system to add machine understanding to spreadsheet data.
[0016] Figure 4 A flowchart of a method for adding machine understanding to spreadsheet data and providing chart recommendations to a user according to an embodiment of the present disclosure is depicted.
[0017] Figure 5 A flowchart of a method for adding machine understanding to spreadsheet data and providing additional information to be presented to a user according to an embodiment of the present disclosure is depicted.
[0018] Figure 6 A block diagram of a computer system operating according to an embodiment of the present disclosure is depicted.
[0019] These drawings can be better understood when observed in conjunction with the following detailed description. Detailed Description
[0020] A spreadsheet software application can store user data as a series of text and numbers. The spreadsheet application can use the series of text and numbers to provide common spreadsheet functions, including formatting, graphing, filtering, etc. However, there is a gap between what the data in the spreadsheet means to a user in the real world and what the spreadsheet software can understand and operate on. For example, a user may know that a column includes a list of countries, yet for the spreadsheet application, the list of countries is just a series of text strings. Although the spreadsheet application may utilize the text strings when implementing the features and functions of the application, what the text strings represent to the computer's understanding may limit these features. Thus, the features and functions provided to the user may be irrelevant to the user's needs and may therefore result in wasted computing resources and a negative user experience.
[0021] Aspects and embodiments of the present disclosure address the above and other deficiencies or problems by providing techniques for adding machine understanding to spreadsheet data. Aspects of the present disclosure can then use the added machine understanding of the spreadsheet data to enhance existing features and functionality of spreadsheet applications and to provide new features to users. For example, by adding machine understanding to spreadsheet data, a spreadsheet application can identify that a column lists country names and can thus recommend a geographical chart to the user. As another example, the spreadsheet application can identify that the spreadsheet includes a list of dates and associated events and can recommend a timeline chart to the user. Adding machine understanding to spreadsheet data may also provide additional new functionality. For example, once the spreadsheet application identifies that a column includes a list of employee identification numbers, the application can automatically populate the employee names, email addresses, and office locations.
[0022] A spreadsheet software application can be provided to a client device over a network from a server. The server can store data and can present the data in the form of a spreadsheet through the spreadsheet application. The spreadsheet application and / or the server can analyze the data to provide valuable features to the user. The spreadsheet application can transmit spreadsheet data over the network to the server, which can trigger a semantic annotator on the server. The semantic annotator can add additional meaning to the spreadsheet data by linking the spreadsheet data to entities on a knowledge graph. The semantic annotator can determine the entity types of each cell and column and use their correspondingly determined types to annotate the cells and columns within the spreadsheet. Loading and / or saving a spreadsheet can automatically trigger the semantic annotator on the server. The semantic annotator can also be triggered at fixed intervals, such as every few milliseconds or every two seconds. The interval can be adjusted, for example, to compensate for the latency caused by the annotator. For a spreadsheet that contains multiple worksheets, the semantic annotator can annotate the columns of the spreadsheet one worksheet at a time during initial loading to balance the resource requirements for loading and processing multiple worksheets with the latency for determining column annotations.
[0023] The semantic annotator can begin by linking each cell of a column to one or more entities on the knowledge graph that are associated with the data in each cell. The knowledge graph can be a data structure that stores ontology data. In other words, the knowledge graph can be a knowledge base with structured information about multiple semantic entities and the relationship connections between the semantic entities. The knowledge graph can store information in the form of nodes and edges, where the nodes are connected by the edges. The nodes in the knowledge can represent entities. Entities can be people, places, items, ideas, topics, abstract concepts, concrete elements, other suitable things, or any combination of these. The entities in the knowledge graph can be related to any number of other entities by the edges, which can represent the relationships between the entities. Each semantic entity (also referred to as an "entity") has one or several types.
[0024] Each cell in a spreadsheet can be annotated with one or more entities associated with the data of that cell. In one embodiment, a semantic annotator can use a machine learning engine to label as many cells as possible with entities on a knowledge graph. For example, spreadsheet data can include a cell that includes the word (a registered trademark of The Coca-Cola Company). The machine learning engine can label the cell with many entities associated with on the knowledge graph, including entities related to product names, beverages, carbonated beverages, parent companies, etc. The machine learning engine can be trained to identify entities on the knowledge graph associated with a data set using historical data collected. The machine learning engine can be supervised or unsupervised. In other embodiments, the semantic annotator can use any other known technique to associate entities on the knowledge graph.
[0025] The semantic annotator can then determine the entity type associated with the column. The entity type can represent the semantic meaning of the data set in the column. In one embodiment, the semantic annotator can aggregate the determined cell entities to determine one or more commonly shared entity types of the column. The commonly shared entity type can be determined by the percentage of cells in the column associated with the same entity type in the knowledge graph. For example, if a column contains more than 95% of the cells associated with the "country" entity on the knowledge graph, then the column can be considered to have the semantic meaning of "country". The percentage or threshold can be adjusted according to, for example, the size of the column or according to the discipline itself. If more than the threshold number of cells are associated with more than one entity type, then the column can be considered to have more than one commonly shared entity type. For example, more than 95% of the cells in a column can reference both the "company name" entity and the "carbonated beverage" entity.
[0026] To determine the entity type associated with a column, the semantic annotator can determine the entity type shared among the majority of the cells in the column among the commonly shared entity types. That is, the column entity type must satisfy both the threshold condition and must be the entity type most shared within the column. For example, if more than 95% of the cells in a column reference both company names and carbonated beverages, but the number of cells referencing carbonated beverages exceeds the number of cells referencing company names, then the semantic annotator will determine that the entity type (and thus the semantic meaning) of the column is "carbonated beverage".
[0027] Then, the semantic annotator can annotate the columns with the determined entity types and annotate each cell in the column with the entities related to the entity type on the knowledge graph. For each cell, the semantic annotator can select the entity associated with the determined entity type of the column on the knowledge graph and link the cell to that entity. Continuing with the above example, once the semantic annotator has determined that the column type is "carbonated beverage", the semantic annotator can first annotate the column with the entity type and then annotate each cell with the specific entity related to the carbonated beverage listed in that cell. Thus, if the cell includes the word then the semantic annotator can use the " carbonated beverage" entity in the knowledge graph (instead of the " company" entity in the knowledge graph) to annotate the cell.
[0028] Once the semantic annotator has annotated the columns and cells in the spreadsheet, the spreadsheet application can then utilize the annotations to provide useful features to the user. For example, after determining that the column has the entity type of "carbonated beverage", the spreadsheet application can suggest to the user to include additional information related to carbonated beverages. As another example, after determining that the column has the entity type of "country", the spreadsheet application can recommend a geographical chart to the user.
[0029] Some technical advantages of the embodiments of the present disclosure include automatically and real-time enhancing spreadsheet software applications. By automatically determining the meaning of cells and columns, it is not required that the user manually provide or annotate individual cells or columns. Although some users may provide titles to describe the columns in the spreadsheet, many users do not. In addition, the title data may be unreliable because there is no standard practice for labeling columns. For example, while one user may find it helpful to label a column containing the names of countries as "country", another user may label the same column as "office" if, for example, the column lists the countries in which the company has offices. Additionally, the machine-readable meanings derived from the systems and methods described herein are up-to-date with the data in the spreadsheet. The spreadsheet application can enhance the operation of the computer by increasing the speed and efficiency of the application that runs on the computer, thereby resulting in faster processing times. The spreadsheet application can also avoid unwanted software recommendations, thereby reducing wasted computer resources.
[0030] The various aspects of the methods and systems cited above are described in detail below by way of example and not limitation.
[0031] Figure 1 FIG. illustrates an example system architecture 100 for adding machine understanding to spreadsheet data for a spreadsheet software application according to an embodiment of the present disclosure.
[0032] The system architecture 100 may include a server 150, which may be connected to one or more client devices 110 and a data store 190 via a network 105. The network 105 may be a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or a wide area network (WAN)), or a combination thereof. The network 105 may include a wireless infrastructure, which may be provided by one or more wireless communication systems, such as Wi-Fi hotspots connected to the network 105 and / or wireless carrier systems that may be implemented using various data processing. Additionally or alternatively, the network 105 may include a wired infrastructure (e.g., Ethernet).
[0033] The client device 110 may include a spreadsheet application 112 for creating, editing, viewing, and / or sharing spreadsheets. In one embodiment, the spreadsheet application 112 may be accessed via the network 105, and users may be allowed to collaborate with other users on a shared spreadsheet. That is, users may be able to simultaneously edit the content or format of a collaborative spreadsheet via their respective spreadsheet applications 112, and the resulting changes may be presented to the users in real time (e.g., with minimal millisecond latency) via their respective spreadsheet applications 112. The spreadsheet application 112 may include a variety of features, including a spreadsheet layout feature 114, an additional information feature 116, and a charting feature 118. The spreadsheet layout feature 114 may determine the layout of the spreadsheet. For example, the spreadsheet layout feature 114 may determine the appropriate positions of page breaks and header rows in the spreadsheet. The additional information feature 116 may determine additional data that may be useful to the client. The client may use the charting feature 118 to create charts and add them to the spreadsheet. Each function may be enhanced by determining the semantic meaning of the data in the spreadsheet. The spreadsheet application may contain additional functions not shown herein. Alternatively, not all of the features shown Figure 1 need to be present in the spreadsheet application. The data in the spreadsheet may be stored as spreadsheet data 194 in the data store 190.
[0034] The server device 150 may include a semantic annotator 152, which may include a cell entity determiner 154, a column annotator 156, and a cell annotator 158. The server device 150 may also include a recommendation generator 160.
[0035] The data store 190 may maintain a knowledge graph 192. The knowledge graph 192 may be a data structure that stores ontology data. In other words, the knowledge graph 192 may be a knowledge base having structured information about multiple semantic entities and the relationship connections between the semantic entities. The knowledge graph 192 may store information in the form of nodes and edges, where the nodes are connected by the edges. The nodes in the knowledge may represent entities. Entities may be people, places, items, ideas, topics, abstract concepts, concrete elements, other suitable things, or any combination of these. The entities in the knowledge graph may be related to any number of other entities by edges, and these edges may represent the relationships between the entities. Each semantic entity (also referred to as an "entity") has one or several types.
[0036] The data store 190 may store spreadsheet data 194, which includes the data contained in each cell of each spreadsheet. The semantic annotator 152 may analyze the data from each cell to determine the semantic meaning of each column. Then, the recommendation generator 160 may trigger the spreadsheet application 112 to make certain recommendations based on the semantic meaning of the columns.
[0037] The semantic annotator 152 may include a cell entity determiner 154, a column annotator 156, and a cell annotator 158. The cell entity determiner 154 may identify one or more entities associated with each cell in the column for the data. To do this, the cell entity determiner 154 may consult the knowledge graph 192 and identify one or more nodes on the knowledge graph associated with the data in the cell. For example, if the cell contains the word "English", the cell entity determiner 154 may consult the knowledge graph 192 and identify the nodes on the graph related to the word "English", which may include language nodes, nationality nodes, school subject nodes, etc. The cell entity determiner 154 assigns the entity to the cell.
[0038] The column annotator 156 may then determine the entity type of the column based on the semantic meaning. To determine the semantic meaning of the column, the column annotator 156 may aggregate the entity types of the cells in the column. Then, the column annotator 156 may determine how many cells share each of the entity types. If more than a threshold number of cells in the column share an entity type, the column annotator 156 may determine that the entity type is the commonly shared entity type of the column. The threshold may be a percentage of all the cells in the column, such as 95%. For example, the threshold may be adjusted according to the size of the column or the subject in the column.
[0039] Once the column annotator 156 has determined one or more commonly shared entity types for the column, it can determine the semantic meaning of the column. The semantic meaning can be based on the commonly shared entity type that is shared among the majority of cells in the column. The column annotator 156 can then annotate the metadata of the column with the semantic meaning or the entity type associated with the semantic meaning of the column. The cell annotator 158 can then annotate each cell in the column with a knowledge graph entity corresponding to the entity type of the column. That is, returning to the example above, if the column annotator 156 determines that the entity type of the column (containing the cell with the word "English") is "language", the cell annotator 158 can annotate the cell by linking the cell to the node in the knowledge graph 192 that represents "English language" (as opposed to the nodes representing "British nationality" or "English subject in school").
[0040] Then, the recommendation generator 160 can consume the meaning of the column and trigger various features in the spreadsheet application 112. For example, the recommendation generator 160 can use the entity type of the column and the assigned meaning of the cells to determine the best chart type to recommend to the user. For example, the recommendation generator 160 can determine that the most relevant chart for representing the data in the column is a geographic chart based on the entity type of the column. In one embodiment, the recommendation generator 160 can rank the available chart types based on the relevance of each type to the entity type of the column and display, for example, the top N relevant charts (where N is an integer greater than 0) to allow the user to select the desired chart. Alternatively or additionally, the recommendation generator 160 can determine that no chart will accurately represent the data in the column and can decide not to recommend a chart at all, thus avoiding making a bad recommendation.
[0041] In some embodiments, the recommendation generator 160 can trigger other additional features in the spreadsheet application 112. For example, the recommendation generator 160 can determine that a specific layout would be most beneficial for displaying the column data based on the determined column entity type. Based on this determination, the recommendation generator 160 can trigger the spreadsheet layout 114 feature of the spreadsheet application 112. In some embodiments, the recommendation generator 160 can trigger the additional information 116 feature of the spreadsheet application 112. For example, the column entity type can be "employee name", and the recommendation generator 160 can identify additional information related to the employee name entity type, such as "employee home office" or "employee ID". The recommendation generator 160 can then add one or more columns with the identified additional information to the spreadsheet.
[0042] In some embodiments, the determined chart and / or added columns can be presented to all users collaborating on a spreadsheet simultaneously, and each of the collaborating users can make changes to the determined chart and / or added columns, and then the other collaborating users can see the determined chart and / or added columns in real time. In some embodiments, different types of charts can be presented to these different users collaborating on a spreadsheet based on different actions / preferences the users have taken in the past.
[0043] Figure 2 Depicts an example of a system architecture 200 for adding machine understanding to spreadsheet data according to an embodiment of the present disclosure.
[0044] System architecture 200 includes one or more processing devices 201 and one or more data stores 190. In the example shown, processing device 201 includes a semantic annotator 152 and a recommendation generator 160. Processing device 201 can be included in Figure 1 server device 150 therein. Semantic annotator 152 can include a cell entity determiner 154, a column annotator 156, and a cell annotator 158. Recommendation generator 160 can include a chart recommendation generator 162 and an additional information generator 164. Data store 190 can maintain a knowledge graph 192, spreadsheet data 194, cell metadata 250, and column metadata 252. Cell metadata 250 can include descriptive attributes of each cell, such as the location of the cell in the spreadsheet, the time of the most recent update, whether the cell is visible, etc. Cell metadata 250 can also include an annotation that links the cell to a node on the knowledge graph 192. Similarly, column metadata 252 can include a description of each column, such as the time of the most recent update to the column, the name of the user who last modified the column, etc. Column metadata 252 can also include an annotation that specifies the semantic meaning of the column.
[0045] The cell entity determiner 154 can identify one or more entities in the data associated with each cell in a column. The data associated with each cell can be stored in the spreadsheet data 194 of the data store 190. The cell entity determiner 154 can include a cell data receiver 202, a knowledge graph (KG) lookupper 204, and an entity assigner 206. The cell data receiver 202 can receive cell data from the spreadsheet data 194. The KG lookupper 204 can look up the knowledge graph 192 to identify one or more entities associated with the cell data. In one embodiment, the KG lookupper 204 can compare the cell data with the nodes in the knowledge graph 192 to determine which nodes best represent the cell data. In one embodiment, machine learning can be used to train the KG lookupper 204 to identify the nodes on the knowledge graph associated with given data. Then, the entity lookupper 206 can label each cell from the knowledge graph with its one or more associated entities and entity types.
[0046] The column annotator 156 can include a cell entity aggregator 212, a semantic meaning determiner 214, and a column metadata updater 216. The cell entity aggregator 212 can aggregate the entity types of the cells in a column. For example, the cell entity aggregator 212 can generate a list of all the entity types referenced by the cells in the column and include the total number of cells referencing each entity type. Then, the semantic meaning determiner 214 can use the aggregated entity type information to determine the semantic meaning of the column. The semantic meaning can represent the machine understanding of each column in the spreadsheet. The machine understanding of a column can refer to the most relevant entity type (as determined in more detail below) among one or more entity types (or entities) that can represent the content of the column. Then the column can be annotated with the entity type associated with the semantic meaning of the column. The column metadata updater 216 can update the column metadata 252 to include the entity type associated with the semantic meaning of the column.
[0047] The cell annotator 158 can include an entity determiner 222 and a cell metadata updater 224. The entity determiner 222 can determine the entities on the knowledge graph 192 that are associated with the cell data and share the same entity type as the column. Although each cell can be associated with more than one entity, the entity determiner 222 can determine the entities that share the determined entity type of the column. The cell metadata updater 224 can then update the cell metadata 250 to link the cell to the determined entities on the knowledge graph 192.
[0048] The recommendation generator 160 may include a chart recommendation generator 162 and an additional information generator 164. The recommendation generator 160 may include additional components not shown herein. The chart recommendation generator 162 may rank available chart types based on the relevance of the available chart types to the entity type of the column being plotted. For example, the chart recommendation generator 162 may assign a score to each chart type based on a rule set that takes into account the entity type of the column or based on the output of a machine learning model. Once scores have been determined for each chart type, the chart recommendation generator 162 may identify the chart with the highest score and recommend that chart to the user and / or add it to the spreadsheet presented to the user. In one embodiment, the chart recommendation generator 162 may rank the chart types based on their scores and recommend the top N charts, where N is an integer, for example. Alternatively, the chart recommendation generator 162 may display all chart options to the user in order of their scores from highest score to lowest score. In one embodiment, the chart recommendation generator 162 may recommend one or more charts having scores that meet a threshold condition. For example, the assigned score range may be from 0 to 10, with 10 indicating the most relevant chart type, and the lowest acceptable value may be set to 7. Thus, a threshold condition may be set to identify charts having a score equal to or higher than 7, and the chart recommendation generator 162 may then recommend charts having a score equal to or higher than 7. The threshold condition may be adjustable.
[0049] In one example, the rule set may direct which charts should be considered for a particular semantic meaning associated with the content of the column, and / or what score should be assigned to each of the charts being considered for a particular semantic meaning associated with the content of the column. The semantic meaning may refer to one or more most relevant entities that may represent the content of the column. The rules may be different for different users or different levels or categories of users (e.g., based on browser type, screen resolution, etc.), and the rules may be predetermined or configurable based on user input.
[0050] In one embodiment, a training engine (not shown) may use machine learning techniques to train a machine learning model to predict what chart a user will select for a particular column. For example, the model can be trained to assign scores to each chart type based on one or more entity types derived from the semantic meaning of the column's content (and optionally based on one or more characteristics of the user to whom the chart recommendation will be provided). The training data can be based on historical data collected and can include training inputs and associated target outputs. The training inputs can include, for example, the entity types of columns in spreadsheets previously used by the user, the characteristics of the user, etc. The training outputs may include previous user actions regarding the charts used for the corresponding columns, such as user usage or chart selection, and / or the user's identified chart preferences. The training data can be associated with the user's account (e.g., based on username and password), and the user can opt out of the processing device that collects such data at any time.
[0051] Once trained, the model can take the entity type as input and provide weights ("scores") for each chart type as output. In some embodiments, the input provided to the model can also include one or more characteristics of a particular user (e.g., browser type, screen resolution, etc.) to recommend relevant charts for that particular user taking into account the different preferences of users at different levels. The chart type with the highest score may be the chart most relevant to the entity type. In some embodiments, given the different preferences and / or past actions of these users (e.g., different collaborators on the same spreadsheet), different chart recommendations can be provided to different users for the same column. The machine learning model can consist of, for example, single-level linear or non-linear operations (such as support vector machines [SVM]), or can be a deep network, i.e., a machine learning model consisting of multi-level non-linear operations. Examples of deep networks are neural networks with one or more hidden layers, and such machine learning models can be trained by adjusting the weights of the neural network, for example, according to the backpropagation learning algorithm, etc.
[0052] The additional information generator 164 can use the entity type of the column to generate additional information associated with that entity type. The additional information generator 164 can automatically populate additional information for the user. For example, the semantic annotator 152 may have determined that the column entity type is "country". The additional information generator 164 can collect additional information (such as population or capital city) related to the countries listed in the column from the knowledge graph 192. In some embodiments, similar to the chart recommendation generator 162 and as described below with respect to Figure 5Further described, the additional information generator 164 can assign a score to each entity type in the knowledge graph based on a ruleset that takes into account the entity type of a column or based on the output of a machine learning model. The additional information generator 164 can extract data from entities of entity types in the knowledge graph that have a score exceeding a threshold condition. Then, the additional information generator 164 can recommend to the user to include such additional information (i.e., the extracted data) in the spreadsheet, and / or can automatically add the information to the spreadsheet.
[0053] Figure 3 Illustrated is a system 300 according to an embodiment of the present disclosure and an example of data being passed between the entities illustrated in Figure 1 for adding machine understanding to spreadsheet data.
[0054] The system 300 can include a client device 110 and a server device 150. At operation 301, a user of the client device can load a spreadsheet in a spreadsheet application, which can trigger the transfer of spreadsheet data from the client device 110 to the server device 150. At operation 303, the user can save or update the spreadsheet in the spreadsheet application, which can trigger the transfer of the modified spreadsheet data from the client device 110 to the server 150. The transfer of spreadsheet data can be triggered by other operations, including scheduling at fixed intervals (e.g., every few milliseconds).
[0055] Once the server device 150 receives spreadsheet data from the client device 110, it can implement the semantic annotator 152. At operation 305, the semantic annotator can first implement the cell entity determiner 154, as described with respect to Figure 2 At operation 307, the semantic annotator 152 can use the determined cell entity data to implement the column annotator 156 and the cell annotator 158, as described with respect to Figure 2 as described.
[0056] At operation 309, the semantic annotator 152 can update the spreadsheet data saved on the server device 150, or can be stored in a data store (not shown). The updated spreadsheet data can now be considered as machine understanding of the spreadsheet data. At operation 311, the server device 150 can then generate recommendations for the user of the client device 110 regarding the machine understanding of the additions to the spreadsheet data. For example, using the machine understanding of the additions to the spreadsheet data, the spreadsheet application can recommend relevant chart types to the user, or can add additional relevant information to the spreadsheet.
[0057] Figure 4 and Figure 5Depicts flowcharts of method 400 and method 500 performed in accordance with some embodiments of the present disclosure. Method 400 and method 500 may be executed by a client-based application running on client device 110. The client-based application may be implemented by a processing device of client device 110.
[0058] For simplicity of illustration, methods 400 and 500 of the present disclosure are depicted and described as a series of acts. However, acts in accordance with the present disclosure may occur in various orders and / or concurrently and have other acts not presented and described herein. In addition, not all illustrated acts may be required to implement methods 400 and 500 in accordance with the disclosed subject matter. Further, those skilled in the art will understand and appreciate that methods 400 and 500 may alternatively be represented as a series of interrelated states via a state diagram or events. Additionally, it should be appreciated that methods 400 and 500 disclosed in this specification can be stored on an article of manufacture to facilitate the transfer and conveyance of such methods to a computing device. As used herein, the term "article of manufacture" is intended to encompass a computer program accessible from any computer-readable device or storage medium.
[0059] Figure 4 Is a flowchart of method 400 for adding machine understanding to spreadsheet data and providing chart recommendations to a user in accordance with some embodiments of the present disclosure.
[0060] Refer to Figure 4 , at operation 402, a processing device of client device 110 may determine a data set, each piece of data including the content of a cell in one or more cells in a column of a spreadsheet presented to the user.
[0061] At operation 404, the processing device may determine the entity type associated with a column based on the data set. The entity type may represent the semantic meaning of the data set in that column of the spreadsheet. Determining the entity type associated with a column may include identifying one or more entities associated with the data of each cell in the column. To identify the one or more entities associated with the data of each cell, the processing device may identify entities from a knowledge graph having a knowledge base that has structured information about multiple entities and the relationship connections between the entities. Once the processing device has identified the one or more entities associated with each cell, the processing device may determine one or more commonly shared entity types of the cells in the column. The commonly shared entity type may be determined based on the number of cells sharing the same entity type. If the number of cells sharing the entity type meets a threshold, then the entity type may be determined to be commonly shared. For example, if 95% of the cells in a column share one entity type, the processing device may determine that the entity type is commonly shared throughout the column. The threshold may be adjustable. Then, the processing device may determine the semantic meaning of the column by determining which commonly shared entity type is shared among the largest number of cells. Then, the entity type most commonly shared is the semantic meaning of the column.
[0062] At operation 406, the processing device may identify at least one chart (e.g., the most relevant one) among the multiple charts that is related to the entity type associated with the column. That is, the processing device may determine a score for each chart among the multiple charts based on a rule set that takes into account the semantic meaning of the column or based on the output of a machine learning model. Then, the processing device may rank the charts and identify the chart with the highest score. In one embodiment, the processing device may identify one or more charts related to the entity type associated with the column. If the scores of the one or more identified charts meet a threshold condition, they may be considered relevant.
[0063] In one example, the rule set may direct which charts should be considered for a particular semantic meaning associated with the content of the column, and / or what score should be assigned to each of the charts considered for a particular semantic meaning associated with the content of the column. The rules may be different for different users or different levels or classes of users (e.g., based on browser type, screen resolution, etc.), and the rules may be predetermined or configurable based on user input.
[0064] In one embodiment, the processing device may identify rules corresponding to entity types associated with columns. Based on the rules, the processing device may determine a subset of charts for the entity type associated with the column, and may determine a score for each chart in the subset. Then, the processing device may identify the chart with the highest score in the determined subset of charts. Additionally or alternatively, the processing device may identify one or more charts in the subset whose scores meet a threshold condition (e.g., exceed a set value).
[0065] As described above, the training engine may use machine learning techniques to train a machine learning model to predict what chart a user will select for a specific column. Once trained, the model may take an entity type as input and provide weights ("scores") for each chart type as output. In some embodiments, the input provided to the model may also include one or more characteristics of a specific user (e.g., browser type, screen resolution, etc.) to take into account different preferences of different levels of users and recommend relevant charts for that specific user. The chart type with the highest score may be the chart most relevant to the entity type.
[0066] At operation 408, the processing device may provide the identified chart for presentation to the user. In one embodiment, the processing device may provide a recommendation to the user based on the identified chart. Then, in response to receiving a user selection of the recommended chart, the processing device may provide the chart for presentation to the user. In another embodiment, the processing device may provide a list of recommended charts to the user, the list ranked by the scores determined at operation 406. In some embodiments, different chart recommendations may be provided to different users for the same column in view of the different preferences and / or past actions of these users (e.g., different collaborators on the same spreadsheet).
[0067] Figure 5 A flowchart of a method 500 for adding machine understanding to spreadsheet data and providing additional information for presentation to a user in accordance with an embodiment of the present disclosure is depicted.
[0068] At operation 502, the processing device may determine a data set, each piece of data including the content of a cell in one or more cells in a column of a spreadsheet presented to the user.
[0069] At operation 504, the processing device may determine an entity type associated with the column based on the data set, the entity type representing the semantic meaning of the data set in the column of the spreadsheet. This process may be the same as the operation 404 described with respect to Figure 4 the description.
[0070] At operation 506, the processing device can identify additional information related to the entity type associated with the column. In some embodiments, to identify additional information related to the entity type of the column, the processing device can identify, based on a knowledge graph, one or more additional entity types that are closely related to the entity type associated with the column. In particular, the processing device can first identify a first node in the knowledge graph that is associated with the entity type of the column. Then, the processing device can identify one or more nodes that are closely related to the first node. For example, if a node is connected by no more than one edge, the node may be closely related to another node. That is, if two nodes are directly connected to each other in the knowledge graph, they can be considered closely related. The data in one or more closely connected nodes can be identified as additional information related to the entity type associated with the column.
[0071] In one embodiment, the processing device can also use a ruleset to guide which entity types should be considered as additional information for a particular semantic meaning associated with the content of the column, and / or what score should be assigned to each of the considered entity types for a particular semantic meaning associated with the content of the column. The rules may be different for different users or different levels or classes of users (e.g., based on browser type, screen resolution, etc.), and the rules can be predetermined or configured based on user input.
[0072] In one implementation, the processing device can identify a rule corresponding to the entity type associated with the column. Based on the rule, the processing device can determine a subset of additional entity types regarding the entity type associated with the column, and can also determine the score of each additional entity type in the subset. Then, the processing device can identify one or more additional entity types in the determined subset of additional entity types whose scores meet a threshold condition. For example, the assigned score can range from 0 to 10, where 10 represents the most closely related entity type, and the lowest acceptable value can be set to 7. Thus, a threshold condition can be set to identify entity types with a score equal to or higher than 7. The threshold condition can be adjustable.
[0073] In some embodiments, to identify one or more additional entity types closely related to the entity type associated with a column, a processing device may use machine learning techniques to train a machine learning model to predict what additional information a user might add to a spreadsheet containing a particular column. For example, the model may be trained to assign weights or scores to each entity type in a knowledge graph based on one or more entity types derived from the semantic meaning of the content of the column (and optionally based on one or more characteristics of the user to whom the additional information will be provided). The training data may be based on historical data collected and may include training inputs and associated target outputs. The training inputs may include, for example, the entity types of the columns in spreadsheets previously used by the user, the characteristics of the user, etc. The training outputs may include: previous user actions regarding the additional information for the corresponding column, such as the user's inclusion of certain types of information, and / or the user's identified information preferences (e.g., the user's preference for using standard units of measurement instead of metric units). The training data may be associated with the user's account (e.g., based on a username and password), and the user may opt out of the processing device collecting such data at any time.
[0074] Once trained, the model can take the entity type as input and provide weights ("scores") for each additional entity type as output. In some embodiments, the input provided to the model may also include one or more characteristics of a particular user (e.g., browser type, screen resolution, etc.) to take into account the different preferences of users at different levels and recommend additional information for that particular user. The entity type with the highest score may be the entity type most relevant to the input entity type. In some embodiments, given the different preferences and / or past actions of different users (e.g., including different collaborators on the same spreadsheet), different entity types may be provided for the same column to different users. The machine learning model may consist of, for example, single-stage linear or non-linear operations (such as support vector machines [SVM]), or may be a deep network, i.e., a machine learning model consisting of multi-stage non-linear operations. An example of a deep network is a neural network with one or more hidden layers, and such machine learning models may be trained by adjusting the weights of the neural network, for example, according to a backpropagation learning algorithm, etc.
[0075] At operation 508, the processing device can add one or more columns with identified additional information to a spreadsheet. In one embodiment, the processing device can provide recommendations to a user based on the identified additional information. Then, in response to receiving a user selection, the processing device can add the identified additional information to one or more columns in the spreadsheet. That is, if there is an entity type closely related to the entity type of the column, the processing device can add an additional column including data from the closely related entity type. For example, if the column entity type is "country", the closely related entity type can be "capital city", and the processing device can add a column listing the capital city of each country listed in the "country" column. If more than one entity type is closely related to the entity type of the column, the processing device can add more than one column. For example, the processing device can add columns listing the capital city of each country and the population of each country. The processing device may be able to determine whether the additional information already exists in the spreadsheet to avoid recommending the addition of duplicate data.
[0076] Figure 6 A block diagram depicting an example computing system operating in accordance with one or more aspects of the present disclosure. In various illustrative examples, computer system 600 can correspond to any of the computing devices within Figure 1 system architecture 100. In one embodiment, computer system 600 can be Figure 1 server device 150 of
[0077] In certain embodiments, computer system 600 can be connected to other computer systems (e.g., via a network such as a local area network (LAN), intranet, extranet, or the Internet). Computer system 600 can operate in the capacity of a server or client computer in a client-server environment, or as a peer computer in a peer-to-peer or distributed network environment. Computer system 600 can be provided by a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), cellular phone, web device, server, network router, switch or bridge, or any device capable of executing a set of instructions (sequentially or otherwise) that specify actions to be taken by that device. Additionally, the term "computer" should include any collection of computers that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods described herein.
[0078] In yet another aspect, computer system 800 can include a processing device 602, volatile memory 604 (e.g., random access memory (RAM)), non-volatile memory 606 (e.g., read-only memory (ROM) or electrically erasable programmable ROM (EEPROM)), and a data storage device 616 that can communicate with each other via a bus 608.
[0079] The processing device 602 can be provided by one or more processors such as: general-purpose processors (such as, for example, complex instruction set computing (CISC) microprocessors, reduced instruction set computing (RISC) microprocessors, very long instruction word (VLIW) microprocessors, microprocessors implementing other types of instruction sets, or microprocessors implementing a combination of various types of instruction sets) or dedicated processors (such as, for example, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), or network processors).
[0080] The computer system 600 can also include a network interface device 622. The computer system 600 can also include a video display unit 610 (e.g., LCD), an alphanumeric input device 612 (e.g., keyboard), a cursor control device 614 (e.g., mouse), and a signal generation device 620.
[0081] The data storage device 616 can include a non-transitory computer-readable storage medium 624 on which instructions 626 encoding any one or more of the methods or functions described herein can be stored, the instructions including implementing methods 400 and 500, and for implementing Figure 1 and Figure 2 the semantic annotator 152 and the recommendation generator 160 of
[0082] The instructions 626 can also reside, in whole or in part, within the volatile memory 604 and / or within the processing device 602 during their execution by the computer system 600, and thus, the volatile memory 604 and the processing device 602 can also constitute machine-readable storage media.
[0083] Although the computer-readable storage medium 624 is shown as a single medium in the illustrative example, the term "computer-readable storage medium" should include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store a set or multiple sets of executable instructions. The term "computer-readable storage medium" should also include any tangible medium that can store or encode a set of instructions executable by a computer, the instructions causing the computer to perform any one or more of the methods described herein. The term "computer-readable storage medium" should include, but is not limited to, solid-state memory, optical media, and magnetic media.
[0084] The methods, components, and features described herein can be implemented by discrete hardware components or can be integrated into the functionality of other hardware components such as ASICs, FPGAs, DSPs, or similar devices. Additionally, the methods, components, and features can be implemented by firmware modules or functional circuits within a hardware device. Further, the methods, components, and features can be implemented using any combination of a hardware device and computer program components or with a computer program.
[0085] Unless specifically stated otherwise, terms such as "initiate", "send", "receive", "analyze", etc. refer to actions and processes performed or implemented by a computer system that manipulate and transform data represented as physical (electronic) quantities within the registers and memory of the computer system into other data similarly represented as physical quantities within the memory or registers of the computer system or other such information storage, transmission, or display devices. Additionally, as used herein, the terms "first", "second", "third", "fourth", etc. are labels used to distinguish different elements and may not have an ordinal meaning based on their numerical names.
[0086] The examples described herein also relate to an apparatus for performing the methods described herein. This apparatus can be specifically constructed to perform the methods described herein, or it can include a general-purpose computer system selectively programmed by a computer program stored in the computer system. Such a computer program can be stored in a computer-readable tangible storage medium.
[0087] The methods and illustrative examples described herein are not inherently related to any particular computer or other device. A variety of general-purpose systems can be used in accordance with the teachings described herein, or it may prove convenient to construct more specialized devices to perform methods 400 and 500, and / or each of their individual functions, routines, subroutines, or operations. Examples of the structures of various such systems are set forth in the foregoing description.
[0088] The foregoing description is intended to be illustrative, not restrictive. Although the present disclosure has been described with reference to specific illustrative examples and embodiments, it should be recognized that the present disclosure is not limited to the examples and embodiments described. The scope of the present disclosure should be determined with reference to the appended claims and the full scope of equivalents to which the claims are entitled.
Claims
1. A method for generating chart recommendations based on machine understanding of spreadsheet data, comprising: determining, by a processing device, a data set, each data including the content of a cell in one or more cells in a column of a spreadsheet presented to a user; annotating the one or more cells in the column with entities representing the semantic meaning of the data in the cell; determining, by aggregating the entities of the one or more cells, an entity type associated with the column, the entity type representing the semantic meaning of the data set in the column; identifying at least one chart among a plurality of charts that is related to the entity type associated with the column, wherein the identifying includes: determining a score for each chart in the plurality of charts based on a rule set associated with the semantic meaning of the column; and identifying the chart with the highest score or one or more charts with scores that meet a threshold condition; and providing the at least one chart for presentation to the user.
2. The method according to claim 1, wherein determining the entity type associated with the column comprises: identifying one or more entities associated with the data of each cell in the column, each entity having an entity type; determining, based on the number of cells in the column that share a common entity type, one or more commonly shared entity types of the column, wherein the number of cells meets a threshold condition; and identifying, from the one or more commonly shared entity types, the commonly shared entity type having the largest number of cells sharing the common entity type as the entity type associated with the column.
3. The method according to claim 2, wherein identifying one or more entities associated with the data of each cell in the column comprises: identifying one or more entities from a knowledge graph including a knowledge base, the knowledge base having structured information about a plurality of entities and the relationship connections between the plurality of entities.
4. The method according to claim 1, wherein identifying at least one chart among the plurality of charts that is related to the entity type associated with the column comprises: using training data including collected historical data to train a machine learning model to assign a score to each chart in the plurality of charts based on the entity type in the knowledge graph; providing the entity type associated with the column as an input to the trained machine learning model; obtaining an output of the trained machine learning model, the output indicating the score of each chart in the plurality of charts for the column; and identifying the chart with the highest score.
5. The method according to claim 1, wherein identifying at least one chart among the plurality of charts that is related to the entity type associated with the column comprises: identifying rules corresponding to the entity type associated with the column; determining, based on the rules, a subset of the plurality of charts related to the entity type associated with the column and the score of each chart in the subset; and identifying the chart with the highest score in the subset.
6. The method according to claim 1, wherein providing the at least one chart for presentation to the user comprises: providing a recommendation to the user based on the at least one chart; and in response to receiving a user selection, providing the at least one chart for presentation to the user.
7. The method according to claim 1, further comprising annotating metadata associated with the column based on the semantic meaning of the column.
8. The method according to claim 3, further comprising annotating metadata associated with each cell in the column based on nodes in the knowledge graph, wherein the nodes represent an entity among the one or more entities that is associated with the data in the corresponding cell and is associated with the semantic meaning of the column.
9. A system for generating additional information for a spreadsheet based on machine understanding of spreadsheet data, the system comprises: a memory; and a processing device communicatively coupled to the memory, the processing device being configured to: determine a data set, each data including the content of a cell in one or more cells in a column of the spreadsheet presented to the user; annotate the one or more cells in the column with entities representing the semantic meaning of the data in the cells; determine an entity type associated with the column by aggregating the entities of the one or more cells, the entity type representing the semantic meaning of the data set in the column; identify additional information related to the entity type associated with the column; and add one or more columns with the additional information to the spreadsheet.
10. The system according to claim 9, wherein in order to determine the entity type associated with the column, the processing device is further configured to: identify one or more entities associated with the data of each cell in the column, each entity having an entity type; determine one or more commonly shared entity types of the column based on the number of cells in the column that share a common entity type, wherein the number of cells meets a threshold condition; and identify, from the one or more commonly shared entity types, the commonly shared entity type having the largest number of cells with the shared common entity type as the entity type associated with the column.
11. The system according to claim 10, wherein in order to identify one or more entities associated with the data of each cell in the column, the processing device is further configured to identify one or more entities from a knowledge graph including a knowledge base having structured information about a plurality of entities and the relationship connections between the plurality of entities.
12. The system according to claim 9, wherein in order to identify additional information related to the entity type associated with the column, the processing device is further configured to: identify one or more additional entity types closely related to the entity type associated with the column; Identify one or more additional entities in a knowledge graph including a knowledge base having structured information about a plurality of entities and relationship connections between the plurality of entities, wherein the one or more additional entities are associated with the one or more additional entity types; and Identify data in the one or more additional entities.
13. The system according to claim 12, wherein, in order to identify one or more additional entity types closely related to the entity type associated with the column, the processing device is further configured to:[[]] Use training data including the collected historical data to train a machine learning model to assign scores to each entity type in the knowledge graph; Provide the entity type associated with the column as an input to the trained machine learning model; Obtain the output of the trained machine learning model, the output indicating the scores of each additional entity type; And Identify the one or more additional entity types having scores that meet a threshold condition.
14. The system according to claim 12, wherein, in order to identify one or more additional entity types closely related to the entity type associated with the column, the processing device is further configured to:[[]] Identify rules corresponding to the entity type associated with the column; Based on the rules, determine a set of additional entity types regarding the entity type associated with the column and the scores of each additional entity type in the set; and Identify one or more additional entity types having scores that meet a threshold condition in the set of additional entity types.
15. The system according to claim 9, wherein, in order to add one or more columns having the additional information to the spreadsheet, the processing device is further configured to:[[]] Provide recommendations to the user based on the additional information; and In response to receiving a user selection, add the additional information to the spreadsheet.
16. The system according to claim 11, wherein the processing device is further configured to:[[]] Annotate the metadata associated with the column based on the semantic meaning of the column; and Annotate the metadata associated with each cell in the column based on the nodes in the knowledge graph, wherein the nodes represent an entity in the one or more entities that is associated with the data in the corresponding cell and is associated with the semantic meaning of the column.
17. A non-transitory machine-readable storage medium including instructions that, when executed, cause a processing device to perform operations for generating chart recommendations based on machine understanding of spreadsheet data, the operations include:[[]] Determine a data set, each piece of data including the content of a cell in one or more cells in a column of a spreadsheet presented to a user; Annotate the one or more cells in the column with entities representing the semantic meaning of the data in the cells; Determine an entity type associated with the column by aggregating the entities of the one or more cells, the entity type representing the semantic meaning of the data set in the column; Identify one or more diagrams among a plurality of diagrams that are related to the entity type associated with the column, where the identification includes: Determine a score for each diagram in the plurality of diagrams based on a rule set associated with the semantic meaning of the column; And Identify the diagram with the highest score or one or more diagrams with scores that meet a threshold condition; And Provide the one or more diagrams for presentation to the user.
18. The non-transitory machine-readable storage medium according to claim 17, wherein determining the entity type associated with the column Includes: Identify one or more entities associated with the data of each cell in the column, where each entity has an entity type; Determine one or more commonly shared entity types of the column based on the number of cells in the column that share a common entity type, where the number of cells meets a threshold condition; And Identify, from the one or more commonly shared entity types, the commonly shared entity type with the largest number of cells sharing the common entity type as the entity type associated with the column.
19. The non-transitory machine-readable storage medium according to claim 18, wherein identifying one or more entities associated with the data of each cell in the column Includes: Identify one or more entities from a knowledge graph including a knowledge base, the knowledge base having structured information about a plurality of entities and the relationship connections between the plurality of entities.
20. The non-transitory machine-readable storage medium according to claim 17, wherein identifying one or more diagrams among the plurality of diagrams that are related to the entity type associated with the column Includes: Use training data including collected historical data to train a machine learning model to assign a score to each diagram in the plurality of diagrams based on the entity type in the knowledge graph; Provide the entity type associated with the column as an input to the trained machine learning model; Obtain the output of the trained machine learning model, the output indicating the score of each diagram in the plurality of diagrams of the column; And Identify one or more diagrams with scores that meet a threshold condition.
21. The non-transitory machine-readable storage medium according to claim 17, wherein identifying one or more diagrams among the plurality of diagrams that are related to the entity type associated with the column Includes: Identify rules corresponding to the entity type associated with the column; Determine a subset of the plurality of diagrams related to the entity type associated with the column and the score of each diagram in the subset based on the rules; And Identify one or more diagrams in the subset with scores that meet a threshold condition.
22. The non-transitory machine-readable storage medium according to claim 17, wherein providing the one or more diagrams for presentation to the user Includes: Provide a recommendation to the user based on the one or more diagrams; Receive a user selection associated with one of the one or more diagrams; And In response to receiving the user selection, provide the selected diagram for presentation to the user.
23. The non-transitory machine-readable storage medium according to claim 17, wherein the operation further comprises annotating metadata associated with the column based on the semantic meaning of the column.
24. The non-transitory machine-readable storage medium according to claim 19, wherein the operation further comprises annotating metadata associated with each cell in the column based on nodes in the knowledge graph, wherein the nodes represent an entity among the one or more entities that is associated with the data in the corresponding cell and is associated with the semantic meaning of the column.
Citation Information
Patent Citations
Generating charts from data in a data table
CN109643329A
Cited By
Dynamically and selectively updated spreadsheets based on knowledge monitoring and natural language processing
US12547823B2