Systems, methods, and computer-readable media for improved table identification using neural networks
By applying machine learning methods in spreadsheets, especially convolutional neural networks and probabilistic graph models, the problem of automatically identifying table boundaries is solved, achieving high accuracy and real-timeness.
Patent Information
- Application Number
- CN201980046843.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-07-13
- Filing Date
- 2019-06-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2039-06-18
AI Technical Summary
The prior art is difficult to automatically and accurately identify table boundaries in spreadsheets, especially in complex environments containing various visual information and data.
Using machine learning methods, especially convolutional neural networks (CNNs), each unit of a spreadsheet is identified, classified as corner or non-corner categories, and corner units are combined through a probability graph model to derive the table.
It realizes the identification of table boundaries with high accuracy in complex spreadsheets, overcomes the problems of fragility and inefficiency of rule-based methods, and can realize real-time table identification on consumer-grade hardware.
Smart Images

Figure CN112424784B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to table identification in a spreadsheet. In particular, the present disclosure relates to improved table identification in a spreadsheet using a neural network. Background Art
[0002] Spreadsheets allow for flexible arrangements of data, calculations, and presentations to provide users ranging in experience from novices to programming experts with the ability to examine, calculate, and make decisions based on the data in the spreadsheet. While the flexible arrangement of data provides many uses, it can also hinder automation tools that can benefit from a well-formed table structure.
[0003] There exists a problem of automatically identifying table boundaries in electronic tables. Although identifying table boundaries may be intuitive to a human, rule-based approaches for identifying table boundaries have met with limited success due to the variety of visual information and data in electronic tables.
[0004] Although the present disclosure specifically discusses table identification in spreadsheets, aspects of the present disclosure may be applicable not only to spreadsheets, but also to other file types including tables and / or flexible data arrangements. Summary of the invention
[0005] According to certain embodiments, systems, methods, and computer-readable media for improved table identification are disclosed.
[0006] According to certain embodiments, a computer-implemented method for improved table identification in an electronic form is disclosed. The method includes: receiving an electronic form including at least one table; identifying one or more of a plurality of categories for each cell of the received electronic form using machine learning, wherein the plurality of categories include corners and non-corners; and deriving at least one table in the received electronic form based on the one or more identified categories for each cell of the received electronic form.
[0007] According to certain embodiments, a system for improved table identification in a spreadsheet is disclosed. A system includes: a data storage device storing instructions for improved table identification in a spreadsheet; and a processor configured to execute the instructions to perform a method, the method including: receiving a spreadsheet including at least one table; using machine learning to identify one or more of a plurality of categories for each cell of the received spreadsheet, wherein the plurality of categories include corners and non-corners; and deriving at least one table in the received spreadsheet based on the one or more identified categories for each cell of the received spreadsheet.
[0008] According to certain embodiments, a computer-readable storage device storing instructions is disclosed that, when executed by a computer, causes the computer to perform a method for improved table identification in a spreadsheet. A method of a computer-readable storage device includes: receiving a spreadsheet including at least one table; identifying one or more of a plurality of categories for each cell of the received spreadsheet using machine learning, wherein the plurality of categories include corners and non-corners; and deriving at least one table in the received spreadsheet based on the one or more identified categories for each cell of the received spreadsheet.
[0009] Other purposes and advantages of the disclosed embodiments will be set forth in part in the following description, and in part will become clear from the description, or may be learned by the practice of the disclosed embodiments. The purposes and advantages of the disclosed embodiments will be realized and obtained by the elements and combinations particularly pointed out in the appended claims.
[0010] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosed embodiments, as claimed. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In the following detailed description, reference will be made to the accompanying drawings. The accompanying drawings illustrate different aspects of the present disclosure, and where appropriate, reference numerals showing similar structures, components, materials and / or elements in different figures are similarly labeled. It should be understood that various combinations of other structures, components and / or elements than those specifically shown are contemplated and within the scope of the present disclosure.
[0012] In addition, many embodiments of the present disclosure are described and shown herein. The present disclosure is neither limited to any single aspect or its embodiments, nor to any combination and / or arrangement of such aspects and / or embodiments. In addition, each aspect of the present disclosure and / or its embodiments may be used alone, or used in combination with one or more other aspects of the present disclosure and / or its embodiments. For the sake of brevity, some arrangements and combinations are not discussed and / or shown separately herein.
[0013] Figure 1 depicts an exemplary spreadsheet having two vertically aligned tables according to an embodiment of the present disclosure;
[0014] Figure 2 depicts an exemplary high-level overview of table identification decomposed into corner identification and table derivation according to an embodiment of the present disclosure;
[0015] Figure 3 Depicts an exemplary architecture of a convolutional neural network for corner identification according to an embodiment of the present disclosure;
[0016] Figure 4An exemplary probabilistic graphical model of event location for a table according to an embodiment of the present disclosure is depicted;
[0017] Figure 5 Another exemplary electronic form having at least one table according to an embodiment of the present disclosure is depicted, the electronic form may have an incorrectly identified table;
[0018] Figure 6 a bar graph depicting the accuracy of a corner identification model and ablation according to an embodiment of the present disclosure;
[0019] Figure 7 a graph depicting the runtime of a table identification phase as a function of spreadsheet size according to an embodiment of the present disclosure;
[0020] Figure 8 Depicted are methods for improved table identification using neural networks according to embodiments of the present disclosure;
[0021] Fig. 9 depicts a high-level diagram of an exemplary computing device that may be used in accordance with the systems, methods, and computer-readable media disclosed herein in accordance with an embodiment of the present disclosure; and
[0022] Fig.10 A high-level diagram of an exemplary computing system that can be used in accordance with the systems, methods, and computer-readable media disclosed herein is depicted in accordance with an embodiment of the present disclosure.
[0023] Again, many embodiments are described and shown here. The present disclosure is neither limited to any single aspect or its embodiment, nor to any combination and / or arrangement of such aspects and / or embodiments. Each aspect of the present disclosure and / or its embodiments may be used alone or in combination with one or more other aspects of the present disclosure and / or its embodiments. For the sake of brevity, many of these combinations and arrangements are not discussed separately herein. DETAILED DESCRIPTION
[0024] Those skilled in the art will recognize that various implementations and embodiments of the present disclosure can be practiced according to the description. All such implementations and embodiments are intended to be included within the scope of the present disclosure.
[0025] As used herein, the terms "comprises," "comprising," "have," "having," "include," "including," or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but may also include other elements not expressly listed or inherent to such process, method, article, or apparatus. The term "exemplary" is used in the sense of an "example" rather than an "ideal." Additionally, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, the phrase "X employs A or B" is intended to mean any natural inclusive permutation. For example, the phrase "X employs A or B" is satisfied in any of the following instances: X employs A; X employs B; or X employs both A and B. Furthermore, the articles "a" and "an" used in the application and the appended claims should generally be construed to mean "one or more," unless otherwise specified or clear from the context to refer to the singular.
[0026] For the sake of brevity, conventional techniques for performing methods and other functional aspects of the systems and servers (and the various operating components of the systems) may not be described in detail herein. In addition, the connecting lines shown in the various figures included herein are intended to represent exemplary functional relationships and / or physical couplings between the various elements. It should be noted that many alternative and / or additional functional relationships or physical connections may exist in embodiments of the subject matter.
[0027] Reference will now be made in detail to exemplary embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. Wherever possible, the same reference numerals will be used throughout the drawings to refer to the same or like components.
[0028] Among other things, the present disclosure generally relates to a method for improving table identification in a spreadsheet using a neural network. Although the present disclosure specifically discusses table identification in a spreadsheet, aspects of the present disclosure may be applicable not only to spreadsheets, but also to other file types including tables and / or flexible data arrangements.
[0029] Go to Figure 1 , Figure 1 An exemplary spreadsheet having two vertically aligned tables according to an embodiment of the present disclosure is depicted. Figure 1As shown, there may be two tables in the spreadsheet (i.e., the upper left corner of the first table is at cell A5 and the lower right corner is at cell G16, represented as A5:G16, and the upper left corner of the second table is at cell A18 and the lower right corner is at cell G29, represented as A18:G29). Figure 1 As shown, the two tables are stacked vertically and include text describing the contents of the columns and rows. Because a spreadsheet includes a variety of information and data, automatically detecting the two tables can be challenging. One possible reason why it is difficult to automatically detect the tables is that blank rows (e.g., rows 7 and 15) may increase readability, but blank rows may also confuse attempts to detect table boundaries. Another possible reason why it is difficult to automatically detect the tables is that the table headers may be irregular (e.g., the cells in column A of rows 6 and 19 are empty), which may confuse the classification of rows as headers. Another possible reason why it is difficult to automatically detect the tables is that the use of cell border lines may be inconsistent. Therefore, the variety of layouts and visual cues may make automatic table identification techniques fragile and ineffective.
[0030] Automatic table extraction from a spreadsheet may have a variety of uses. For example, extracting individual tables from a spreadsheet as database relations may allow database queries to be applied to a data set, and may potentially include filters and joins across multiple tables, as well as natural language based queries. Additionally, data type analysis based on header information of one or more tables in a spreadsheet may be used for type inference and / or consistency verification.
[0031] In an embodiment of the present disclosure, a table in a spreadsheet may be a rectangular area including rows and columns with data cells, and the table may include one or more additional rows at the top of the table, including a string (such as text) and descriptive information about the data in the column (which may be referred to as a table header). A table in a spreadsheet may also include one or more subtotal rows, which may summarize the data in a column using an accumulation function and / or a percentage.
[0032] An exemplary embodiment of the present disclosure describes a table identification technique that uses a convolutional neural network to identify independent corners of one or more tables, such as a patch-based convolutional neural network. However, the present disclosure may not be limited to convolutional neural networks, and embodiments of the present disclosure may use other types of neural networks for table identification. After identifying candidates for each corner, the corners may be stitched together to construct one or more tables.
[0033] As discussed in more detail below, embodiments of the present disclosure provide a table identification method that applies a convolutional neural network ("CNN") to a rectangular context surrounding a cell to classify the cell as a potential corner, and then combines the corner prediction results according to a graphical model to identify the table. The table identification method is end-to-end data-driven, is several orders of magnitude faster than full spreadsheet convolution and / or object detection, and is capable of generating large training datasets from a limited number of spreadsheets (such as 1,638 spreadsheets), as described below.
[0034] In an embodiment of the present disclosure, S may be a spreadsheet having R×C cells, where R is the number of rows and C is the number of columns. Spreadsheet S may include one or more tables. These tables have one or more visual cues and other features, such as borders, colors, font styles, and / or spacing conventions. Each table can be a rectangular or square region of a spreadsheet S, identified by six (6) corner cells, including the top left corner ("TL"), the top right corner ("TR"), the bottom left corner ("BL"), the bottom right corner ("BR"), the header bottom left corner ("HBL"), and the header bottom right corner ("HBR"). In spreadsheet notation, a table can be defined as a region TL:BR and its header region TL:HBR.
[0035] return Figure 1 , Figure 1 An exemplary spreadsheet with multiple tables is depicted. Given a spreadsheet S such as Figure 1 For a spreadsheet S, the table identification problem can be as follows: (a) identify whether the spreadsheet S includes any tables; and (b) identify the location of each table in the spreadsheet S and the corresponding column headers.
[0036] like Figure 1 As shown, the table identification problem can be challenging for a number of reasons, including: a table can be specified using various visual and other attributes of one or more cells of a spreadsheet, which can identify a table in different combinations. Building rule-based table identification based on various visual and other attributes can be fragile and cumbersome. In addition, tables in a spreadsheet can be sparse by utilizing empty rows and / or empty columns to lay out the content in the table of the spreadsheet so that it looks clearer on the page. Therefore, the common heuristic approach of "largest dense rectangle / square around a given cell" (used in industrial spreadsheet software, such as Microsoft Excel and Google Sheets) may not be valid in many spreadsheets. Table identification should be as precise as possible, because errors in a few rows and / or columns may ignore important data in the spreadsheet.
[0037] To train the neural network, each spreadsheet with zero or more tables is used as a data set. The following discussion is exemplary technical details of embodiments of the present disclosure. The training data used to train the neural network can come from multiple sources, including but not limited to: an expert annotated corpus of spreadsheets, a crowdsourced annotated corpus of spreadsheets, and / or a mixture of different methods.
[0038] For example, exemplary embodiments of the present disclosure are tested using a collection of spreadsheets from a publicly available dataset that includes spreadsheets from different fields and / or a variety of different sources. For example, the table identification dataset can be an open source and / or closed source collection of personal and / or commercial spreadsheets.
[0039] A preliminary service can be used to examine an electronic form to determine whether the electronic form includes a table with data. For example, a set of preliminary annotators can be used to identify electronic forms with tables with data. Then, a predetermined number of annotations can be performed on the electronic form. After performing the predetermined number of annotations, the remaining electronic forms in the annotated data set can include at least one table.
[0040] After processing the electronic form to determine whether the electronic form includes a table with data, the second service can annotate the rest of the electronic form. For example, for each electronic form, the annotator can identify one or more tables in the electronic form and mark the corners of each table and any header rows of the table. Figure 1 , table identification may be challenging for the service and / or annotators, and the service and / or annotators may not agree on the table, corner and / or header rows. Therefore, majority consensus may be used among annotators when applicable. Alternatively and / or additionally, the labels of the most experienced annotators may be decisive.
[0041] According to an embodiment of the present disclosure, a data set for training a neural network can be processed to generate an annotated data set before training the neural network. The processing can be performed by an automated service, one or more annotators, and / or annotators of crowdsourcing annotation. Crowdsourcing annotation can establish a human performance baseline table with respect to table identification.
[0042] Spreadsheets automatically annotated by the service can be merged with spreadsheets annotated by other methods (such as expert annotators). The combined annotated spreadsheets can constitute a training dataset for table identification. When using a combined dataset, for overlapping files, the annotations made by expert annotators are preferred as the ground truth.
[0043] As described above, embodiments of the present disclosure use a trained convolutional neural network ("CNN") for table identification. CNNs can overcome the challenges of processing spreadsheets, which can contain tens of thousands of cells. For example, on a consumer-grade machine, 10,000 wide full-image convolutions with more than 50 channels may be too slow to be practical. In addition, as used in embodiments of the present disclosure, CNNs can overcome the challenges of requiring hundreds of thousands of data points to learn high-level features of the data. For example, a labeled table identification dataset of this magnitude may not exist.
[0044] To address these two challenges, table identification can be decomposed into two stages: (a) corner identification and (b) table derivation. Corner identification can apply CNN to detect cells in a spreadsheet that are likely to be table headers or corners. Table derivation uses a probabilistic graphical model to group sets of corner cells into candidate tables. Figure 2 An exemplary high-level overview of table identification decomposed into corner identification and table derivation according to an embodiment of the present disclosure is depicted. Reducing table identification to corner identification can allow the problem to be reproduced by cell classification, allowing more available data to be used and making the convolution of local blocks faster.
[0045] The corner identification stage of table identification can solve the following problem: given a context block of W×W cells, classify the center cell of the block as one of the six table corners (i.e., TL, TR, BL, BR, HBL, or HBR (as described above) or a non-corner (“NaC”). To avoid the multi-class classification problem, an HBL corner or an HBR corner can be defined as a cell directly below the visual table title that appears on the left table boundary or the right table boundary, respectively.
[0046] To classify a cell of a spreadsheet as a table corner, a neural network can consider various visual and other features of the cell as well as the immediate surrounding context of the cell. Each cell in a block can be represented by a predetermined number of channels that encode visual and other cell features, e.g., F = 51 features including font style, color, border, alignment, data type formatting, etc.
[0047] Figure 3 An exemplary architecture of a convolutional neural network for corner identification according to an embodiment of the present disclosure is depicted. Figure 3As shown in the process, the CNN for corner identification can process W×W×F input volumes to predict a 7-way softmax distribution over the corner category C = {TL, TR, BL, BR, HBL, HBR, NaC}. The CNN architecture can apply bottleneck convolutions down to 24 features. The CNN architecture can then upsample to 40 features. Thereafter, the CNN can apply a stack of three (3) convolutional layers with an exponential linear unit ("ELU") activation function, followed by two fully connected layers with dropout. The convolutional neural network can be trained to not overfit. Deeper convolutional neural networks with more fully connected layers can be used. The convolutional neural network model can be constructed to include a plurality of neurons and can be configured to output one or more of a plurality of categories for each cell of a received spreadsheet. The plurality of neurons can be arranged in a plurality of layers including at least one hidden layer and can be connected by a plurality of connections. Although Figure 3 A specific exemplary CNN architecture is depicted, but any CNN architecture that obtains a probability distribution over classes may be used.
[0048] According to an exemplary embodiment of the present disclosure, in order to train a neural network for corner identification, as described above, training blocks can be generated for each table from a labeled real benchmark data set. Specifically, for each real benchmark table T in the training data set, corner blocks and a predetermined number of NaC block samples, such as N=30 NaC blocks, can be extracted from various key positions in the table. The key positions of the NaC blocks in the table include one or more of the following: the center of the table, the midpoint of the edge, the gap between the title and the data, the cells around the table corners, and random cell samples that are generally distributed near the target corners. The neural network model can be trained using a categorical cross entropy objective on a predetermined number of categories (e.g., seven categories of TL, TR, BL, BR, HBL, HBR, and NaC).
[0049] A key challenge in the training process can be class imbalance, as there may be many more non-corner ("NaC") cells than corner cells. For example, each table provides only a single example of each corner type, but may include multiple examples of NaC cells. To correct for class imbalance, one or more regularization techniques may be applied. One regularization technique includes undersampling the possible NaC cells for each table by η·N, where η can be a hyperparameter that can be optimized on a validation set, as described below. Another regularization technique includes applying dropout with a predetermined probability p, where p can be a hyperparameter. As Figure 3As shown in the CNN model of , the predetermined probability p can be, for example, p=0.5. Another regularization technique includes applying data augmentation to introduce additional corner examples with synthetic noise, where each synthetic example can add small random noise around the corner cell only in the out-of-table cell. The noise can conform to a symmetric bimodal normal distribution with peaks at the edges of the input block and σ=(W / 2) 1 / 4 Thus, the total number of corner examples for each table can match the number of examples of the NaC. Yet another regularization technique includes rescaling the weights of the training examples for each table so that the total weight of the corner examples matches the total weight of the NaC examples.
[0050] For corner identification, high recall may be more important than high precision. The purpose of corner identification is to suggest candidates for subsequent table derivation of corners. Table derivation may be robust to false positive corners, because it is unlikely that all suggested corners are false. However, table derivation may be sensitive to false negatives. Therefore, one or more of the above regularization techniques may increase the likelihood that the neural network model identifies all true corners, while weakening the NaC classification.
[0051] Turning to the second stage of table identification, table derivation can use the candidate corners suggested by the corner identification of the neural network model to derive the most likely set of tables that appear in a given spreadsheet. The main challenge of table derivation can be uncertainty. The corner candidates identified by the CNN may be noisy and include false positives. In order to accurately model the table derivation problem from noisy corner candidates, a probabilistic graphical model ("PGM") can be used.
[0052] Figure 4 An exemplary probabilistic graphical model for event location of a table according to an embodiment of the present disclosure is depicted. T=t, B=b, L=l and R=r may be independent events, indicating that the top, bottom, left and right boundary lines of the target table T are equal to t, b, l and r, respectively. θ (x, y; γ) can be a unit<x,y> The probability of belonging to the corner class γ∈C is given by applying the unit<x,y> The corner identification model of the block centered at f θ Estimated. According to f θ , we can observe the estimated probability of a table corner appearing at the intersection of the corresponding boundary lines, for example, Pr(T=t, L=l)=f θ (t, l; T L ). The spatial confidence score of the candidate table T can be calculated as follows.
[0053]
[0054] Algorithm 1 shown below provides a table derivation algorithm according to an embodiment of the present disclosure. The table derivation algorithm may start with, if f θ The predicted probability of a cell is at least a predetermined amount (such as 0.5), then the table is for each cell<x,y> ∈S assigns category γ∈C. The table derivation algorithm can then score each table candidate using the spatial confidence of the table according to formula (1) above. The table derivation algorithm can also suppress overlapping tables with lower spatial confidence scores. To ensure that the tables match the real-world semantics, the table derivation algorithm can vertically merge tables that do not have at least one well-defined title. Finally, the table derivation algorithm can select tables with spatial confidence scores above a predetermined threshold τ, which can be a hyperparameter optimized on the validation dataset, as described above.
[0055]
[0056] As described above, the data set used to train the neural network model can be split into multiple data sets. For example, the data set can be split and / or divided into a training data set, a validation data set, and a test data set. In an exemplary embodiment of the present disclosure, a spreadsheet annotated by a preliminary service can be used for training. In the annotated spreadsheet, various predetermined percentages can be used for different data sets. For example, 50% can be delegated as a training data set, 30% can be delegated as a validation data set, and 20% can be delegated as a test data set. Smaller tables (such as 2×2 cells or less) can be deleted because such annotated tables may result in false positives.
[0057] In order to evaluate various aspects of exemplary embodiments of the present disclosure, different performance indicators can be used. One indicator that can be used is Jaccard accuracy / recall. Average Jaccard accuracy / recall can be used for object detection. Average Jaccard accuracy and recall can be defined as follows: accuracy=true positive area / predicted area, recall=true positive area / true reference area, and area can be the number of cells covered by the proposed table. Average tables of accuracy, recall, and F-values can be calculated for different tables.
[0058] Another metric that can be used is the mean normalized mutual information (“NMI”), which is a metric commonly used to evaluate clustering techniques. As discussed in this paper, the table identification problem can be viewed as the clustering of spreadsheet cells, where tables correspond to clusters. Formally, for a true benchmark table on a spreadsheet S, and prediction tables T1, ..., T m , each cell c∈S can be associated with the corresponding true cluster tc(c)∈{0,1,...,n} and predicted cluster pc(c)∈{0,1,...,m}. NMI can be calculated as NMI([tc(c1),...,tc(c|S| )],[pc(c1),...,pc(c |S| )]). For simplicity, in embodiments of the present disclosure, only cells in a bounding box around all predicted and true reference tables may be considered.
[0059] Both the average Jaccard precision / recall and the average normalized mutual information can be used in different scenarios. Precision / recall can be a common evaluation metric for each table prediction. However, if Figure 1 As shown, a spreadsheet may include multiple tables, and disambiguating multiple tables may be part of the challenge of the table identification problem. The NMI may be a clustering metric of the accuracy of a set of predicted tables relative to a true baseline for each spreadsheet.
[0060] The corner identification neural network model can be trained on various types of computer hardware, such as by using two NVIDIA GTX 1080Ti GPUs, using open source software libraries such as TensorFlow for a predetermined amount of time, such as 16 epochs (10-12 hours per model). As shown below, experiments for embodiments of the present disclosure were run on an Intel Xeon 3.60GHz 6-core CPU with 32GB RAM (for feature extraction and table export) and a single NVIDIA GTX 1080Ti GPU (for corner identification).
[0061] Table identification techniques using neural networks according to exemplary embodiments of the present disclosure are evaluated against baselines including built-in table detection in Microsoft Excel, rule-based detection, and unit classification via random forests. Performance is evaluated using accurate table identification, normalized mutual information based on unit clustering, and Jaccard metric based on overlap, among other metrics.
[0062] For example, a collection of spreadsheets from various data sets was used to test the exemplary embodiments of the present disclosure. Using the Jaccard metric, the overall accuracy of the exemplary embodiments of the present disclosure was 74.9%, the recall was 93.6%, and the F value was 83.2%, which is comparable to the accuracy of 86.9%, the recall of 93.2%, and the F value of 89.9% of human performance.
[0063] As discussed herein, data-driven approaches can detect tables more accurately than rule-based approaches. Due to the inherent difficulty of the table identification problem, table identification using neural networks can be evaluated against two rule-based baselines. As described above, the first baseline is the "Select Table" feature in Microsoft Excel, which selects the maximum density rectangle around a given cell. The second baseline is a separate rule-based implementation of table detection, which finds the largest contiguous block of rows with compatible data type signatures (assuming that empty cells are "compatible" with any data type), while ignoring gaps. Both the first and second baselines detect tables around a given cell, which can be user-selected. The interaction can be simulated by selecting m=5 random cells in the real benchmark table and averaging the metrics over m. However, such simulations may make it impossible to measure NMI for the rule-based baselines.
[0064] Table 1 summarizes the results from our hyperparameter sweeps and ablation tests using exemplary embodiments of the present disclosure, as discussed in detail below. In particular, η=1 / 4τ=3 / 4, and W-11.
[0065] As shown below, exemplary embodiments of the present disclosure may outperform rule-based methods in terms of Jaccard recall and F-value. Rule-based methods may have higher precision and lower recall due to conservative implementation, and the method only covers certain patterns of real-world tables. Data-driven table identification such as embodiments of the present disclosure may be able to cover a more diverse range of sparse table layouts.
[0066] Table 1
[0067]
[0068] As shown in Table 1, P represents Jaccard precision, R represents Jaccard recall, F represents F value, and I represents NMI. Each value shown in Table 1 is in percentage (%).
[0069] The DeExcelerator project is another type of data-driven approach that can use random forests and support vector machines ("SVMs") to classify spreadsheet cells into various data / title / metadata types, and then can apply rule-based algorithms to merge geometric regions derived from these classifications into tables. Comparisons of embodiments of the present disclosure with other types of data-driven approaches can be performed by evaluating the table identification of the present disclosure on its dataset against the claimed results of other types of data-driven approaches. For this evaluation, the following metrics can be used:
[0070]
[0071] As shown in formula (2), the above indicators and the data associated with them may not match the free-form real-world definition of the table. The definition of the table and the corresponding annotations are continuous with respect to regions of similar types, which is applicable to other types of rule-based table derivation algorithms.
[0072] Figure 5 Another exemplary spreadsheet having at least one table according to an embodiment of the present disclosure is depicted, which may have misidentified tables. Specifically, according to other types of data-driven methods, the spreadsheet is misidentified as having four (4) tables. Embodiments of the present disclosure may identify a single table with empty rows and empty columns. In such divergent cases, additional evaluation may be performed. Human agreement between our expert annotations and the original annotations may be measured, and the two annotations agree with a precision of 72.3%, a recall of 90.0%, and an F-value of 80.2%.
[0073] Table 2 summarizes the claimed results of exemplary embodiments of the present disclosure and other types of data-driven methods as described above for the original dataset and the re-labeled dataset as described above, with respect to the metrics provided by formula (2) and the comparison of the above metrics. For the re-labeled dataset, the exemplary embodiments of the present disclosure can have better performance than the original dataset, and for itself, the performance for the re-labeled dataset is better than the original dataset. The improvement in accuracy and F-value based on formula (2) can indicate that neural networks can better utilize the rich visual structure of spreadsheets than shallow classification methods.
[0074] Table 2
[0075]
[0076] As shown in Table 2, P DE , R DE and F DE They respectively represent the precision, recall and F value when there is a match between the two tables recorded according to formula (2), P represents Jaccard precision, R represents Jaccard recall, F represents F value, and I represents NMI. Each value shown in Table 2 is in percentage (%).
[0077] As described above, according to the embodiments of the present disclosure, the regularization technology can improve the accuracy of corner identification. Figure 6Bar graphs depicting the accuracy of corner identification models and ablations according to embodiments of the present disclosure. The results of corner identification using neural network models according to embodiments of the present disclosure can be evaluated as F-values of their predictive accuracy on different corner categories on validation and test datasets. The groups on the x-axis of the bar graph can depict different values of the hyperparameter η, which can control the number of non-corner training examples. Different values of the block size W can be depicted by the hashing of the bars. Each neural network model can represent consistently high accuracy, with the relative accuracy of each corner category varying. The values η=1 / 4 and W=11 can be selected based on their average performance over all corner categories on the validation set.
[0078] Based on exemplary embodiments of the present disclosure, table identification is possible for real-time use on consumer-grade hardware. Figure 7 A graph depicting the runtime of the table identification phase as a function of the size of the spreadsheet according to an embodiment of the present disclosure. Figure 7 As shown, the amount of time required to identify the table is linear in the number of cells in the spreadsheet, with a median run time of 71 milliseconds. Most of the time is probably spent extracting features from the spreadsheet and identifying corners. Corner identification for multiple blocks can be run in parallel on a GPU and automatically benefits from any further parallelization.
[0079] Although Figure 1 and 2 A neural network framework is depicted, but those skilled in the art will appreciate that neural networks can be performed with respect to models and can include the following stages: model creation (neural network training), model validation (neural network testing), and model utilization (neural network evaluation), although these stages may not be mutually exclusive. According to an embodiment of the present disclosure, a neural network can be implemented through training, inference, and evaluation stages. At least one server can execute the machine learning component of the table identification system described herein. As those skilled in the art will appreciate, machine learning can be performed with respect to models and can include multiple stages: model creation, model testing, model validation, and model utilization, although these stages may not be mutually exclusive. In addition, model creation, testing, validation, and utilization can be a continuous process of machine learning.
[0080] In the first machine learning phase, the model creation phase can involve extracting features from a spreadsheet of training data sets. As will be appreciated by those skilled in the art, these extracted features and / or other data can be derived from statistical analysis and / or machine learning techniques on large amounts of data collected over time based on patterns. Based on observations of this monitoring, the machine learning component can create a model (i.e., a set of rules or heuristics) for identifying corners from a spreadsheet. As described above, a neural network can be trained to de-emphasize false negatives.
[0081] In the second machine learning phase, the accuracy of the model created in the model creation phase can be tested and / or verified. In this phase, the machine learning component can receive a spreadsheet from a test data set and / or a validation data set, extract features from the test data set and / or the validation data set, and compare these extracted features to the predicted labels made by the model. By continuously tracking and comparing this information and for a period of time, the machine learning component can determine whether the model accurately predicts which parts of the spreadsheet may be corners. This testing and / or verification is often expressed in terms of accuracy: that is, what percentage of the time the model predicts a label. Information about the success or failure of predictions made by the model can be fed back to the model creation phase to improve the model, thereby improving the accuracy of the model.
[0082] The third machine learning stage can be based on a model that is verified to reach a predetermined threshold accuracy. For example, a model determined to have an accuracy rate of at least 50% may be suitable for the use stage. According to an embodiment of the present disclosure, in this third utilization stage, the machine learning component can identify corners from a spreadsheet where the model suggests that there are corners. When encountering a corner, the model suggests what type of corner exists and can store data. Of course, information based on the confirmation or rejection of each stored corner can be returned to the previous stage (testing, verification, and creation) as data for improving the model to improve the accuracy of the model.
[0083] Although the present disclosure specifically discusses spreadsheet processing, aspects of the present disclosure may be applicable not only to spreadsheets, but also to flexible data arrangement and / or classification problems.
[0084] As described above and below, embodiments of the present disclosure allow for reduced computational complexity and memory requirements, which can also reduce power consumption. Embodiments of the present disclosure can be implemented on mobile devices such as smartphones, tablets and / or even wearable items, smart speakers, computers, laptops, car entertainment systems, and the like.
[0085] Figure 8 A method 800 for improved table identification, such as by using machine learning including neural networks, according to an embodiment of the present disclosure is depicted. The method 800 may begin at step 802, where a data set including a plurality of spreadsheets may be received, wherein at least one of the plurality of spreadsheets includes at least one table. Each cell of the received spreadsheet may include a predetermined number of channels encoding visual and other features of the cell, and the features may include font style, color, border, alignment, data type format, etc.
[0086] Then, at step 804, the data set can be processed to produce an annotated data set, which includes a plurality of annotated electronic tables with identified tables, corners, and header rows. Prior to annotation, the electronic tables can be processed by the service to determine whether each electronic table includes a table with data. The second service can then annotate the remaining electronic tables. In addition, for each electronic table, the annotator can identify one or more tables in the electronic table and mark the corners of each table and any header rows of the table.
[0087] In step 806, the corner cells of the table and a predetermined number of non-corner cells of the table can be extracted from each table of the annotated data set. Then, in step 808, at least one regularization technique can be applied to the annotated data set based on the extracted corner cells of each table and the predetermined number of non-corner cells of each table. The predetermined number of non-corner cells of the table may include non-corner cells of the table from predetermined positions in the table. The predetermined positions of the non-corner cells of the table may include one or more of the following: the center of the table, the midpoint of the edge of the table, the gap between the table title and the data, the cells around the table corners of the table, and random cells that are generally distributed near the target corners of the table.
[0088] At least one regularization technique may include one or more of the following: i) undersampling a number of non-corner cells for each table based on hyperparameters that can be optimized on a validation set, ii) applying dropout in the architecture of the neural network model, iii) applying data augmentation to introduce additional corner examples with synthetic noise, where each synthetic example adds a small random noise around the corner cell to cells that are not in the table, and iv) rescaling the weights of the annotated dataset so that the total weight of the corner cells matches the total weight of the non-corner cells. The synthetic noise may follow a symmetric bimodal normal distribution with peaks at the edges of the input cells and σ = (W / 2) 1 / 4 .
[0089] In step 810, a neural network model including a plurality of neurons may be constructed, the neural network model being configured to output one or more of a plurality of categories for each cell of the received spreadsheet, the plurality of neurons being arranged in a plurality of layers including at least one hidden layer and being connected by a plurality of connections. At least one hidden layer of the neural network model may include at least one long short-term memory layer. Of course, the construction of the neural network model may occur at any time before the use of the neural network model, and the construction of the neural network model may not be limited to occurring before and / or after at least the above steps.
[0090] At step 812, a neural network model may be trained using the annotated data set, and the neural network model may be configured to output one or more of a plurality of categories for each cell of the received spreadsheet. In addition, the trained neural network model may include a convolutional neural network. The plurality of categories may include upper left corner, upper right corner, lower left corner, lower right corner, title lower left corner, title lower right corner, and non-corner, as described above. After training the neural network model, at step 814, a trained neural network model configured to output one or more of a plurality of categories for each cell of the received spreadsheet may be output.
[0091] After training the neural network model, at step 816, a test data set and / or a validation data set including spreadsheets, each spreadsheet having zero or more tables, may be received. Then, at step 818, the trained neural network may be evaluated using the received test data set. Alternatively, steps 816 and 818 may be omitted, and / or may be performed at different times. Once evaluated as passing a predetermined threshold, the trained neural network may be utilized. Additionally, in certain embodiments of the present disclosure, the steps of method 800 may be repeated to generate a plurality of trained neural networks. The plurality of trained neural networks may then be compared to each other and / or to other neural networks.
[0092] At step 820, a spreadsheet including at least one table may be received, wherein each cell of the received spreadsheet includes a predetermined number of channels encoding visual and other features of the cell. Then, at step 822, a neural network model may be used to identify one or more of a plurality of categories for each cell of the received spreadsheet, wherein the plurality of categories may include at least corners and non-corners. The identification may include classifying each cell of the received spreadsheet based on the predetermined number of channels encoding visual and other features of the cell. Additionally and / or alternatively, the classification of each cell of the received spreadsheet may also be based on an immediate surrounding context of the cell.
[0093] Then, at step 824, at least one table in the received spreadsheet can be derived based on the one or more categories identified for each cell of the received spreadsheet, wherein the deriving can be accomplished using a probabilistic graphical model. Using the probabilistic graphical model to derive at least one table in the received spreadsheet can include estimating a probability of a corner appearing at an intersection of corresponding boundary lines, calculating a spatial confidence score for each candidate table, and selecting each candidate table as a table when the spatial confidence score for each candidate table is greater than a predetermined threshold. The predetermined threshold can be a hyperparameter optimized based on a validation dataset.
[0094] Fig. 9A high-level diagram of an exemplary computing device 900 that can be used in accordance with the systems, methods, modules, and computer-readable media disclosed herein, according to an embodiment of the present disclosure is depicted. For example, according to an embodiment of the present disclosure, the computing device 900 can be used in a system that uses a neural network to process data such as a spreadsheet. The computing device 900 may include at least one processor 902 that executes instructions stored in a memory 904. The instructions may be, for example, instructions for implementing the functions described as being performed by one or more components discussed above or instructions for implementing one or more of the methods described above. The processor 902 may access the memory 904 via a system bus 906. In addition to storing executable instructions, the memory 904 may also store data, spreadsheets, one or more neural networks, and the like.
[0095] The computing device 900 may further include a data repository (also referred to as a database) 908 that the processor 902 can access via the system bus 906. The data repository 908 may include executable instructions, data, examples, features, etc. The computing device 900 may also include an input interface 910 that allows external devices to communicate with the computing device 900. For example, the input interface 910 may be used to receive instructions from an external computer device, a user, etc. The computing device 900 may also include an output interface 912 that interfaces the computing device 900 with one or more external devices. For example, the computing device 900 may display text, images, etc. through the output interface 912.
[0096] It is conceivable that external devices communicating with the computing device 900 via the input interface 910 and the output interface 912 can be included in an environment providing substantially any type of user interface with which the user can interact. Examples of user interface types include graphical user interfaces, natural user interfaces, and the like. For example, a graphical user interface can accept input from a user using input devices such as a keyboard, a mouse, a remote controller, and can provide output on output devices such as a display. In addition, a natural user interface can enable a user to interact with the computing device 900 in a manner that is not constrained by input devices such as a keyboard, a mouse, a remote controller, and the like. Instead, a natural user interface can rely on voice identification, touch and stylus identification, gesture identification on the screen adjacent to the screen, air gestures, head and eye tracking, sound and voice, vision, touch, gestures, machine intelligence, and the like.
[0097] In addition, although shown as a single system, it should be understood that the computing device 900 can be a distributed system. Thus, for example, several devices can communicate via a network connection and can jointly perform the tasks described as being performed by the computing device 900.
[0098] Go to Fig.10 , Fig.10A high-level diagram of an exemplary computing system 1000 that can be used in accordance with the systems, methods, modules, and computer-readable media disclosed herein, according to an embodiment of the present disclosure is depicted. For example, computing system 1000 can be or include computing device 900. Additionally and / or alternatively, computing device 900 can be or include computing system 1000.
[0099] Computing system 1000 may include multiple server computing devices, such as server computing device 1002 and server computing device 1004 (collectively referred to as server computing devices 1002-1004). Server computing device 1002 may include at least one processor, and memory; at least one processor executes instructions stored in the memory. The instructions may be, for example, instructions for implementing the functions described as being performed by one or more components discussed above or instructions for implementing one or more of the methods described above. Similar to server computing device 1002, in addition to server computing device 1002, at least a subset of server computing devices 1002-1004 may each include at least one processor, and memory, respectively. In addition, at least a subset of server computing devices 1002-1004 may include corresponding data repositories.
[0100] The processor of one or more server computing devices 1002-1004 may be or may include a processor such as processor 902. Furthermore, one or more memories of one or more server computing devices 1002-1004 may be or may include a memory such as memory 904. Furthermore, one or more data stores of one or more server computing devices 1002-1004 may be or may include a data store such as data store 908.
[0101] The computing system 1000 may also include various network nodes 806 for transmitting data between the server computing devices 1002-1004. In addition, the network nodes 1006 may transmit data from the server computing devices 1002-1004 to an external node (e.g., external to the computing system 1000) via a network 1008. The network nodes 1002 may also transmit data from an external node to the server computing devices 1002-1004 via the network 1008. The network 1008 may be, for example, the Internet, a cellular network, etc. The network nodes 1006 may include switches, routers, load balancers, etc.
[0102] The fabric controller 1010 of the computing system 1000 can manage the hardware resources of the server computing devices 1002-1004 (e.g., the processors, memory, data storage repositories, etc. of the server computing devices 1002-1004). The fabric controller 1010 can also manage the network nodes 1006. In addition, the fabric controller 1010 can manage the creation, provisioning, de-provisioning, and supervision of managed runtime environments instantiated on the server computing devices 1002-1004.
[0103] As used herein, the terms "component" and "system" are intended to encompass a computer-readable data storage device configured with computer-executable instructions that, when executed by a processor, cause certain functions to be performed. The computer-executable instructions may include routines, functions, etc. It should also be understood that a component or system may be located on a single device or distributed among multiple devices.
[0104] The various functions described herein can be implemented in hardware, software, or any combination thereof. If implemented in software, the function can be stored on a computer-readable medium and / or transmitted via a computer-readable medium as one or more instructions or codes. Computer-readable media may include computer-readable storage media. Computer-readable storage media may be any available storage medium that a computer can access. As an example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, disk storage, or other magnetic storage devices, or any other hardware medium that can be used to store desired program codes in the form of instructions or data structures and can be accessed by a computer. As used herein, disks and optical disks may include compact disks ("CDs"), laser optical disks, optical disks, digital versatile disks ("DVDs"), floppy disks, and blue-ray disks ("BDs"), wherein disks typically copy data magnetically, and optical disks typically copy data optically using lasers. In addition, propagation signals are not included in the scope of computer-readable storage media. Computer-readable media may also include communication media, including any media that helps transfer a computer program from one place to another. Connections may be, for example, communication media. For example, if the software is transmitted from a website, server or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line ("DSL") or wireless technologies (such as infrared, radio and microwave), then the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technologies (such as infrared, radio and microwave) are included in the definition of communication media. Combinations of the above may also be included within the scope of computer-readable media.
[0105] Alternatively and / or additionally, the functions described herein may be performed at least in part by one or more hardware logic components. For example, and not limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays ("FPGAs"), application specific integrated circuits ("ASICs"), application specific standard products ("ASSPs"), systems on chip ("SOCs"), complex programmable logic devices ("CPLDs"), and the like.
[0106] The above description includes examples of one or more embodiments. Of course, it is not possible to describe every possible modification and variation of the above-mentioned apparatus or method for the purpose of describing the above-mentioned aspects, but a person of ordinary skill in the art will recognize that many further modifications and arrangements of various aspects are possible. Therefore, the various aspects described are intended to cover all such changes, modifications and variations that fall within the scope of the appended claims.
Claims
1. A computer-implemented method for improved table identification in a spreadsheet, the method comprising: receiving a data set comprising a plurality of electronic spreadsheets, wherein at least one electronic spreadsheet of the plurality of electronic spreadsheets comprises at least one table; processing the data set to produce an annotated data set comprising a plurality of annotated spreadsheets having identified tables, corners, and header rows; as well as training a neural network model using the annotated data set, wherein the neural network model is configured to output one or more categories of a plurality of categories for each cell of the received electronic form; receiving a spreadsheet including at least one table; for each cell of the received electronic table, identifying one or more of a plurality of categories using the neural network model, wherein the plurality of categories include an upper left corner, an upper right corner, a lower left corner, a lower right corner, a title lower left corner, a title lower right corner, and a non-corner, the title lower left corner being a cell directly below a visual table title on a left table border, and the title lower right corner being a cell directly below a visual table title on a right table border; as well as At least one table in the received electronic form is derived based on the one or more categories identified for each cell of the received electronic form.
2. The method according to claim 1, further comprising: Extracting, from each table of the annotated data set, corner cells of the table and a predetermined number of non-corner cells of the table; as well as At least one regularization technique is applied to the labeled dataset based on the extracted corner cells of each table and the predetermined number of non-corner cells of each table. 3 . The method of claim 2 , wherein the predetermined number of non-corner cells of the table comprises non-corner cells of the table from predetermined positions in the table.
4. A method according to claim 3, wherein the predetermined positions of the non-corner cells of the table include one or more of the following: the table center of the table, the midpoint of the edge of the table, the gap between the title and data of the table, the cells around the table corners of the table, and random cells normally distributed near the target corner of the table.
5. The method of claim 2, wherein the at least one regularization technique comprises one or more of: i) undersampling a plurality of non-corner cells for each table based on hyperparameters optimized on a validation set, ii) applying dropout in the architecture of the neural network model, iii) applying data augmentation to introduce additional corner examples with synthetic noise, wherein each synthetic example adds a small random noise around the corner cell to cells that are not in the table, and iv) rescaling the weights of the annotated dataset to provide a total weight of the corner cells that matches the total weight of the non-corner cells. The method of claim 5 , wherein the synthetic noise follows a symmetric bimodal normal distribution.
7. The method of claim 1, wherein the trained neural network model comprises a convolutional neural network.
8. The method according to claim 1, further comprising: Constructing the neural network model including a plurality of neurons, the neural network model is configured to output the one or more of the plurality of categories for each cell of the received spreadsheet, the plurality of neurons being arranged in a plurality of layers including at least one hidden layer and connected through a plurality of connections.
9. The method of claim 1 , wherein each cell of the received electronic form comprises a predetermined number of channels encoding one or more of visual and other characteristics of the cell, and Wherein identifying one or more of the plurality of categories for each cell of the received electronic form using the neural network model comprises: Each cell of the received electronic form is classified based on the predetermined number of channels encoding the one or more of the visual and other characteristics of the cell.
10. The method of claim 9, wherein the classification of each cell of the received electronic form is further based on the immediate surrounding context of the cell.
11. The method of claim 1 , wherein deriving the at least one table in the received electronic form based on the one or more categories identified for each cell of the received electronic form comprises: The at least one table is derived using a probabilistic graphical model.
12. The method of claim 11, wherein deriving the at least one table in the received spreadsheet using the probabilistic graphical model comprises: Estimate the probability that a corner occurs at the intersection of corresponding boundary lines; Calculate the spatial confidence score of each candidate table; as well as When the spatial confidence score of each candidate table is greater than a predetermined threshold, each candidate table is selected as a table. The method according to claim 12 , wherein the predetermined threshold is a hyperparameter optimized based on a validation dataset.
14. The method according to claim 1, further comprising: receiving a test data set, the test data set comprising a plurality of spreadsheets, each spreadsheet having at least one table; as well as The trained neural network is evaluated using the received test data set.
15. A system for improved table identification in an electronic form, the system comprising: a data storage device storing instructions for improved table identification in a spreadsheet; as well as A processor is configured to execute the instructions to implement the following method, the method comprising: receiving a data set comprising a plurality of electronic spreadsheets, wherein at least one electronic spreadsheet of the plurality of electronic spreadsheets comprises at least one table; processing the data set to produce an annotated data set comprising a plurality of annotated spreadsheets having identified tables, corners, and header rows; and training a neural network model using the annotated dataset, the neural network model being configured to output one or more of a plurality of categories for each cell of the received electronic form; receiving a spreadsheet including at least one table; for each cell of the received electronic table, identifying one or more of the plurality of categories using the neural network model, wherein the plurality of categories include upper left corner, upper right corner, lower left corner, lower right corner, title lower left corner, title lower right corner, and non-corner, the title lower left corner being a cell directly below a visual table title on a left table border, and the title lower right corner being a cell directly below a visual table title on a right table border; and At least one table in the received electronic form is derived based on the one or more categories identified for each cell of the received electronic form.
16. The system of claim 15, wherein for the one or more categories identified based on each cell of the received electronic form, deriving the at least one table in the received electronic form comprises: The at least one table is derived using a probabilistic graphical model, and wherein deriving the at least one table in the received spreadsheet using the probabilistic graphical model comprises: Estimate the probability that a corner occurs at the intersection of corresponding boundary lines; Calculating a spatial confidence score for each candidate list; and When the spatial confidence score of each candidate table is greater than a predetermined threshold, each candidate table is selected as a table.
17. A computer-readable storage device storing instructions which, when executed by a computer, cause the computer to implement a method for improved table identification in a spreadsheet, the method comprising: receiving a data set comprising a plurality of electronic spreadsheets, wherein at least one electronic spreadsheet of the plurality of electronic spreadsheets comprises at least one table; processing the data set to produce an annotated data set comprising a plurality of annotated spreadsheets having identified tables, corners, and header rows; as well as training a neural network model using the annotated dataset, the neural network model being configured to output one or more of a plurality of categories for each cell of the received electronic form; receiving a spreadsheet including at least one table; for each cell of the received electronic table, identifying one or more of the plurality of categories using the neural network model, wherein the plurality of categories include upper left corner, upper right corner, lower left corner, lower right corner, title lower left corner, title lower right corner, and non-corner, the title lower left corner being a cell directly below a visual table title on a left table border, and the title lower right corner being a cell directly below a visual table title on a right table border; as well as At least one table in the received electronic form is derived based on the one or more categories identified for each cell of the received electronic form.
Citation Information
Patent Citations
Method and apparatus for automatically processing data in a cell format
WO2012017056A1