A method, apparatus, electronic device and storage medium for processing tabular data
By identifying the table structure and disassemblying the molecular tables according to similarity, the problem of inconsistent format of the table data content is solved, and automated data processing and format uniformity are achieved, which improves processing efficiency and accuracy.
Patent Information
- Application Number
- CN202111222425.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-10-20
AI Technical Summary
In the prior art, the input format of the table data content is not unified, resulting in the inability to automatically perform subsequent citation, statistics and calculation processing, and manual cumbersome format unified processing is required.
By identifying the table structure, the table is split into multiple subtables according to the similarity between the table structure and the data content, and the data content of the subtable is normalized to generate a consistent display format.
It realizes the intelligent format unified processing of table data content, reduces manual participation, saves labor costs, and improves the efficiency and accuracy of data processing.
Smart Images

Figure CN114036914B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method, apparatus, electronic device and storage medium for processing tabular data. Background Art
[0002] A table supports input, output and display of data content, and can help users create various complex tabular documents. However, when the input formats of the data content in the table are not unified, subsequent processing such as reference, statistics and calculation cannot be performed based on the data content.
[0003] In the prior art, in order to utilize the data content with non-uniform input formats, it is necessary to manually perform format-unifying processing on the data content. However, this manual method of processing data content is cumbersome and inefficient. Summary of the Invention
[0004] The present invention provides a method, apparatus, electronic device and storage medium for processing tabular data, so as to solve the problem in the prior art that the format of the data content in the table can only be unified manually.
[0005] The present invention provides a method for processing tabular data, including:
[0006] Identifying the table structure, and splitting the area of the table according to the similarity between the table structure and the data content in the table to generate a plurality of sub-tables;
[0007] Performing normalization processing on the data content of the sub-tables to obtain processed sub-tables.
[0008] Preferably, the splitting the area of the table according to the similarity between the table structure and the data content in the table to generate a plurality of sub-tables includes:
[0009] Obtaining the first field composition and the first data content of the current table, as well as the second field composition and the second data content of the prior table;
[0010] Calculating the similarity value between the prior table and the current table;
[0011] If the similarity value is less than the first threshold, splitting the table into a prior sub-table and a current sub-table.
[0012] Preferably, the calculating the similarity value between the prior table and the current table includes:
[0013] Determining the first type value of the current table according to the first field composition and the first data content, and determining the second type value of the prior table according to the second field composition and the second data content;
[0014] Calculate the similarity value according to the first type value and the second type value.
[0015] Preferably, the calculating the similarity value according to the first type value and the second type value includes:
[0016] Use the ratio of the first type value and the second type value as the similarity value.
[0017] Preferably, it further includes:
[0018] Identify the blank areas in the table. If the number of blank rows / blank columns corresponding to the blank area is greater than a second threshold, split the table into multiple sub-tables.
[0019] Preferably, the normalizing the data content of the sub-table to obtain a processed sub-table includes:
[0020] Traverse all the data content of the sub-table according to the data fields of the row headers of the sub-table, and display the data content of each data content column in a format corresponding to the data fields of the row headers of the sub-table; or, traverse all the data content of the sub-table according to the data fields of the column headers of the sub-table, and display the data content of each data content row in a format corresponding to the data fields of the column headers of the sub-table.
[0021] Preferably, the normalizing the data content of the sub-table to obtain a processed sub-table includes:
[0022] Traverse each data content row of the sub-table, determine the first type of the display format of the data content of each data content row and the first number of data content corresponding to the first type of the display format, and display the data content based on the display format with the largest number of data content;
[0023] Or, traverse each data content column of the sub-table, determine the second type of the display format of the data content of each data content column and the second number of data content corresponding to the second type of the display format, and display the data content based on the display format with the largest number of data content.
[0024] Preferably, after normalizing the data content of the sub-table to obtain a processed sub-table, it further includes:
[0025] Check according to the data content in the processed sub-table;
[0026] Mark the data content that does not conform to the specification as abnormal data content, and output a prompt message for the abnormal data content.
[0027] The present invention also provides a table data processing device, including:
[0028] A first processing module, configured to identify a table structure, and split regions of the table according to the similarity between the table structure and data content in the table, so as to generate a plurality of sub-tables;
[0029] A second processing module, configured to perform normalization processing on data content of the sub-tables to obtain processed sub-tables.
[0030] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of any one of the above-mentioned table data processing methods are implemented.
[0031] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned table data processing methods are implemented.
[0032] The table data processing method, device, electronic device, and storage medium provided by the present invention split regions of a table according to the similarity between the table structure and data content in the table by identifying the table structure, so as to generate a plurality of sub-tables, and then perform normalization processing on data content of the sub-tables to obtain processed sub-tables. It can be seen that according to the similarity between the table structure and table data content, the present invention reorganizes data content with higher similarity into sub-tables, and then performs normalization processing on data content in the sub-tables, that is, making data content in each row of the sub-table be displayed in a consistent display format, or making data content in each column of the sub-table be displayed in a consistent display format, so as to realize intelligent format unification processing of data content in the table, without the need for manual participation in processing the display format of data content, saving labor costs, reducing the degree of manual participation, and ensuring the regularity and referability of data content in the sub-tables. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0034] Figure 1 is one of the flow diagrams of the table data processing method provided by the present invention;
[0035] Figure 2 is another flow diagram of the table data processing method provided by the present invention;
[0036] Figure 3 It is one of the processing schematic diagrams of the table data processing method provided by the present invention;
[0037] Figure 4 It is the second of the processing schematic diagrams of the table data processing method provided by the present invention;
[0038] Figure 5 It is the third of the processing schematic diagrams of the table data processing method provided by the present invention;
[0039] Figure 6 It is the fourth of the processing schematic diagrams of the table data processing method provided by the present invention;
[0040] Figure 7 It is the fifth of the processing schematic diagrams of the table data processing method provided by the present invention;
[0041] Figure 8 It is the sixth of the processing schematic diagrams of the table data processing method provided by the present invention;
[0042] Figure 9 It is the seventh of the processing schematic diagrams of the table data processing method provided by the present invention;
[0043] Figure 10 It is the eighth of the processing schematic diagrams of the table data processing method provided by the present invention;
[0044] Figure 11 It is the ninth of the processing schematic diagrams of the table data processing method provided by the present invention;
[0045] Figure 12 It is the tenth of the processing schematic diagrams of the table data processing method provided by the present invention;
[0046] Figure 13 It is the structural schematic diagram of the table data processing device provided by the present invention;
[0047] Figure 14 It is the structural schematic diagram of the electronic device provided by the present invention. Detailed implementation manners
[0048] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0049] A table processing method provided by the present invention is applicable to electronic devices, such as smartphones, tablets, computers, etc. Optionally, the execution subject of a table processing method provided by the present invention may be a table processing client running on an electronic device, such as a table processing application APP (Application). This method is applied to tables in table documents, text documents, and presentation documents.
[0050] Generally, a table has many cells arranged continuously in a matrix form, and data content can be displayed in each cell of the table. The vertical arrangement of cells is called "column", and the horizontal arrangement is called "row".
[0051] In one or more embodiments of this specification, the terms "horizontal", "vertical", "horizontal line", and "vertical line" are used. The direction of "horizontal" and "horizontal line" refers to the horizontal direction parallel to the horizontal ground with respect to the data content displayed on the display interface of the client, where the display interface is the one that can be recognized by the naked eye when the user looks straight ahead and can display the data content normally and accurately. The direction of "vertical" and "vertical line" refers to the vertical direction perpendicular to the horizontal ground with respect to the data content displayed on the display interface of the client, where the display interface is the one that can be recognized by the naked eye when the user looks straight ahead and can display the data content normally and accurately.
[0052] The data content is arranged in each cell that makes up a row and / or in each cell that makes up a column. With respect to the data content displayed on the display interface of the client, where the display interface is the one that can be recognized by the naked eye when the user looks straight ahead and can display the data content normally and accurately, a row is the area composed of horizontal lines; a column is the area composed of vertical lines; and the table formed by the intersection of rows and columns is a cell.
[0053] The following combines Figure 1 and Figure 2 to describe the table processing method of the present invention. As Figure 1 shown, a table processing method provided by the present invention includes:
[0054] Step 101: Identify the table structure, and split the area of the table according to the similarity between the table structure and the data content in the table to generate multiple sub-tables;
[0055] In this step, it should be noted that first, the table structure is obtained based on the table structure recognition algorithm. Optionally, the first field composition and the first data content of the current table, as well as the second field composition and the second data content of the previous table, are obtained based on the table structure recognition algorithm. Among them, the first field composition and the second field composition refer to the data content types corresponding to each field in the horizontal table or the data content types corresponding to each field in the vertical table. Here, the field can refer to each row or each column in the table. The data content type corresponding to each field is also the total cell types included in the data content in the table. The cell types can be: text, number, date, time, etc. The first data content and the second data content are the total cell sets included in the data content in the table.
[0056] In this step, it can be understood that the table may contain multiple sub-tables. Therefore, when performing format unification processing on the data content of the table, it is necessary to first split the sub-tables included in the table, and then perform normalization processing on each sub-table separately. Specifically, the table is split by the similarity of the data content in the table to generate multiple sub-tables. Among them, the similarity can be obtained by calculating the similarity value between different data contents in the table, which can indicate the similarity degree between different data contents in the table, so as to split the table according to whether the similarity value is less than the preset threshold. Among them, the area of the table refers to the area where the data content in the table is similar.
[0057] The similarity refers to the degree of similarity between the display format of the data content and the character composition of the data content. The area of the table refers to the data content row, or the data content column, or the set formed by multiple data content rows, or the set formed by multiple data content columns in the table. Splitting means disassembling the table based on the selected row as the benchmark, or disassembling it based on the selected column as the benchmark. It can also mean reorganizing multiple spaced data content rows into a new table, or reorganizing multiple spaced data content columns into a new table.
[0058] Step 102: Perform normalization processing on the data content of the sub-table to obtain the processed sub-table.
[0059] In this step, it should be noted that after splitting the table into multiple sub-tables, the data content of each sub-table is normalized respectively. Normalization refers to traversing all the data content of the sub-table according to the data fields of the row headers of the sub-table. The numeric fields include integer type, floating-point type, character type, date type, enumeration, and set. The data content of each data content column is displayed in a format corresponding to the data fields of the row headers of the sub-table; or, according to the data fields of the column headers of the sub-table, traverse all the data content of the sub-table, and display the data content of each data content row in a format corresponding to the data fields of the column headers of the sub-table; or, traverse each data content row of the sub-table, determine the first type of the display format of the data content of each data content row and the first number of data content of the corresponding first type of the display format, and display the data content based on the display format with the largest number of data content. For example, modify the minority data content to be consistent with the majority data content, and / or traverse each data content column of the sub-table, determine the second type of the display format of the data content of each data content column and the second number of data content of the corresponding second type of the display format, and display the data content based on the display format with the largest number of data content; or delete the data content in the sub-table that has no direct corresponding relationship with the row headers. After normalizing the data content of each sub-table, the unified processing of the data content format of the table is completed.
[0060] The table data processing method provided by the present invention splits the area of the table according to the table structure and the similarity of the data content in the table by identifying the table structure, generates multiple sub-tables, and then normalizes the data content of the sub-tables to obtain the processed sub-tables. It can be seen that according to the similarity of the table structure and the table data content, the present invention reorganizes the data content with higher similarity into sub-tables, and then normalizes the data content in the sub-tables, that is, makes the data content of each row in the sub-table be displayed in a consistent display format, or makes the data content of each column in the sub-table be displayed in a consistent display format, so as to realize the intelligent unified processing of the data content format of the table, without the need for manual participation in processing the display format of the data content, saving labor costs, reducing the degree of manual participation, and ensuring the regularity and referability of the data content in the sub-tables.
[0061] Based on the content of the above embodiments, in this embodiment, splitting the area of the table according to the table structure and the similarity of the data content in the table to generate multiple sub-tables includes:
[0062] Obtain the first field composition and the first data content of the current table, and the second field composition and the second data content of the prior table;
[0063] Calculate the similarity value between the prior table and the current table;
[0064] If the similarity value is less than the first threshold, the table is split into a prior sub-table and a current sub-table.
[0065] In this embodiment, the first field composition and the second field composition refer to the data content types corresponding to each field of the horizontal table, or the data content types corresponding to each field of the vertical table. Here, the field can refer to each row or each column in the table. The data content type corresponding to each field here is also the total cell types included in the data content in the table. The cell types can be: text, number, date, and time, etc. The first data content and the second data content are the total cell sets included in the data content in the table. The current table refers to the table area corresponding to the data content to be confirmed for similarity, which can be one row or one column, or multiple rows or multiple columns; the prior table refers to the table area corresponding to the data content for which similarity has been confirmed, which can be one row or one column, or multiple rows or multiple columns.
[0066] In this embodiment, the specific method for calculating the similarity value between the prior table and the current table is as follows: Determine the first type value of the data content of the current horizontal table according to the data content types corresponding to each field of the current horizontal table and the data content types corresponding to each field of the prior horizontal table, and determine the second type value of the data content of the current horizontal table according to the data content types corresponding to each field of the current horizontal table and the data content types corresponding to each field of the prior horizontal table; or determine the first type value of the data content of the current vertical table according to the data content types corresponding to each field of the current vertical table and the data content types corresponding to each field of the prior vertical table, and determine the second type value of the data content of the current vertical table according to the data content types corresponding to each field of the current vertical table and the data content types corresponding to each field of the prior vertical table; as Figure 12 shown, the data content types corresponding to each field of the current vertical table are Chinese, Chinese, Chinese, date, and blank respectively, and the corresponding first field composition is ([Chinese], [Chinese], [Chinese], [Date], [None]). Similarly, the data content types corresponding to each field of the prior vertical table are number + English, Chinese, number and English, Chinese, Chinese respectively, and the corresponding second field composition is ([Number, English], [Chinese], [Number, English], [Chinese], [Chinese]).
[0067] In this embodiment, according to the first field composition and the second field composition, determining the first type value of the first data content and determining the second type value of the second data content can be understood as determining the number of dissimilar column types and the number of comparable column types of the first data content and the second data content according to the first field composition and the second field composition. Furthermore, according to the ratio of the number of dissimilar column types to the number of comparable column types, determining the similarity value of the first data content and the second data content. When the similarity value is less than a preset threshold, the table is split into a prior sub-table and a current sub-table.
[0068] Based on the content of the above embodiment, in this embodiment, it further includes:
[0069] Identifying the blank area in the table. If the number of blank rows / blank columns corresponding to the blank area is greater than a second threshold, the table is split into multiple sub-tables.
[0070] In this embodiment, the present invention also provides a way to split a table, that is, identifying the blank area in the table. When the number of rows or columns corresponding to the blank area is greater than a preset number of rows or columns, the table is split into multiple sub-tables, and after splitting, each sub-table is normalized.
[0071] The blank area refers to a blank cell row or a blank cell column composed of blank cells without any data content in the cells. Exemplarily, the second threshold can be 3 rows, 6 rows, 5 columns, 8 columns, etc., and the present invention does not limit this.
[0072] Based on the content of the above embodiment, in an exemplary embodiment of the present invention, normalizing the data content of the sub-table to obtain the processed sub-table includes:
[0073] According to the data fields of the row headers of the sub-table, traversing all the data content of the sub-table, and displaying the data content of each data content column in a format corresponding to the data fields of the row headers of the sub-table; or, according to the data fields of the column headers of the sub-table, traversing all the data content of the sub-table, and displaying the data content of each data content row in a format corresponding to the data fields of the column headers of the sub-table.
[0074] In another exemplary embodiment of the present invention, normalizing the data content of the sub-table to obtain the processed sub-table includes:
[0075] Traversing each data content row of the sub-table, determining the first type of the display format of the data content of each data content row and the number of the first data content corresponding to the first type of the display format, and displaying the data content based on the display format with the largest number of data content;
[0076] Alternatively, traverse each data content column of the sub-table to determine the second type of the display format of the data content in each data content column and the second number of data contents of the corresponding display format, and display the data content based on the display format with the largest number of data contents.
[0077] That is, according to the number of types of field compositions, modify the minority types to be consistent with the majority types, and / or delete the data contents in the sub-table that have no direct corresponding relationship with the row titles.
[0078] In this embodiment, when the data fonts or sizes in the table are not uniform, the data contents are displayed in the format corresponding to the data fields of the sub-table row titles; if most of the data contents corresponding to a certain column in the table are numerical values, but there is a Chinese character content, it can be regarded as a writing error for this Chinese character content, and modify this Chinese character type to be consistent with the majority numerical types; for the invalid data contents that may exist in the table, they need to be deleted during the normalization process. For example, if the last row of the table is "Remarks Information" but the corresponding cell content is blank, then delete this row.
[0079] Based on the content of the above embodiments, in this embodiment, after normalizing the data contents of the sub-table and obtaining the processed sub-table, it further includes:
[0080] Proofread the data contents in the processed sub-table according to the pre-established proofreading instructions;
[0081] Mark the data contents that do not meet the specifications as abnormal data contents, and output the prompt information of the abnormal data contents.
[0082] In this embodiment, after normalizing each sub-table, the data contents in the processed sub-table can be proofread according to the pre-established proofreading instructions, mark the data contents that do not meet the specifications as abnormal data contents, and output the prompt information of the abnormal data contents. Among them, the preset proofreading instructions can be written according to the actual situation, and no detailed description is made here.
[0083] The following is illustrated by specific embodiments:
[0084] The first embodiment:
[0085] In this embodiment, as Figure 3 shown, it can be understood that the cell types in the table can be divided into: big title, row title, content, and others. Among them, the big titles include: High School Class 1-3 Semester Grade Statistics and High School Class 2-1 Semester Grade Statistics, the row titles include Cultural Course Grades, Extracurricular Activity Grades, Serial Number, Label, Chinese, Mathematics, English, Physical Education, and Social Practice, others include "Remarks", and the content is the cell content except for "big title", "row title", and "others".
[0086] In this embodiment, based on the table structure division method, a merged row list of table data is obtained; the merged row list includes a merged row area row interval, a merged row area column interval, and a merged row type; the merged row type includes the main title, row title, and content of the table data.
[0087] In this embodiment, it should be noted that since the layout of the user's recorded or written table is relatively random, the data analysis function cannot be directly adopted. Therefore, the present invention needs to regenerate the initial table data in the form of a data table. Specifically, first, a merged row list of table data is obtained by using a table structure division algorithm. As Figure 4 shown, it can be seen from Figure 4 that it includes a merged row area row interval, a merged row area column interval, and a merged row type, and the merged row type includes the main title, row title, and content of the table data.
[0088] In this embodiment, each column of the table data is traversed, and based on the method of tiling merged cells, the cells with the merged row type of row title are merged to obtain a field name list of the table data.
[0089] In this step, it should be noted that in most cases, the table data to be generated into a data table will include multiple sub-tables. When the table data to be generated into a data table is a single table data, each column of data in the table is traversed, and based on the method of tiling merged cells, the merged cells with the merged row type of row title are tiled after merging to obtain the corresponding field names. If the current merged row type is not a row title, the default row title_index name is output. After traversing the cell data corresponding to each column of the table, a field name list of the table data is obtained. For example Figure 3 in, based on the method of tiling merged cells, the merged cells with the merged row type of row title are tiled after merging to obtain the corresponding field names including: Cultural Course Score_Mathematics, Cultural Course Score_Chinese, Extracurricular Activity Score_Sports, and Extracurricular Activity Score_Social Practice.
[0090] In this embodiment, the blank columns in the table data are identified, and based on the blank columns and the merged rows with the merged row types of row title and content, the merged row list is divided to obtain a list of "row title + content" or "content" as the sub-table list divided by the row title. As Figure 5 shown. After obtaining the sub-table list of the table data, each column of the sub-table list is traversed, and the cells with the merged row type of row title are merged and tiled to obtain a tiled list of merged cells of the row title of the table data. As Figure 6 shown, and then the field name list of the table data can be obtained according to Figure 6 , as Figure 7 shown.
[0091] In this embodiment, according to the field name list and the cells with the merged row type as content, the table data is reorganized in the form of a data table to generate the data table corresponding to the table data.
[0092] In this embodiment, it should be noted that when the table data to be generated into a data table is single table data, after obtaining the field name list, the field name list and the cells with the merged row type as content are merged to realize reorganizing the table data in the form of a data table and generating the data table corresponding to the table data. When the table data to be generated into a data table contains at least two sub-table data, first determine the type of each column in the content merged row of the sub-table list. Specifically, traverse each row in the content merged row of the sub-table list, count the number of occurrences of each cell type in each column, and then determine the type of each column in the content merged row of the sub-table list according to the number of occurrences of each cell type in each column and the number of rows in the content merged row. After determining the type of each column in the content merged row of the sub-table list, calculate the dissimilarity between the content merged rows according to the type of each column in the content merged row of the sub-table list, so as to divide the sub-table list according to the dissimilarity between the content merged rows to obtain the content similarity list of the table data, as Figure 8 shown. After obtaining the content similarity list, traverse the content similarity list of the table data, merge the content merged rows in the content similarity list with the field name list to generate the data table corresponding to the table data, as Figure 9 and Figure 10 shown, thus realizing the logic of identifying and splitting the data table according to the content.
[0093] Based on the content of the above embodiment, in this embodiment, based on the table structure division method, obtain the merged row list of the table data, including:
[0094] Based on the table structure division method, determine the row interval of the merged row area and the column interval of the merged row area of the table data;
[0095] According to the row interval of the merged row area and the column interval of the merged row area of the table data, determine the merged row type of the table data, so as to generate the merged row list of the table data based on the row interval of the merged row area, the column interval of the merged row area and the merged row type of the table data.
[0096] In this embodiment, it should be noted that the table structure is divided by using a preset table structure division algorithm. First, the row interval and column interval of the merged row area of the table data are determined. Then, according to the row interval and column interval of the merged row area of the table data, the type to which the current merged row belongs is judged. Finally, the above data is summarized to generate a merged row list of the table data. The merged row list obtained by the table structure division algorithm of the present invention provides a guarantee for the subsequent reorganization of the table data, improving the efficiency and reducing the operation time for the user during the process of structurally standardizing and adjusting the table data.
[0097] Based on the content of the above embodiment, in this embodiment, before traversing each column of the table data and merging the cells with the merged row type of row title in a tiled manner based on merged cells to obtain a list of field names of the table data, it further includes:
[0098] Dividing the merged row list according to the merged row type and the blank columns in the table data to obtain a sub-table list of the table data.
[0099] Among them, dividing the merged row list according to the merged row type and the blank columns in the table data to obtain a sub-table list of the table data includes:
[0100] Obtaining the blank columns in the table data, and dividing the merged row list according to the blank columns and the merged rows with the merged row type of row title and content to obtain a sub-table list of the table data.
[0101] In this embodiment, it should be noted that when the table data of the data table to be generated contains at least two sub-table data, first identify the blank columns in the table data, and divide the merged row list according to the blank columns and the merged rows with the merged row type of row title and content to obtain a list of "row title + content" or "content" as the sub-table list divided by the row title. By dividing the merged row list according to the blank columns and the merged rows with the merged row type of row title and content, the present invention can accurately identify each sub-table data contained in the table, providing a basis for the subsequent splitting of the data table.
[0102] Based on the content of the above embodiment, in this embodiment, traversing each column of the table data and merging the cells with the merged row type of row title in a tiled manner based on merged cells to obtain a list of field names of the table data includes:
[0103] Traversing each column of the sub-table list, merging the cells with the merged row type of row title and then tiling to obtain a tiled list of merged cells of the row titles of the table data, and obtaining a list of field names of the table data according to the tiled list of merged cells of the row titles.
[0104] Based on the content of the above embodiment, in this embodiment, it further includes:
[0105] Determine the type of each column in the content merging rows of the sub-table list;
[0106] According to the type of each column in the content merging rows of the sub-table list, calculate the dissimilarity between the content merging rows, so as to divide the sub-table list according to the dissimilarity between the content merging rows, and obtain a content similarity list of the table data.
[0107] In this embodiment, it should be noted that by traversing each row of the content merging rows in the sub-table list, counting the number of occurrences of each cell type in each column, and then determining the type of each column in the content merging rows of the sub-table list according to the number of occurrences of each cell type in each column and the number of rows of the content merging rows. Optionally, traverse each row of the content merging rows in the sub-table list, count the number of occurrences of each cell type in each column, and summarize all cell types whose (number of occurrences of cell type / number of rows) is greater than the preset threshold to obtain the type of each column in the content merging rows of the sub-table list. The cell types include: Chinese, English, numbers, time, and date.
[0108] In this embodiment, by traversing the type of each column in the content merging rows of the sub-table list, counting the number of comparable column types and the number of dissimilar column types of the first content merging row and the second content merging row included in the sub-table list, and then calculating the dissimilarity between the first content merging row and the second content merging row according to the number of comparable column types and the number of dissimilar column types of the first content merging row and the second content merging row included in the sub-table list, so as to divide the sub-table list according to the dissimilarity between the content merging rows, and obtain a content similarity list of the table data. Optionally, the number of comparable column types and the number of dissimilar column types of the first content merging row and the second content merging row included in the sub-table list are counted in the following manner:
[0109] 1. If any one of the comparison column types of the first content merging row and the second content merging row is empty, then add +0 to the number of comparable column types and the number of dissimilar column types of the first content merging row and the second content merging row respectively;
[0110] 2. If the comparison column types of the first content merging row and the second content merging row are the same, and the comparison column types only include one of time, date, and number, then add +1 to the number of comparable column types of the first content merging row and the second content merging row, and add +0 to the number of dissimilar column types;
[0111] 3. If the comparison column types of the first content merging row and the second content merging row are the same, and both are time or date, then add +1 to the number of comparable column types of the first content merging row and the second content merging row, and add +0 to the number of dissimilar column types;
[0112] 4. If the cells of the first content merged row and the second content merged row both contain Chinese, then add 0 to the number of comparable column types and the number of dissimilar column types of the first content merged row and the second content merged row respectively;
[0113] 5. If the comparison column types of the first content merged row and the second content merged row are different, then calculate the difference value of the comparison column types of the first content merged row and the second content merged row according to the following formula. If the difference value is greater than 0.5, then add 1 to the number of comparable column types and the number of dissimilar column types of the first content merged row and the second content merged row respectively. Otherwise, add 1 to the number of comparable column types of the first content merged row and add 0 to the number of dissimilar column types; the formula is:
[0114] Difference value = (the number of comparison column types of the second content merged row that do not appear in the comparison column types of the first content merged row + the number of comparison column types of the first content merged row that do not appear in the comparison column types of the second content merged row) / the maximum number of comparison column types.
[0115] In this embodiment, optionally, if the dissimilarity between content merged rows is greater than a preset threshold, and / or the interval between content merged rows is greater than a preset number of rows, then divide the sub-table list to obtain a content similarity list of the table data. It can be seen that the embodiment of the present invention calculates whether to split the data table by using the similarity of row content, increasing the possibility of processing complex scenarios while reducing the human participation.
[0116] Based on the content of the above embodiment, in this embodiment, according to the field name list and the cells with the merged row type as content, reorganize the table data in the form of a data table to generate a data table corresponding to the table data, including:
[0117] Traverse the content similarity list of the table data, and merge the content merged rows in the content similarity list with the field name list to generate a data table corresponding to the table data.
[0118] Based on the content of the above embodiment, in this embodiment, determine the type of each column of the content merged row in the sub-table list, including:
[0119] Traverse each row of the content merged row in the sub-table list, and count the number of times each cell type appears in each column; the cell types include: Chinese, English, numbers, time, and date;
[0120] The cell type with the number of times each cell type appears in each column / the number of rows of the content merged row greater than 0.2 is used as the type of each column of the content merged row.
[0121] In this embodiment, it should be noted that by traversing each row of the content merged rows in the sub-table list, counting the number of occurrences of each cell type in each column, and then determining the type of each column of the content merged rows in the sub-table list based on the number of occurrences of each cell type in each column and the number of rows of the content merged rows. Optionally, traverse each row of the content merged rows in the sub-table list, count the number of occurrences of each cell type in each column, and summarize all types where (the number of occurrences of the cell type / the number of rows) is greater than 0.2 to obtain the type of each column of the content merged rows in the sub-table list. For example, if the content merged row is [2-15], after traversing the 14 cells corresponding to the first column, assuming that the cell types of all 14 rows are numbers, then the cell type of number appears 14 times and the number of rows is 14. At this time, the calculation result according to the above formula is 1, which is greater than the threshold of 0.2, so the type of the first column of the content merged row [2-15] is regarded as Number. Another example, if the content merged row is [2-15], after traversing the 14 cells corresponding to the first column, assuming that the cell types of 8 rows are numbers and the cell types of 6 rows are English, then the cell type of number appears 8 times and the number of rows is 14. At this time, the calculation result according to the above formula is 0.57, which is greater than the threshold of 0.2. At the same time, the cell type of English appears 6 times and the number of rows is 14. At this time, the calculation result according to the above formula is 0.28, which is greater than the threshold of 0.2, so the type of the first column of the content merged row [2-15] is Number and English.
[0122] Based on the content of the above embodiment, in this embodiment, according to the type of each column of the content merged rows in the sub-table list, calculate the dissimilarity between the content merged rows, including:
[0123] Traverse the type of each column of the content merged rows in the sub-table list, and count the number of comparable column types and the number of dissimilar column types of the first content merged row and the second content merged row included in the sub-table list;
[0124] According to the number of comparable column types and the number of dissimilar column types of the first content merged row and the second content merged row included in the sub-table list, calculate the dissimilarity between the first content merged row and the second content merged row.
[0125] In this embodiment, after obtaining the number of comparable column types and the number of dissimilar column types of the first content merged row and the second content merged row, calculate the dissimilarity between the first content merged row and the second content merged row based on the following formula:
[0126]
[0127] Based on the content of the above embodiment, in this embodiment, according to the dissimilarity between the content merged rows, divide the sub-table list to obtain a content similarity list of the tabular data, including:
[0128] If the dissimilarity between content merged rows is greater than a preset threshold, and / or the interval between content merged rows is greater than a preset number of rows, then partition the sub-table list to obtain a list of content similarities of the tabular data.
[0129] In this embodiment, optionally, if the dissimilarity between content merged rows is greater than 0.5, and / or the interval between content merged rows is greater than 4 rows, then partition the sub-table list to obtain a list of content similarities of the tabular data.
[0130] Second Embodiment:
[0131] In this embodiment, first obtain tabular data, and then use a tabular structure partitioning algorithm to obtain the recognition types of the merged row list, where the types include: big title, row title, content, others.
[0132] In this embodiment, obtain the sub-table list partitioned using the row title. Specifically, use blank columns and row titles to partition the table area, and the list of "row title + content" or "content" obtained is the sub-table list partitioned using the row title.
[0133] In this embodiment, traverse the sub-table list partitioned by the row title, and traverse the sub-tables to obtain the list of field names of the row title. Specifically, flatten the merged cells of the row title merged rows to obtain the field names. If there is no row title, give the default row title column_index name, so as to obtain the list of field names.
[0134] In this embodiment, collect the content merged rows with numbers or time, and traverse the content merged rows. Specifically, collect the merged rows whose cells contain dates or numbers, and then obtain the list of content similarity tables. Specifically, traverse the content merged rows with numbers or time, and determine that the dissimilarity between the current row and all the previous merged rows is greater than 0.5, and / or the previous interval is 4 rows, then split them into 2 tables to obtain the content similarity list.
[0135] In this embodiment, obtain the list of tables that conform to the data table structure, and traverse the content similarity list, and merge the content merged row cells and the list of field names of the content similarity list to obtain the list of data tables.
[0136] Thus, it can be seen that a method for generating a data table based on tabular data provided by an embodiment of the present invention eliminates non-standard content by collecting and merging content containing numbers or time, and splits the table according to the similarity of the content, realizing the logic of identifying and splitting the data table according to the content.
[0137] Third Embodiment:
[0138] In this embodiment, the method for calculating the dissimilarity between content merged rows is specifically as follows:
[0139] 1. Traverse all the content to merge lines, and use Method 1 to collect the types of each column in the merged lines.
[0140] Method 1: Traverse each line in the merged lines, count the number of occurrences of each cell type in each column, and summarize all the types where (the number of occurrences of the cell type / the number of lines) is greater than 0.2 to obtain the type of the column.
[0141] 2. Traverse the types of columns in the merged lines, and use Method 2 to compare the minimum similarity value between the current merged line and all the previous merged lines of this table. If this value is greater than 0.5, and / or is 4 rows apart from the previous one, then split it into 2 tables.
[0142] Method 2: To compare the similarity between the current merged line and the previous merged lines in the same column, two values need to be counted: the number of comparable column types and the number of dissimilar column types. Use Method 3 to calculate the dissimilarity of the two column types = the number of dissimilar column types / the number of comparable column types.
[0143] Method 3:
[0144] 1) If any one of the comparison column types of the first content merged line and the second content merged line is empty, then add 0 to the number of comparable column types and the number of dissimilar column types of the first content merged line and the second content merged line respectively;
[0145] 2) If the comparison column types of the first content merged line and the second content merged line are the same, and the comparison column types only contain one of time, date, and number, then add 1 to the number of comparable column types of the first content merged line and the second content merged line, and add 0 to the number of dissimilar column types;
[0146] 3) If the comparison column types of the first content merged line and the second content merged line are the same, and both are time or date, then add 1 to the number of comparable column types of the first content merged line and the second content merged line, and add 0 to the number of dissimilar column types;
[0147] 4) If the cells of the first content merged line and the second content merged line both contain Chinese, then add 0 to the number of comparable column types and the number of dissimilar column types of the first content merged line and the second content merged line respectively;
[0148] 5) If the comparison column types of the first content merged line and the second content merged line are different, then calculate the difference value of the comparison column types of the first content merged line and the second content merged line according to the following formula. If the difference value is greater than 0.5, then add 1 to the number of comparable column types and the number of dissimilar column types of the first content merged line and the second content merged line respectively, otherwise add 1 to the number of comparable column types of the first content merged line and the second content merged line, and add 0 to the number of dissimilar column types; The formula is:
[0149] Difference value = (Number of comparison column types in the second content merged row that do not appear in the comparison column types of the first content merged row + Number of comparison column types in the first content merged row that do not appear in the comparison column types of the second content merged row) / Maximum number of comparison column types.
[0150] For example, as Figure 11 and Figure 12 shown, content merged rows [2 - 4], [5 - 8] belonging to content containing numbers or time are obtained through Method 1, and the comparison of the merged row content is through Method 3 to obtain the number of comparable column types and the number of dissimilar column types for each column type. For example, the comparison column types of the first column of content merged rows [2 - 4] and [5 - 8] are [Number, English] and [Chinese] respectively, that is, the comparison column type of the first column of content merged row [2 - 4] contains numbers and English, which is different from the comparison column type (Chinese) of the first column of content merged row [5 - 8], then 5) in Method 3 above is adopted. Specifically, first calculate the difference value of the comparison column types of content merged rows [2 - 4] and [5 - 8], Difference value = (Number of comparable column types in content merged row [5 - 8] that do not appear in the comparable column types of content merged row [2 - 4] + Number of comparable column types in content merged row [2 - 4] that do not appear in the comparable column types of content merged row [5 - 8]) / Maximum value of the length of comparable column types, that is, Difference value = (2 + 1) / 2 = 1.5, which is greater than the threshold 0.5. Then, when counting the number of comparable column types and the number of dissimilar column types of content merged rows [2 - 4] and [5 - 8], the number of comparable column types and the number of dissimilar column types of content merged rows [2 - 4] and [5 - 8] are each incremented by 1. Another example, the comparison column types of the second column of content merged rows [2 - 4] and [5 - 8] are both [Chinese], that is, the comparison column types of the second column of content merged rows [2 - 4] and [5 - 8] are both Chinese, then 4) in Method 3 above is adopted, and the number of comparable column types and the number of dissimilar column types of content merged rows [2 - 4] and [5 - 8] are each incremented by 0. The comparison column types of the third and fourth columns of content merged rows [2 - 4] and [5 - 8] are all different. Similarly, 5) in Method 3 above is used to calculate the difference degree, and then the number of comparable column types and the number of dissimilar column types are counted according to the difference degree, which will not be elaborated here. After traversing all column types of the content merged row, the number of comparable column types is 3 and the number of dissimilar column types is 3. Then, the dissimilarity degree of content merged rows [2 - 4] and [5 - 8] = Number of dissimilar column types / Number of comparable column types = 3 / 3 = 1, which is greater than 0.5, and the Figure 11 table in
[0151] The table data processing device provided by the present invention will be described below. The table data processing device described below can be correspondingly referred to the table data processing method described above.
[0152] As Figure 13 shown, a table data processing device provided by the present invention includes:
[0153] A first processing module 1, configured to identify a table structure, and split the area of the table according to the similarity between the table structure and the data content in the table to generate a plurality of sub-tables;
[0154] A second processing module 2, configured to perform normalization processing on the data content of the sub-tables to obtain processed sub-tables.
[0155] In this embodiment, it should be noted that first, the table structure is obtained based on a table structure recognition algorithm. Optionally, the first field composition and the first data content of the current table, as well as the second field composition and the second data content of the previous table are obtained based on the table structure recognition algorithm. Among them, the first field composition and the second field composition refer to the data content types corresponding to each field in the horizontal table or the data content types corresponding to each field in the vertical table. Here, the field can refer to each row or each column in the table. The data content type corresponding to each field is also the entire cell type included in the data content in the table. The cell type can be: text, number, date, time, etc. The first data content and the second data content are the entire cell sets included in the data content in the table.
[0156] In this embodiment, it can be understood that the table may contain multiple sub-tables. Therefore, when performing format unification processing on the data content of the table, it is necessary to first split to obtain the sub-tables included in the table, and then perform normalization processing on each sub-table separately. Specifically, the table is split through the similarity of the data content in the table to generate a plurality of sub-tables. Among them, the similarity can be obtained by calculating the similarity value between different data contents in the table, which can indicate the similarity degree between different data contents in the table, so as to split the table according to whether the similarity value is less than a preset threshold. Among them, the area of the table refers to the area where the data content in the table is similar.
[0157] The similarity refers to the degree of similarity between the display format of the data content and the character composition of the data content. The area of the table refers to the data content row of the table, or the data content column, or the set formed by multiple data content rows, or the set formed by multiple data content columns. Splitting means splitting the table based on a selected row as a benchmark, or splitting based on a selected column as a benchmark, and can also refer to reorganizing multiple spaced data content rows into a new table, or reorganizing multiple spaced data content columns into a new table.
[0158] In this embodiment, it should be noted that after splitting the table into multiple sub-tables, the data content of each sub-table is normalized respectively. The normalization process refers to traversing all the data content of the sub-table according to the data fields of the row headers of the sub-table. The digital fields include integer type, floating-point type, character type, date type, enumeration, and set. The data content of each data content column is displayed in a format corresponding to the data fields of the row headers of the sub-table; or, according to the data fields of the column headers of the sub-table, traverse all the data content of the sub-table, and display the data content of each data content row in a format corresponding to the data fields of the column headers of the sub-table; or, traverse each data content row of the sub-table, determine the first type of the display format of the data content of each data content row and the first number of data content corresponding to the first type of the display format, and display the data content based on the display format with the largest number of data content. For example, modify the minority data content to be the same as the majority data content, and / or traverse each data content column of the sub-table, determine the second type of the display format of the data content of each data content column and the second number of data content corresponding to the second type of the display format, and display the data content based on the display format with the largest number of data content; or delete the data content in the sub-table that has no direct corresponding relationship with the row headers. After normalizing the data content of each sub-table, the unified processing of the data content format of the table is completed.
[0159] The table data processing device provided by the present invention splits the area of the table by identifying the table structure and according to the similarity between the table structure and the data content in the table to generate multiple sub-tables, and then normalizes the data content of the sub-tables to obtain the processed sub-tables. It can be seen that according to the similarity between the table structure and the table data content, the present invention reorganizes the data content with higher similarity into sub-tables, and then normalizes the data content in the sub-tables, that is, makes the data content of each row in the sub-table be displayed in a consistent display format, or makes the data content of each column in the sub-table be displayed in a consistent display format, so as to realize the intelligent unified processing of the data content format of the table, without the need for manual participation in processing the display format of the data content, saving labor costs, reducing the degree of manual participation, and ensuring the regularity and referability of the data content in the sub-tables.
[0160] Based on the content of the above embodiment, in this embodiment, the first processing module is specifically used for:
[0161] Obtain the first field composition and the first data content of the current table, as well as the second field composition and the second data content of the prior table;
[0162] Calculate the similarity value between the prior table and the current table;
[0163] If the similarity value is less than the first threshold, the table is split into a prior sub-table and a current sub-table.
[0164] Based on the content of the above embodiments, in this embodiment, the first processing module is further specifically configured to:
[0165] Determine the first type value of the current table according to the first field composition and the first data content, and determine the second type value of the prior table according to the second field composition and the second data content;
[0166] Calculate the similarity value according to the first type value and the second type value.
[0167] Based on the content of the above embodiments, in this embodiment, the first processing module is further specifically configured to:
[0168] Use the ratio of the first type value and the second type value as the similarity value.
[0169] Based on the content of the above embodiments, in this embodiment, the first processing module is further specifically configured to:
[0170] Identify the blank area in the table. If the number of blank rows / blank columns corresponding to the blank area is greater than the second threshold, the table is split into multiple sub-tables.
[0171] Based on the content of the above embodiments, in this embodiment, the second processing module is specifically configured to:
[0172] Traverse all the data content of the sub-table according to the data field of the row header of the sub-table, and display the data content of each data content column in a format corresponding to the data field of the row header of the sub-table; or, traverse all the data content of the sub-table according to the data field of the column header of the sub-table, and display the data content of each data content row in a format corresponding to the data field of the column header of the sub-table.
[0173] Based on the content of the above embodiments, in this embodiment, the second processing module is specifically configured to:
[0174] Traverse each data content row of the sub-table, determine the first type of the display format of the data content of each data content row and the number of the first data content corresponding to the first type of the display format, and display the data content based on the display format with the largest number of data content;
[0175] Or, traverse each data content column of the sub-table, determine the second type of the display format of the data content of each data content column and the number of the second data content corresponding to the second type of the display format, and display the data content based on the display format with the largest number of data content.
[0176] Based on the content of the above embodiments, in this embodiment, after obtaining the processed sub-table, the second processing module is further specifically configured to:
[0177] Proofread the data content in the processed sub-table;
[0178] Mark the data content that does not meet the specification as abnormal data content and output a prompt message for the abnormal data content.
[0179] Figure 14 An example of the physical structure diagram of an electronic device is shown as Figure 14 shown. The electronic device may include: a processor 1410, a communication interface 1420, a memory 1430, a communication bus 1440, and a computer program stored on the memory and executable on the processor. Among them, the processor 1410, the communication interface 1420, and the memory 1430 communicate with each other through the communication bus 1440. The processor 1410 can call the logical instructions in the memory 1430 to execute the table data processing method, which includes: identifying the table structure, splitting the area of the table according to the similarity between the table structure and the data content in the table to generate multiple sub-tables; performing normalization processing on the data content of the sub-tables to obtain the processed sub-tables.
[0180] A memory, a processor, and a computer program stored on the memory and executable on the processor.
[0181] In addition, when the logical instructions in the above-mentioned memory 1430 can be implemented in the form of software function units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0182] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the table data processing method provided by each of the above methods. The method includes: identifying a table structure, and splitting the area of the table into multiple sub-tables according to the similarity between the table structure and the data content in the table; and performing normalization processing on the data content of the sub-tables to obtain processed sub-tables.
[0183] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the table data processing method provided by each of the above methods. The method includes: identifying a table structure, and splitting the area of the table into multiple sub-tables according to the similarity between the table structure and the data content in the table; and performing normalization processing on the data content of the sub-tables to obtain processed sub-tables.
[0184] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0185] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0186] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A method for processing tabular data, characterized in that, Including: Identifying the table structure, and splitting the area of the table into multiple sub-tables according to the similarity between the table structure and the data content in the table; wherein, the identifying the table structure includes: obtaining a merged row list of table data based on a table structure division method; the merged row list includes a merged row area row interval, a merged row area column interval, and a merged row type. Normalizing the data content of the sub-tables to obtain processed sub-tables.
2. The method for processing tabular data according to claim 1, characterized in that, The merged row type includes the main title, row title, and content of the table. The splitting the area of the table into multiple sub-tables according to the similarity between the table structure and the data content in the table includes: Traversing each column of the table data, and merging the cells with the merged row type of row title based on the method of tiling merged cells to obtain a field name list of the table data. Reorganizing the table data in the form of a data table according to the field name list and the cells with the merged row type of content to generate a data table corresponding to the table data.
3. The method for processing tabular data according to claim 2, characterized in that, The splitting the area of the table into multiple sub-tables according to the similarity between the table structure and the data content in the table includes: Obtaining the first field composition and the first data content of the current table, and the second field composition and the second data content of the previous table. Calculating the similarity value between the previous table and the current table. If the similarity value is less than a first threshold, splitting the table into a previous sub-table and a current sub-table. Wherein, the calculating the similarity value between the previous table and the current table includes: Determining a first type value of the current table according to the first field composition and the first data content, and determining a second type value of the previous table according to the second field composition and the second data content. Calculating the similarity value according to the first type value and the second type value. Alternatively, determining the type of each column of the content merged rows in the sub-table list; calculating the dissimilarity between the content merged rows according to the type of each column of the content merged rows in the sub-table list, and dividing the sub-table list according to the dissimilarity between the content merged rows to obtain multiple sub-tables.
4. The method for processing tabular data according to claim 3, characterized in that, The calculating the similarity value according to the first type value and the second type value includes: Taking the ratio of the first type value and the second type value as the similarity value.
5. The method for processing tabular data according to claim 1, characterized in that, Also including: Identifying the blank area in the table. If the number of blank rows / blank columns corresponding to the blank area is greater than a second threshold, splitting the table into multiple sub-tables. Alternatively, dividing the merged row list according to the merged row type and the blank columns in the table data, and splitting the table into multiple sub-tables; including: obtaining the blank columns in the table data, and dividing the merged row list according to the blank columns and the merged rows with the merged row type of row title and content to obtain a sub-table list of the table data.
6. The method for processing tabular data according to claim 1, characterized in that, The normalizing the data content of the sub-tables to obtain processed sub-tables includes: Traverse all the data content of the sub-table according to the data fields of the row headers of the sub-table, and display the data content of each data content column in a format corresponding to the data fields of the row headers of the sub-table; or, traverse all the data content of the sub-table according to the data fields of the column headers of the sub-table, and display the data content of each data content row in a format corresponding to the data fields of the column headers of the sub-table.
7. The method for processing tabular data according to claim 1, characterized in that, The normalization processing of the data content of the sub-table to obtain a processed sub-table includes: Traverse each data content row of the sub-table, determine the first type of the display format of the data content of each data content row and the first number of data content corresponding to the first type of the display format, and display the data content based on the display format with the largest number of data content; Or, traverse each data content column of the sub-table, determine the second type of the display format of the data content of each data content column and the second number of data content corresponding to the second type of the display format, and display the data content based on the display format with the largest number of data content.
8. The method for processing tabular data according to claim 6 or 7, characterized in that, After the normalization processing of the data content of the sub-table to obtain a processed sub-table, it further includes: Proofread the data content in the processed sub-table; Mark the data content that does not meet the specification as abnormal data content, and output the prompt information of the abnormal data content.
9. A device for processing tabular data, characterized in that, It includes: A first processing module, configured to identify the table structure, and split the area of the table into multiple sub-tables according to the table structure and the similarity of the data content in the table. The identification of the table structure includes: obtaining a merged row list of the table data based on the table structure division method; the merged row list includes a merged row area row interval, a merged row area column interval, and a merged row type; A second processing module, configured to perform normalization processing on the data content of the sub-table to obtain a processed sub-table.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the table data processing method according to any one of claims 1 to 8.
11. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the steps of the table data processing method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Spreadsheet-oriented table clone automatic detection method and electronic device
CN112364786A
Method and device for identifying table direction, and electronic equipment
CN112560545A