Method for identifying complex table in document
By identifying and analyzing the similarity of table areas in PDF financial reports and combining similarity transferability assumptions, the problem of misjudgment of cross-page tables and merged cells is solved, efficient and accurate table data extraction and reconstruction are achieved, and the degree of automation and accuracy of data processing is improved.
Patent Information
- Application Number
- CN202510498402.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
In the prior art, when dealing with complex tables, especially spreadsheets and merging cells, there are problems of identifying misjudgment and data loss, which affects the accuracy and completeness of the table data.
Through image processing and feature analysis based on PDF financial reports, the table area is identified and the similarity of adjacent tables is calculated. Combined with the similarity transferability assumption, span page recognition is performed, and blank cells and merged cells are type determination and range recognition are performed to complete the identification and reconstruction of the complete table.
It improves the accuracy and completeness of cross-page table recognition, ensures the continuity and accuracy of tabular data, reduces manual intervention, and improves the quality and efficiency of data extraction.
Smart Images

Figure CN120014662A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of table recognition, and in particular to a method for recognizing complex tables in documents. Background Art
[0002] With the rapid development of information technology and artificial intelligence, document image processing and analysis has become an important research direction. This field covers technologies such as image recognition, optical character recognition and natural language processing, and is applied to automated data extraction, information understanding and processing in different scenarios. With the widespread use of various electronic documents, how to efficiently and accurately extract structured information from these documents, especially the complex tabular data in financial reports, has become a research focus. Especially in common financial reports such as PDF financial reports, tables not only have complex layouts, but also contain special cases such as merged cells and cross-page tables, which poses a great challenge to automated recognition.
[0003] Although existing document table recognition methods have made some progress in processing standard table layouts, they still have many shortcomings when facing complex tables, especially cross-page tables and merged cells. Most existing methods rely on fixed regular pattern matching, which is difficult to adapt to the diversity of different table styles and layouts. When identifying blank cells, especially when blank cells are part of merged cells, existing methods often have difficulty accurately distinguishing between true blank cells and those that belong to the merged area, which directly leads to data parsing errors. In addition, for cross-page tables, since the table content may be split on multiple pages, the current cross-page recognition method lacks the ability to accurately transmit the cross-page information of the table, which may lead to data loss or inconsistency, thereby affecting the final table data reconstruction and analysis results.
[0004] In the process of table recognition, the lack of blank cell type determination and cross-page table analysis will lead to the following major problems: 1. Misjudgment of blank cells: If a blank cell is mistakenly determined to be part of a merged cell, or vice versa, the merged state and data structure of the table may not be correctly restored, affecting the accuracy of the final data. For example, the content of some merged cells may not be correctly transferred to the corresponding cell location, resulting in data loss or errors. 2. Incorrect splicing of cross-page tables: In the process of processing cross-page tables, if the cross-page information is not correctly transferred, it may cause misalignment of table content between pages, especially when the merged area of multi-page tables is not correctly connected, the contextual relationship of the data will be lost, which will affect the analysis of the table and subsequent decision support. Summary of the invention
[0005] In view of the above problems existing in the prior art, the present application provides a method for identifying complex tables in a document.
[0006] The present disclosure provides a method for identifying a complex table in a document, comprising the following steps:
[0007] S1. Based on the PDF financial report, multiple groups of image files are obtained, and the table area in each group of image files is identified to obtain the table data;
[0008] S2. Based on the table data, analyze the similarity between the table areas, combine the similarity transitivity hypothesis, perform cross-page recognition, and complete the complete table recognition task;
[0009] S3. Based on S2, preliminarily determine the relevant data of the complete table, and determine the type of blank cells based on the form type of the merged cells, and identify the range of the merged cells, and further improve the relevant data of the complete table. Based on the improved relevant data of the complete table, perform the restoration task on the table.
[0010] Optionally, the specific steps of S1 include:
[0011] S11. According to the company's finance department, the PDF financial report is obtained in advance, and the PDF financial report document is converted into multiple sets of image files, and high-resolution images and text layer information are extracted from the multiple sets of image files, wherein the text layer information includes but is not limited to text content, text position, font and font size;
[0012] S12. Use the edge detection kernel to perform convolution operation on the high-resolution image to obtain the edge strength of each pixel , the edge strength The way to obtain is as follows: ;
[0013] In the formula, For the image at coordinates The pixel value at ; is the weight of the edge detection kernel, which is defined by the selected edge detection kernel (such as Sobel kernel); and are the horizontal and vertical offsets of the ith position in the edge detection kernel relative to the center position; i is an index variable, indicating the ith position in the edge detection kernel currently being processed; and n is the total number of elements in the edge detection kernel.
[0014] Optionally, the specific steps of S1 further include:
[0015] S13, using the Otsu algorithm, the optimal threshold is set according to the maximum inter-class variance, and the threshold is combined with the edge strength of each pixel. Compare the size of the corresponding pixel points. When the set threshold is exceeded, the corresponding pixel point is used as the edge point of the table, and the straight lines in the image are detected by Hough transform, the horizontal and vertical lines of the table in the image are found, and the boundary lines in the table are determined according to the intersections between the horizontal and vertical lines of the table to generate table data, wherein the table data includes the horizontal line positions, vertical line positions and intersection positions of the table in the image in each group of image files.
[0016] Optionally, the specific steps of S2 include:
[0017] S21, based on the table data, determining the number of columns, the number of rows, the column width and the row width of the table in each group of image files, and combining the text layer information, constructing the feature vector H of each table area, and determining the adjacent relationship between the tables according to the order of the multiple groups of image files converted from the PDF financial report;
[0018] S22, identifying the similarity between adjacent tables based on the feature vector H of each table area to obtain the similarity , specifically:
[0019] In the formula, is the eigenvector of the jth table region, is the eigenvector of the j+1th table region, is the feature vector and the eigenvector The dot product of is the feature vector The norm of .
[0020] Optionally, the specific steps of S2 further include:
[0021] S23, pre-set the similarity threshold, and Compare to determine whether there is cross-page behavior between corresponding adjacent tables. The specific comparison content is as follows:
[0022] If the similarity If the similarity threshold is exceeded, it is determined that there is cross-page behavior between the corresponding adjacent tables, indicating that the current adjacent tables belong to the same group of tables, and the current table is marked as a subtable;
[0023] If the similarity If the similarity threshold is not exceeded, it is determined that there is no cross-page behavior between the corresponding adjacent tables, indicating that the current adjacent tables do not belong to the same group of tables.
[0024] Optionally, the specific steps of S2 further include:
[0025] S24. Based on the comparison result in S23 and in combination with the similarity transitivity hypothesis, if Table A and Table B belong to the same group of tables, and Table B and Table C belong to the same group of tables, then Table A and Table C also belong to the same group of tables. By summarizing, multiple groups of complete tables can be identified in multiple groups of image files, and the complete tables include at least one group of sub-tables.
[0026] Optionally, the S3 specific steps include:
[0027] S31, splicing the sub-tables in the complete table to determine the coordinates of each intersection in the complete table, and dividing the complete table into a plurality of groups of cells according to the coordinates of each intersection in the complete table.
[0028] Optionally, the S3 specific steps also include:
[0029] S32, traversing all cells in the complete table, if there are merged cells in the complete table, and the form type of the merged cells is type 1, identifying the merged cells by the following steps:
[0030] S321, starting from the starting point in the upper left corner of the complete table, expanding to the right until encountering the boundary line, and recording the number of horizontally merged columns Hs, specifically: ;
[0031] In the formula, For cells In line and the boundary line at column p+r; For cells The number of columns to merge horizontally; To obtain the maximum r value, There are no boundary lines in the location; For the cell The boundary line of the position does not exist, where the subscript Translated into Chinese, it means non-existence; r is the number of columns to check to the right from the current position.
[0032] S322, starting from the starting point in the upper left corner of the complete table, extending downward until encountering a boundary line, recording the number of vertically merged rows Zs, specifically: ;
[0033] In the formula, For cells In line and columns The boundary line of the location, For cells The number of rows to merge vertically; To obtain the maximum value of e, There are no boundary lines in the location; For the cell The boundary line of the position does not exist, where the subscript Translated into Chinese, it means non-existence; e is the number of rows to check from the current position downwards;
[0034] S323: Based on the contents of S321 and S322, determine that the form type of the merged cells in the complete table is a range F of type 1, the range .
[0035] Optionally, the S3 specific steps also include:
[0036] S33, traversing all cells in the complete table, if there are merged cells in the complete table, and the form type of the merged cells is type 2, identifying the merged cells in the following manner:
[0037] In the formula, is the result of the judgment, indicating that the cell Is it a blank cell in the merged cell? Represents a cell is part of a merged cell, Represents a cell It is an independent, truly blank cell that does not belong to any merged cell area; Represents a cell It does not span other columns and is an independent cell; Represents a cell spans multiple columns; Represents a cell spans multiple rows; Represents a cell No other rows are crossed; Represents a cell The merge status of Represents a cell It is not the first cell of the merged cell. Represents a cell is the first cell of the merged cell;
[0038] , , indicating a cell Blank cells that are part of a merged cell;
[0039] , , indicating a cell Is a true blank cell that does not belong to any merged area.
[0040] Optionally, the S3 specific steps also include:
[0041] S34. Based on the type of blank cells in the complete table determined in S33, gradually determine that the form type of merged cells in the complete table is a range of type two, and re-perfect the complete table related data based on the merged cell range of type one and the merged cell range of type two, wherein the complete table related data includes the positions of horizontal lines, vertical lines, points, number of columns, number of rows, column width, row width, text layer information and merged cell range in the complete table; based on the complete table related data, convert the complete table involved in the PDF financial report into structured table data, wherein the structured table data includes but is not limited to Excel, JSON and CSV, and in the conversion process, perform a restoration task on the table, specifically: ;
[0042] In the formula, S is the reconstructed table, For cells content, For cells The row index in the complete table ranges from 1 to Q; Q is the number of rows in the complete table. For cells The column index in the complete table ranges from 1 to P, where P is the number of columns in the complete table.
[0043] like , which means that the cell Participated in the merge, cell Contents will be added to the reconstruction table;
[0044] like , which means that the cell Not involved in the merge, cell Contents No re-editing will be done in the rebuilt form.
[0045] The present invention provides a method for identifying complex tables in a document, which has the following beneficial effects:
[0046] (1) By using the similarity transitivity hypothesis, this method can efficiently determine whether there is cross-page behavior in different table areas by calculating the similarity between adjacent tables (by calculating the dot product and norm of the feature vectors). By setting a similarity threshold, the method can automatically identify whether adjacent tables belong to the same group of tables, thereby determining whether there is a cross-page table. This process not only reduces the need for manual intervention, but also significantly improves the accuracy of cross-page table recognition, ensuring the continuity and integrity of table data. By combining similarity analysis and transitivity assumptions, the present invention can automatically identify complete cross-page tables in multiple groups of image files and successfully splice multiple sub-tables into a complete table. In this process, the method can correctly identify the continuity of tables between different pages by summarizing the feature information of adjacent tables. Through this automated cross-page table reconstruction, the problem of cross-page table information loss or misalignment in traditional methods is solved, thereby improving the reliability of table data and analysis effect. In summary, the method of the present invention provides an efficient and accurate solution for the processing of table data, especially in the identification of cross-page tables and the reconstruction of merged tables. This makes the table recognition and data extraction process in complex documents such as PDF financial reports more intelligent and automated, effectively improving the quality and efficiency of information extraction.
[0047] (2) By traversing all cells in the complete table and judging whether each cell is part of a merged cell, the present invention can accurately distinguish between blank merged cells and real blank cells. This process makes the judgment of merged cells more detailed and accurate, avoiding data parsing errors caused by misjudgment of merged cells in traditional methods. In particular, in type II merged cells, it is possible to accurately judge whether each blank cell is part of a merged cell based on the merged state and positional relationship of the cells, thereby ensuring the integrity and accuracy of the table data. In the process of identifying merged cells, the system can not only judge whether the merged cell belongs to the merged area, but also automatically identify the range of horizontal and vertical merged cells based on the number of rows and columns spanned by the cells. Through refined identification, the system can accurately distinguish which cells are independent and which are parts of merged cells. For horizontal and vertical merged cells, the system can automatically identify their spans, avoiding damage to the table structure caused by incorrect calculation of the merged area range. The method of the present invention can ensure that the complete table is effectively restored after merged cell identification by judging the merged state and type of each cell. By accurately identifying each cell in the table, especially the type 2 merged cells, the table can be accurately restored during the merged cell processing. In this way, the data in the table can be accurately restored according to the original layout, and the position of the blank cells can also be correctly calibrated.
[0048] (3) The present invention can accurately identify different types of merged cells in a table by determining the type of blank cells and identifying the range of type one and type two merged cells. In particular, it reduces misjudgment in traditional methods when processing merged cells in complex financial statements. This process ensures the structural integrity and content accuracy of the table by gradually improving the range of merged cells, and avoids data loss or misalignment caused by incorrect merged cell ranges. Structured table data conversion: By converting complete table data into structured table data (such as Excel, JSON or CSV), the present invention achieves efficient extraction and reconstruction of document content. This process not only improves the accuracy of data extraction, but also provides a clear and standardized table format for subsequent data processing and analysis, facilitating further automated analysis and processing. For example, the table data in a PDF financial report is automatically converted to CSV format, so that the data can be directly used for financial analysis, data visualization and other applications. The present invention ensures accurate reconstruction of the content of the merged cells by weighted summing the content of each cell in the table. By combining the merge status of the merged cells (whether they participate in the merge), the content of the merged cells is reasonably adjusted and filled when the table is reconstructed, ensuring the restoration of the table structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the present application or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application.
[0050] Figure 1 The figure is a flow chart of a method for identifying complex tables in a document according to the present invention. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0052] Example 1
[0053] See also Figure 1 The present invention provides a method for identifying a complex table in a document, comprising the following steps:
[0054] S1. Based on the PDF financial report, multiple groups of image files are obtained, and the table area in each group of image files is identified to obtain the table data; PDF Financial Report refers to a financial report published in the form of a PDF (Portable Document Format) file.
[0055] S1 specific steps include:
[0056] S11. According to the company's finance department, the PDF financial report is obtained in advance, and the PDF financial report document is converted into multiple sets of image files, and high-resolution images and text layer information are extracted from the multiple sets of image files, wherein the text layer information includes but is not limited to text content, text position, font and font size;
[0057] S12. Use the edge detection kernel (Sobel kernel) to perform convolution operation on the high-resolution image to obtain the edge strength of each pixel , the edge strength The way to obtain is as follows: ;
[0058] In the formula, For the image at coordinates The pixel value at is obtained by reading the pixel matrix of the image; is the weight of the edge detection kernel, which is defined by the selected edge detection kernel (such as Sobel kernel); and are the horizontal and vertical offsets of the ith position in the edge detection kernel relative to the center position, respectively. They are determined by the size and shape of the kernel. For example, for a 3x3 Sobel kernel, and The value range is ; i is the index variable, indicating the i-th position in the edge detection kernel currently being processed. For example, for a 3x3 Sobel kernel, the value range of i is 1-9; n is the total number of elements in the edge detection kernel. For a k×k edge detection kernel, , for example, for a 3x3 Sobel kernel, n=9, and for a 5x5 Sobel kernel, n=25.
[0059] In the table region detection module, the edge detection kernel is a mathematical tool used to detect edges in an image, usually a matrix (also called a convolution kernel or filter).
[0060] The specific steps of S1 also include:
[0061] S13, using the Otsu algorithm, the optimal threshold is set according to the maximum inter-class variance, and the threshold is combined with the edge strength of each pixel. Compare the size of the corresponding pixel points. When the set threshold is exceeded, the corresponding pixel point is used as the edge point of the table, and the straight lines in the image are detected by Hough transform, the horizontal and vertical lines of the table in the image are found, and the boundary lines in the table are determined according to the intersections between the horizontal and vertical lines of the table to generate table data, wherein the table data includes the horizontal line positions, vertical line positions and intersection positions of the table in the image in each group of image files.
[0062] In this embodiment, the table data in the PDF financial report is efficiently extracted: the method can accurately extract the table area in the document by converting the PDF financial report into a high-resolution image file and combining the text layer information. The use of the text layer information can not only obtain the text content, but also further accurately determine the position, font and font size of the text, thereby effectively improving the accuracy of the table data extraction. Edge detection and table structure recognition: by using the Sobel edge detection kernel, the image can be convolved to identify the edge strength in the image. This process can effectively capture the horizontal and vertical lines of the table, and provide clear edge information for the subsequent table structure recognition. Edge detection technology can find the boundary of the table in the image, providing reliable support for the subsequent table area division and data restoration. Accurate table boundary detection: by combining the Otsu algorithm for threshold setting, the optimal threshold can be automatically determined according to the principle of maximizing the inter-class variance, which helps to distinguish the table edge in the image from the background noise. Hough transform is further used to detect the straight lines in the image, so as to accurately find the horizontal and vertical lines of the table. This technology ensures that the boundary detection of the table can be completed efficiently and accurately, especially in the case of low image quality or noise, and can effectively filter out irrelevant interference information.
[0063] Accurate recognition of cross-page tables: After identifying the horizontal and vertical lines of the table and their intersection positions, the structural data of the table can be further obtained, including the distribution of rows and columns, the content distribution of cells, etc. This method can accurately connect and identify the data of cross-page tables in cross-page PDF financial reports, ensuring that the table data between different pages can be completely and correctly merged. Enhance the accuracy of table data reconstruction and subsequent processing: By accurately identifying and extracting the boundaries, rows and columns and intersection positions of the table, the subsequent table data reconstruction process is further improved. In the processing of merged cells and complex table structures, the layout and content of the original table can be accurately restored, thereby providing efficient support for financial analysis, data processing and automated report generation. This method successfully realizes the table area recognition and data extraction in PDF financial reports through precise edge detection, threshold optimization and straight line detection technology. It not only improves the accuracy of table recognition, but also effectively solves the problems of cross-page tables and merged cells, ensuring the integrity and accuracy of subsequent data processing. These advantages make this method have significant advantages in the application of automated financial report analysis and large-scale data processing.
[0064] Example 2
[0065] Please refer to Figure 1 ,Specifically: S2, based on the table data, analyze the similarity between the ,areas of each table, and perform cross-page recognition in combination with the ,similarity transitivity hypothesis to complete the complete table recognition ,task;
[0066] The specific steps of S2 include:
[0067] S21, based on the table data, determining the number of columns, the number of rows, the column width and the row width of the table in each group of image files, and combining the text layer information to construct a feature vector H of each table area, and determining the adjacent relationship between the tables according to the order of the multiple groups of image files converted from the PDF financial report;
[0068] S22, identifying the similarity between adjacent tables based on the feature vector H of each table area to obtain the similarity , specifically: ;
[0069] In the formula, is the eigenvector of the jth table region, is the eigenvector of the j+1th table region, is the feature vector and the eigenvector The dot product of is the feature vector The norm of (i.e. the length of the vector).
[0070] The specific steps of S2 also include:
[0071] S23, pre-set the similarity threshold, and Compare to determine whether there is cross-page behavior between corresponding adjacent tables. The specific comparison content is as follows:
[0072] If the similarity If the similarity threshold is exceeded, it is determined that there is cross-page behavior between the corresponding adjacent tables, indicating that the current adjacent tables belong to the same group of tables, and the current table is marked as a subtable;
[0073] If the similarity If the similarity threshold is not exceeded, it is determined that there is no cross-page behavior between the corresponding adjacent tables, indicating that the current adjacent tables do not belong to the same group of tables.
[0074] The specific steps of S2 also include:
[0075] S24. Based on the comparison result in S23 and in combination with the similarity transitivity hypothesis, if Table A and Table B belong to the same group of tables, and Table B and Table C belong to the same group of tables, then Table A and Table C also belong to the same group of tables. By summarizing, multiple groups of complete tables can be identified in multiple groups of image files, and the complete tables include at least one group of sub-tables.
[0076] In this embodiment, by extracting the number of columns, rows, column width, row width and other information based on the table data in the PDF financial report, combined with the text layer information, the system can accurately construct the feature vector of each table area. This method ensures that the key features of each table are effectively extracted, and can provide high-quality data support for subsequent similarity analysis and cross-page recognition. Table similarity analysis and cross-page recognition: Using the dot product and norm of the feature vector, the system can efficiently calculate the similarity between adjacent tables, and determine whether there is cross-page behavior by setting a similarity threshold. The degree of similarity directly determines whether adjacent tables belong to the same group, thereby accurately identifying cross-page tables. This method not only improves the accuracy of cross-page table recognition, but also can handle complex multi-page table layouts, so that the information in the table can remain consistent after crossing pages. Application of the similarity transitivity hypothesis: Through the similarity transitivity hypothesis, the system can recursively identify the relationship between cross-page tables. For example, if Table A and Table B belong to the same group of tables, and Table B and Table C also belong to the same group, then Table A and Table C will also be identified as belonging to the same group. This reasoning mechanism not only improves the accuracy of table group recognition, but also enhances the consistency of table data in the entire document. Through this transitivity analysis, the system can successfully identify and reconstruct complete tables in multi-page PDF financial reports, avoiding omissions and misidentifications that may occur when a single table is processed separately.
[0077] Improve the integrity and consistency of cross-page table data: This method completes the accurate recognition of cross-page tables by combining the table features of multiple image files, ensuring that the table data on multiple pages can be successfully spliced and restored. No matter how many parts the table is split into, the system can ensure that these parts can correctly belong to the same group of tables through the similarity of feature vectors and the similarity transfer mechanism, and merge the related sub-tables into a complete table. This ability to recognize cross-page tables effectively avoids data loss and ensures the integrity and consistency of table data. Automation and efficiency: Through the above-mentioned automated steps, including feature vector extraction, similarity calculation, cross-page judgment, and the application of transitivity assumptions, the system can efficiently and automatically process complex financial report data. This method reduces manual intervention, avoids the time and cost overhead in traditional manual processing methods, and improves the overall efficiency and accuracy of table recognition. In summary, the present invention significantly improves the recognition effect of complex tables in multi-page documents by combining high-precision table feature analysis, similarity calculation, cross-page table recognition and similarity transfer mechanism, and provides reliable technical support for the extraction and structured processing of table data in PDF financial reports.
[0078] Example 3
[0079] Please refer to Figure 1 Specifically: S3, based on S2, preliminarily determine the relevant data of the complete table, and determine the type of blank cells according to the form type of the merged cells, and identify the range of the merged cells, and then improve the relevant data of the complete table again, and perform the restoration task on the table based on the improved relevant data of the complete table.
[0080] The specific steps of S3 include:
[0081] S31, splicing the sub-tables in the complete table to determine the coordinates of each intersection in the complete table, and dividing the complete table into a plurality of groups of cells according to the coordinates of each intersection in the complete table.
[0082] The specific steps of S3 also include:
[0083] S32, traversing all cells in the complete table, if there are merged cells in the complete table, and the form type of the merged cells is type 1, identifying the merged cells by the following steps:
[0084] Among them, the merged cells of type 1 refer to standard merged cells, and their characteristics are: this merged cell form is usually the most common type. Cells are merged together through horizontal or vertical merging operations to form an area.
[0085] S321, starting from the starting point in the upper left corner of the complete table, expanding to the right until encountering the boundary line, and recording the number of horizontally merged columns Hs, specifically: ;
[0086] In the formula, For cells In line and the boundary line at the column p+r position, that is, it refers to the Whether the cell boundary at row and column p+r exists; For cells The number of columns to merge horizontally; To obtain the maximum r value, There is no boundary line at the position, which means that the cell has merged so many rows horizontally; For the cell The boundary line of the position does not exist, where the subscript Translated into Chinese, it means does not exist; it means that the cell at this position is part of a merged cell. If there is no boundary line at this position, it means that it is merged with the previous cell; r is the number of columns checked to the right from the current position. For horizontally merged cells, the number of merged columns is determined by checking whether there is a boundary line in the column to the right of the current cell. r is a gradually increasing integer until a place without a boundary line on the right is found, which means that the horizontal merge has ended.
[0087] S322, starting from the starting point in the upper left corner of the complete table, extending downward until encountering a boundary line, recording the number of vertically merged rows Zs, specifically: ;
[0088] In the formula, For cells In line and columns The boundary line of the location, For cells The number of rows to merge vertically; To obtain the maximum value of e, If there is no boundary line at the position, it means that the cell has merged so many rows vertically; e is the number of rows to be checked downward from the current position. For vertically merged cells, the number of merged rows is determined by checking whether there is a boundary line in the row below the current cell. e is a gradually increasing integer until a place without a boundary line is found below, which means that the vertical merge has ended; For the cell The boundary line of the position does not exist, where the subscript Translated into Chinese, it means non-existence;
[0089] S323: Based on the contents of S321 and S322, determine that the form type of the merged cells in the complete table is a range F of type 1, the range ; When calculating the range of merged cells, 1 is subtracted to correctly represent the index range of rows and columns. Assume =2, =3, which means the merged cell starts from row 2 and spans 3 rows. We want to get the row range of the merged cell, which is from row 2 to row 4. If 1 is not subtracted, the range we will get is , which means from row 2 to row 5.
[0090] In this embodiment, precise table splicing and cell division: by splicing multiple sub-tables, the various intersections of the complete table can be effectively identified, ensuring the consistency of the coordinates and structure of the table when it crosses pages or crosses regions. This splicing and division method can ensure that all cells of the table are accurately located, avoiding the problem of data loss in cross-page tables or split tables, thereby achieving complete reconstruction of the table. Automatic identification of merged cells and their ranges: the present invention can accurately determine the position and range of merged cells by automatically identifying the number of horizontally merged columns and vertically merged rows. Using a boundary line-based calculation method, the type and range of merged cells can be accurately determined, especially by checking the merged area column by column and row by row, which can effectively distinguish which cells belong to the merged area, avoiding the situation of inaccurate processing or missed recognition of merged cells in traditional methods. Perfect data processing and table restoration: when calculating the merged cell range, the step of subtracting 1 ensures the correct row and column index range, avoiding the problem of misalignment of table data or inaccurate recognition of merged cells caused by index errors. Through the improved complete table data, it can be ensured that the final restored table is not only complete in content and accurate in layout, but also that all merged cells and blank cells can be accurately identified and processed, providing a reliable basis for subsequent data analysis and decision-making. Efficient processing of cross-page tables and complex layouts: This method can handle tables containing multiple pages to ensure the continuity and consistency of cross-page tables. By performing similarity analysis on cross-page tables and combining the feature vectors of adjacent tables for cross-page identification, the sub-tables in the cross-page tables can be automatically spliced, which improves the efficiency and accuracy of table identification and reconstruction in complex documents (especially financial statements). In summary, the method of the present invention significantly improves the automation and accuracy of table processing through accurate merged cell identification, table splicing and restoration, and automatic identification of cross-page tables. It is particularly suitable for extracting tabular data from complex PDF financial reports and other similar documents, greatly improving the efficiency and reliability of data processing.
[0091] Example 4
[0092] Please refer to Figure 1 , specifically: S3 specific steps also include:
[0093] S33, traversing all cells in the complete table, if there are merged cells in the complete table, and the form type of the merged cells is type 2, identifying the merged cells in the following manner:
[0094] Among them, merged cells of form type 2 refer to blank merged cells, and their characteristics are: this type of merged cells usually involve blank cells, which may have no content themselves, but these blank cells are still part of the merged cells, and their blank state only means that they have no data, but are part of the merged area.
[0095] ;
[0096] In the formula, is the result of the judgment, indicating that the cell Is it a blank cell in the merged cell? Represents a cell is part of a merged cell (i.e. it is blank and part of a merged area), Represents a cell It is an independent, truly blank cell that does not belong to any merged cell area; Represents a cell It does not span other columns and is an independent cell; Represents a cell It spans multiple columns, that is, it is part of a horizontally merged cell; Represents a cell It spans multiple rows, that is, it is part of a vertically merged cell; Represents a cell It does not span other rows and is an independent cell; Represents a cell The merge status of Represents a cell It is not the first cell of the merged cell. Represents a cell is the first cell of the merged cell; , , indicating a cell Blank cells that are part of a merged cell; , , indicating a cell Is a true blank cell that does not belong to any merged area.
[0097] In this embodiment, the difference between merged cells and blank cells is accurately identified: by specially processing the type 2 of merged cells (i.e., blank cells in the merged area), the present invention can accurately distinguish between blank cells in the merged area and real blank cells. In traditional table recognition methods, blank cells are often misjudged as part of merged cells or mistakenly identified as independent cells, while the present invention can classify each blank cell according to the merged state of the cell and the number of rows and columns spanned, effectively avoiding these misjudgments. Dynamic determination of the range of merged cells: the present invention can dynamically determine the precise range of merged cells by combining the merged state of the cell (whether it spans multiple rows or columns) and whether it is the first cell of the merged cell. By determining the blank cells in the merged cells, the actual range of the merged cells can be accurately restored, and data dislocation or omission caused by misjudgment can be prevented. Enhanced table structure recognition capability: by accurately classifying merged cells and blank cells, the present invention can better identify the structure and data distribution of the table, especially for those tables containing complex merged areas (such as financial statements, etc.). This recognition capability makes the table parsing more stable, can avoid structural errors caused by the loss of part of the information in the merged cells, and improves the reliability of table reconstruction. Adapt to different table layouts: The present invention can adapt to different table layouts and merged cell forms, especially when there are a large number of horizontally and vertically merged cells in the document. By accurately identifying the blank part of the merged cells and combining the row and column span information, it can efficiently perform cross-page recognition and merged cell restoration tasks in different types of tables, ensuring that the final identified table structure is complete and in line with reality. In traditional methods, manual inspection and correction of merged cells in tables are often time-consuming and error-prone, while the present invention can automatically complete this process efficiently and accurately, improving the efficiency and accuracy of table data processing. Through accurate determination of the types of merged cells and blank cells, the method of the present invention improves the accuracy of table reconstruction and the integrity of cross-page table recognition, especially when processing documents containing complex merged cells. It has significant advantages.
[0098] Example 5
[0099] Please refer to Figure 1 , specifically: S3 specific steps also include:
[0100] S34. Based on the type of blank cells in the complete table determined in S33, gradually determine that the form type of merged cells in the complete table is a range of type two, and re-perfect the complete table related data based on the merged cell range of type one and the merged cell range of type two, wherein the complete table related data includes the positions of horizontal lines, vertical lines, points, number of columns, number of rows, column width, row width, text layer information and merged cell range in the complete table; based on the complete table related data, convert the complete table involved in the PDF financial report into structured table data, wherein the structured table data includes but is not limited to Excel, JSON and CSV, and in the conversion process, perform a restoration task on the table, specifically: ;
[0101] In the formula, S is the reconstructed table, For cells content, For cells The row index in the complete table ranges from 1 to Q; Q is the number of rows in the complete table. For cells The column index in the complete table ranges from 1 to P, where P is the number of columns in the complete table.
[0102] like , which means that the cell Participated in the merge, cell Contents will be added to the reconstruction table;
[0103] like , which means that the cell Not involved in the merge, cell Contents No re-editing will be done in the rebuilt form;
[0104] The purpose of this step is to reconstruct the content and structure of the merged cells by taking a weighted sum of the merged status of the contents of each cell in the table.
[0105] In this embodiment, by determining the type of blank cells and identifying the range of type 2 merged cells, the present invention can gradually and accurately determine the exact range of each merged cell. Whether it is a horizontal merge or a vertical merge, the merge state of each cell can be clearly identified. This method ensures that the range calculation of the merged cells is accurate and will not affect the subsequent data reconstruction due to misjudgment. Automatic restoration and reconstruction of table data: Based on the accurate identification of merged cells, the present invention can automatically convert the table in the PDF financial report into a structured data format (such as Excel, JSON or CSV). The restoration task in this process can ensure that the data of the merged cells is correctly delivered to the target location, and the table content will not be lost or misplaced due to the confusion of the merged area, thereby improving the accuracy of data extraction. Through this method, the generation process of structured data is highly automated, greatly reducing the occurrence of manual intervention and errors. Efficient table content recovery and merge state judgment: By weighted summing the merge state of each cell content in the table, the method of the present invention can accurately restore the merged cell content in the original table to the reconstructed structured table. The content of the merged cell involved will be correctly added to the reconstructed table, and the cells that are not involved in the merge will not be re-edited. This step ensures the efficiency and accuracy of the table recovery process and ensures the integrity and reliability of the financial report data. Accurate data conversion and format compatibility: The reconstruction process of the present invention supports the conversion of financial report tables into a variety of structured data formats (such as Excel, JSON, CSV), which facilitates subsequent data analysis, storage and sharing. This multi-format support enables table data to be easily integrated into various analysis systems, providing a convenient operation experience for enterprises or financial personnel.
[0106] Enhanced processing capabilities for tables that span multiple pages: The present invention can efficiently process tables that span multiple pages, ensuring that the contents of tables on multiple pages can be accurately spliced and restored. This advantage is particularly suitable for processing multi-page documents such as PDF financial reports, avoiding the problem of inconsistent or lost data due to tables being scattered across multiple pages. The present invention provides an efficient and automated solution for tabular data in PDF financial reports by accurately identifying merged cells and accurately reconstructing table structures. Through weighted summation of merged cells, automated data conversion, and reconstruction of structured tables, the accuracy and efficiency of extracting and recovering tabular data are significantly improved, while reducing the risk of manual intervention and errors, ensuring the integrity and efficiency of data processing.
[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for identifying complex tables in a document, characterized in that: The following steps are involved: S1. Based on the PDF financial report, multiple groups of image files are obtained, and the table area in each group of image files is identified to obtain the table data; S2. Based on the table data, analyze the similarity between the table areas, combine the similarity transitivity hypothesis, perform cross-page recognition, and complete the complete table recognition task; S3. Based on S2, preliminarily determine the relevant data of the complete table, and determine the type of blank cells based on the form type of the merged cells, and identify the range of the merged cells, and further improve the relevant data of the complete table. Based on the improved relevant data of the complete table, perform the restoration task on the table.
2. The method for identifying complex tables in a document according to claim 1, characterized in that: S1 specific steps include: S11. According to the company's finance department, the PDF financial report is obtained in advance, and the PDF financial report document is converted into multiple sets of image files, and high-resolution images and text layer information are extracted from the multiple sets of image files, wherein the text layer information includes but is not limited to text content, text position, font and font size; S12. Use the edge detection kernel to perform convolution operation on the high-resolution image to obtain the edge strength of each pixel , the edge strength The way to obtain is as follows: ; In the formula, For the image at coordinates The pixel value at ; is the weight of the edge detection kernel, which is defined by the selected edge detection kernel; and are the horizontal and vertical offsets of the ith position in the edge detection kernel relative to the center position; i is an index variable, indicating the ith position in the edge detection kernel currently being processed; and n is the total number of elements in the edge detection kernel.
3. The method for identifying complex tables in a document according to claim 2, characterized in that: The specific steps of S1 also include: S13, using the Otsu algorithm, the optimal threshold is set according to the maximum inter-class variance, and the threshold is combined with the edge strength of each pixel. Compare the size of the corresponding pixel points. When the set threshold is exceeded, the corresponding pixel point is used as the edge point of the table, and the straight lines in the image are detected by Hough transform, the horizontal and vertical lines of the table in the image are found, and the boundary lines in the table are determined according to the intersections between the horizontal and vertical lines of the table to generate table data, wherein the table data includes the horizontal line positions, vertical line positions and intersection positions of the table in the image in each group of image files.
4. The method for identifying complex tables in a document according to claim 3, characterized in that: The specific steps of S2 include: S21, based on the table data, determining the number of columns, the number of rows, the column width and the row width of the table in each group of image files, and combining the text layer information, constructing the feature vector H of each table area, and determining the adjacent relationship between the tables according to the order of the multiple groups of image files converted from the PDF financial report; S22, identifying the similarity between adjacent tables based on the feature vector H of each table area to obtain the similarity , specifically: ; In the formula, is the eigenvector of the jth table region, is the eigenvector of the j+1th table region, is the feature vector and the eigenvector The dot product of is the feature vector The norm of .
5. The method for identifying complex tables in a document according to claim 4, characterized in that: The specific steps of S2 also include: S23, pre-set the similarity threshold, and Compare to determine whether there is cross-page behavior between corresponding adjacent tables. The specific comparison content is as follows: If the similarity If the similarity threshold is exceeded, it is determined that there is cross-page behavior between the corresponding adjacent tables, indicating that the current adjacent tables belong to the same group of tables, and the current table is marked as a subtable; If the similarity If the similarity threshold is not exceeded, it is determined that there is no cross-page behavior between the corresponding adjacent tables, indicating that the current adjacent tables do not belong to the same group of tables.
6. The method for identifying complex tables in a document according to claim 5, characterized in that: The specific steps of S2 also include: S24. Based on the comparison result in S23 and in combination with the similarity transitivity hypothesis, if Table A and Table B belong to the same group of tables, and Table B and Table C belong to the same group of tables, then Table A and Table C also belong to the same group of tables. By summarizing, multiple groups of complete tables can be identified in multiple groups of image files, and the complete tables include at least one group of sub-tables.
7. The method for identifying complex tables in a document according to claim 6, characterized in that: The specific steps of S3 include: S31, splicing the sub-tables in the complete table to determine the coordinates of each intersection in the complete table, and dividing the complete table into a plurality of groups of cells according to the coordinates of each intersection in the complete table.
8. The method for identifying complex tables in a document according to claim 7, characterized in that: The specific steps of S3 also include: S32, traversing all cells in the complete table, if there are merged cells in the complete table, and the form type of the merged cells is type 1, identifying the merged cells by the following steps: S321, starting from the starting point in the upper left corner of the complete table, expanding to the right until encountering the boundary line, and recording the number of horizontally merged columns Hs, specifically: ; In the formula, For cells In line and the boundary line at column p+r; For cells The number of columns to merge horizontally; To obtain the maximum r value, There are no boundary lines in the location; For the cell The boundary line of the position does not exist, where the subscript Translated into Chinese, it means non-existence; r is the number of columns to check to the right from the current position; S322, starting from the starting point in the upper left corner of the complete table, extending downward until encountering a boundary line, recording the number of vertically merged rows Zs, specifically: ; In the formula, For cells In line and column The boundary line of the location, For cells The number of rows to merge vertically; To obtain the maximum value of e, There are no boundary lines in the location; For the cell The boundary line of the position does not exist, where the subscript Translated into Chinese, it means non-existence; e is the number of rows to check from the current position downwards; S323: Based on the contents of S321 and S322, determine that the form type of the merged cells in the complete table is a range F of type 1, the range .
9. The method for identifying complex tables in a document according to claim 8, characterized in that: The specific steps of S3 also include: S33, traversing all cells in the complete table, if there are merged cells in the complete table, and the form type of the merged cells is type 2, identifying the merged cells in the following manner: ; In the formula, Is the judgment result, indicating that the cell Is it a blank cell in the merged cell? Represents a cell is part of a merged cell, Represents a cell It is an independent, truly blank cell that does not belong to any merged cell area; Represents a cell It does not span other columns and is an independent cell; Represents a cell spans multiple columns; Represents a cell spans multiple rows; Represents a cell No other rows are crossed; Represents a cell The merge status of Represents a cell It is not the first cell of the merged cell. Represents a cell is the first cell of the merged cell; , , indicating a cell Blank cells that are part of a merged cell; , , indicating a cell Is a true blank cell that does not belong to any merged area.
10. The method for identifying complex tables in a document according to claim 9, characterized in that: The specific steps of S3 also include: S34. Based on the type of blank cells in the complete table determined in S33, gradually determine that the form type of merged cells in the complete table is a range of type two, and re-perfect the complete table related data based on the merged cell range of type one and the merged cell range of type two, wherein the complete table related data includes the positions of horizontal lines, vertical lines, points, number of columns, number of rows, column width, row width, text layer information and merged cell range in the complete table; based on the complete table related data, convert the complete table involved in the PDF financial report into structured table data, wherein the structured table data includes but is not limited to Excel, JSON and CSV, and in the conversion process, perform a restoration task on the table, specifically: ; In the formula, S is the reconstructed table, For cells content, For cells The row index in the complete table ranges from 1 to Q; Q is the number of rows in the complete table. For cells The column index in the complete table ranges from 1 to P, where P is the number of columns in the complete table. like , which means that the cell Participated in the merge, cell Contents will be added to the reconstruction table; like , which means that the cell Not involved in the merge, cell Contents No re-editing will be done in the rebuilt form.
Citation Information
Patent Citations
Table recognition method, device and storage medium
CN110532968A
Announcement document table data identification method and terminal
CN112434496A
Table standardization processing method and device, equipment and storage medium
CN113033170A
Spreadsheet data processing method and device, computer equipment and storage medium
CN113553326A
PDF (Portable Document Format) document-oriented table data extraction method and device
CN117454851A