A method for recognizing tables in a document
By identifying and analyzing the table areas in the PDF financial report, combining the similarity transferability assumptions for cross-page recognition, and determining the scope of merged cells, the shortcomings of cross-page tables and merged cells recognition in the existing technology are solved, and efficient and accurate tabular data extraction and reconstruction are achieved.
Patent Information
- Application Number
- CN202510498402.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-21
AI Technical Summary
Existing document table recognition methods have shortcomings in processing spreadsheets and merging cells, making it difficult to accurately identify blank cells and spread information, resulting in data parsing errors and data loss.
By obtaining image files in PDF financial reports, identifying table areas, analyzing similarity between tables, combining similarity transferability assumptions, performing span recognition, determining the range of merged cells, and performing restore tasks on the table.
Improve the accuracy of cross-page table recognition, ensure the continuity and integrity of the table data, accurately distinguish between blank merged cells and real blank cells, avoid data parsing errors, and enhance the reliability and analysis effect of the table data.
Smart Images

Figure CN120014662B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of table recognition, and particularly to a method for recognizing tables in documents. Background Art
[0002] In the context of the rapid development of information technology and artificial intelligence, document image processing and analysis have become an important research direction. This field covers technologies such as image recognition, optical character recognition, and natural language processing, and is applied to automated data extraction, information understanding, and processing in different scenarios. With the widespread use of various electronic documents, how to efficiently and accurately extract structured information from these documents, especially table data in financial reports, has become the focus of research. Especially in common financial statements such as PDF financial reports, tables not only have complex layouts but also contain special cases such as merged cells and multi-page tables, which pose extremely high challenges to automated recognition.
[0003] Although existing document table recognition methods have made certain progress in dealing with standard table layouts, there are still many deficiencies when faced with tables, especially multi-page tables and merged cells. Most existing methods rely on fixed rule-based pattern matching and are difficult to adapt to the diversity of different table styles and layouts. When identifying blank cells, especially when a blank cell is part of a merged cell, existing methods often have difficulty accurately distinguishing between true blank cells and those that belong to the merged area, which directly leads to errors in data parsing. In addition, for multi-page tables, since table content may be split across multiple pages, current multi-page recognition methods lack the ability to accurately transfer multi-page table information, which may result in data loss or inconsistency, thus affecting the final table data reconstruction and analysis results.
[0004] In the process of table recognition, deficiencies in the type determination of blank cells and the analysis of multi-page tables will lead to the following main problems: 1. Incorrect judgment of blank cells: If a blank cell is misjudged as part of a merged cell, or vice versa, it may cause the merged state and data structure of the table to not be correctly restored, affecting the accuracy of the final data. For example, the content of some merged cells may not be correctly transferred to the corresponding cell positions, resulting in data loss or errors. 2. Incorrect splicing of multi-page tables: In the process of processing multi-page tables, if multi-page information fails to be correctly transferred, it may cause misalignment of table content between pages. Especially when the merged area of a multi-page table fails to be correctly connected, the context relationship of the data will be lost, thus affecting the analysis of the table and subsequent decision-making support. Summary of the Invention
[0005] In view of the above problems existing in the prior art, the present application provides a method for recognizing tables in documents.
[0006] An embodiment of the present disclosure provides a method for recognizing tables in a document, including the following steps:
[0007] S1. Based on the PDF financial report, obtain multiple groups of image files, recognize the table areas in each group of image files, and obtain table data;
[0008] S2. Based on the table data, analyze the similarity between each table area, and combine the similarity transitivity hypothesis to perform cross-page recognition to complete the complete table recognition task;
[0009] S3. On the basis of S2, preliminarily determine the data related to the complete table, determine the types of blank cells according to the form types of merged cells, identify the ranges of merged cells, further improve the data related to the complete table, and perform a restoration task on the table according to the improved data related to the complete table.
[0010] Optionally, the specific steps of S1 include:
[0011] S11. Obtain the PDF financial report in advance according to the company's finance department, convert the PDF financial report document into multiple groups of image files, and extract high-resolution images and text layer information from the multiple groups of image files, where the text layer information includes but is not limited to text content, text position, font, and font size;
[0012] S12. Use edge detection to perform convolution operations on the high-resolution images to obtain the edge intensity of each pixel point , and the acquisition method of the edge intensity is as follows:
[0013] ;
[0014] In the formula, is the pixel value of the image at the coordinate ; is the weight of the edge detection kernel, defined by the selected edge detection kernel (such as the Sobel kernel); and are the offsets of the i-th position in the edge detection kernel relative to the center position in the horizontal direction and the vertical direction, respectively; i is an index variable representing the i-th position in the currently processed edge detection kernel; n is the total number of elements in the edge detection kernel.
[0015] Optionally, the specific steps of S1 further include:
[0016] S13. Use the Otsu algorithm to set the optimal threshold according to the maximum between-class variance, compare the threshold with the edge intensity of each pixel point. If the edge intensity of the corresponding pixel point When it exceeds the set threshold, the corresponding pixel points are regarded as the edge points of the table at this time, and the straight lines existing in the image are detected by the Hough transform to find the horizontal and vertical lines of the table in the image. According to the intersection points between the horizontal and vertical lines of the table, the boundary lines in the table are determined to generate table data. Among them, the table data includes the horizontal line positions, vertical line positions, and intersection positions of the tables existing in the images of each group of image files.
[0017] Optionally, the specific steps of S2 include:
[0018] S21. Based on the table data, determine the number of columns, number of rows, column width, and row width of the tables in each group of image files, and combine the text layer information to construct the feature vector H of each table area. According to the order of the multiple groups of image files converted from the PDF financial report, determine the adjacent relationship between the tables;
[0019] S22. Identify the similarity between adjacent tables according to the feature vector H of each table area to obtain the similarity , specifically:
[0020]
[0021] In the formula, is the feature vector of the jth table area, is the feature vector of the (j + 1)th table area, is the feature vector and the feature vector is the dot product of, is the feature vector is the norm of.
[0022] Optionally, the specific steps of S2 further include:
[0023] S23. Preset a similarity threshold, and compare the similarity threshold with the similarity to determine whether there is a page-crossing behavior between the corresponding adjacent tables. The specific comparison content is as follows:
[0024] If the similarity exceeds the similarity threshold, it is determined that there is a page-crossing behavior between the corresponding adjacent tables, indicating that the current adjacent tables belong to the same group of tables, and the current table is marked as a sub-table;
[0025] If the similarity does not exceed the similarity threshold, it is determined that there is no page-crossing behavior between the corresponding adjacent tables, indicating that the current adjacent tables do not belong to the same group of tables.
[0026] Optionally, the specific steps of S2 further include:
[0027] S24. According to the comparison result in S23 and combining with the similarity transitivity hypothesis, if table A and table B belong to the same group of tables, and table B and table C belong to the same group of tables, then table A and table C also belong to the same group of tables. Through summarization, multiple complete tables are identified in multiple groups of image files, and the complete table includes at least one group of sub-tables.
[0028] Optionally, the specific steps of S3 include:
[0029] S31. Stitch the sub-tables in the complete table to determine the coordinates at each intersection in the complete table. According to the coordinates at each intersection in the complete table, divide the complete table into several groups of cells.
[0030] Optionally, the specific steps of S3 also include:
[0031] S32. Traverse all the cells in the complete table. If there are merged cells in the complete table and the form type of the merged cell is type one, then identify the merged cell through the following steps:
[0032] S321. Starting from the starting point in the upper left corner of the complete table, expand to the right until a boundary line is encountered, and record the number of horizontally merged columns Hs. Specifically: ;
[0033] In the formula, is the boundary line at the position of the cell in row and column p + r; is the number of horizontally merged columns of the cell ; is to take the maximum r value such that there is no boundary line at the position of ; means that there is no boundary line at the position of the cell . Among them, the subscript is translated into Chinese and means non-existent; r is the number of columns checked to the right starting from the current position.
[0034] S322. Starting from the starting point in the upper left corner of the complete table, expand downwards until a boundary line is encountered, and record the number of vertically merged rows Zs. Specifically:
[0035] ;
[0036] In the formula, is the boundary line at the position of the cell in row and column , is the number of vertically merged rows of the cell ; To obtain the maximum value of e such that there is no boundary line at the position; For the boundary line at the position of the cell not to exist, where the subscript translated into Chinese means non-existence; e is the number of rows to be checked downward starting from the current position;
[0037] S323. Based on the content of S321 and S322, determine the range F of the form type of the merged cells in the complete table as type one, and the range .
[0038] Optionally, the specific steps of S3 further include:
[0039] S33. Traverse all cells in the complete table. If there are merged cells in the complete table and the form type of the merged cells is type two, then identify the merged cells by the following method:
[0040]
[0041] In the formula, is the determination result, indicating whether the cell is a blank cell in the merged cell, indicates that the cell is part of the merged cell, indicates that the cell is an independent, truly blank cell and does not belong to any merged cell area; indicates that the cell does not span other columns and is an independent cell; indicates that the cell spans multiple columns; indicates that the cell spans multiple rows; indicates that the cell does not span other rows; indicates the merged state of the cell , indicates that the cell is not the first cell of the merged cell, indicates that the cell is the first cell of the merged cell;
[0042] , , indicates that the cell is a blank cell belonging to the merged cell;
[0043] , , indicating a cell is a truly blank cell and does not belong to any merged area.
[0044] Optionally, the specific steps of S3 further include:
[0045] S34. Based on the determination of the types of blank cells in the complete table in S33, gradually determine the range of the form type of the merged cells in the complete table that is type two, and based on the range of the merged cells of type one and the range of the merged cells of type two, re-improve the relevant data of the complete table. The relevant data of the complete table includes the positions of each horizontal line, vertical line, dot position, number of columns, number of rows, column width, row width, text layer information, and the range of merged cells in the complete table; based on the relevant data of the complete table, convert the complete table involved in the PDF financial report into structured table data, where the structured table data includes but is not limited to Excel, JSON, and CSV, and during the conversion process, perform a restoration task on the table, specifically:
[0046] ;
[0047] In the formula, S is the reconstructed table, is the cell content, is the cell row index in the complete table, and the value range is from 1 to Q; Q is the number of rows of the complete table, is the cell column index in the complete table, and the value range is from 1 to P; P is the number of columns of the complete table;
[0048] If , it means that this cell participates in the merger, and the content of the cell will be added to the reconstructed table; will be added to the reconstructed table;
[0049] If , it means that this cell does not participate in the merger, and the content of the cell will not be re-edited in the reconstructed table. will not be re-edited in the reconstructed table.
[0050] The present invention provides a method for identifying a table in a document, having the following beneficial effects:
[0051] (1) By leveraging the assumption of similarity transitivity, this method can efficiently determine whether there is a behavior of spanning pages in different table regions by calculating the similarity between adjacent tables (through the dot product and norm of feature vectors). By setting a similarity threshold, this method can automatically identify whether adjacent tables belong to the same group of tables, thereby determining whether there are tables spanning pages. This process not only reduces the need for manual intervention but also significantly improves the accuracy of identifying tables spanning pages, ensuring the continuity and integrity of table data. Through the combination of similarity analysis and the assumption of transitivity, the present invention can automatically identify complete tables spanning pages in multiple groups of image files and successfully splice multiple sub-tables into a complete table. In this process, by summarizing the feature information of adjacent tables, the method can correctly identify the continuity of tables between different pages. Through this automated reconstruction of tables spanning pages, the problem of lost or misaligned information of tables spanning pages in traditional methods is solved, thereby improving the reliability and analysis effect of table data. In summary, the method of the present invention provides an efficient and accurate solution in the processing of table data, especially showing significant advantages in the identification of tables spanning pages and the reconstruction of merged tables. This makes the process of identifying tables and extracting data in complex documents such as PDF financial reports more intelligent and automated, effectively improving the quality and efficiency of information extraction.
[0052] (2) By traversing all cells within a complete table and determining whether each cell is part of a merged cell, the present invention can accurately distinguish between blank merged cells and truly blank cells. This process makes the determination of merged cells more meticulous and accurate, avoiding data parsing errors caused by misjudging merged cells in traditional methods. Especially in type-two merged cells, it is possible to accurately determine whether each blank cell is part of a merged cell based on the merging status and positional relationship of the cells, thereby ensuring the integrity and accuracy of table data. During the identification of merged cells, the system can not only determine whether a merged cell belongs to a merged area but also automatically identify the ranges of horizontally and vertically merged cells according to the number of rows and columns spanned by the cell. Through refined identification, the system can accurately distinguish which cells are independent and which are components of merged cells. For horizontally and vertically merged cells, the system can automatically identify their spans, avoiding damage to the table structure caused by incorrect calculation of the merged area range. The method of the present invention can ensure the effective restoration of a complete table after the identification of merged cells by judging the merging status and type of each cell. Through the precise identification of each cell within the table, especially the identification of type-two merged cells, the table can be accurately restored during the process of handling merged cells. In this way, the data in the table can be accurately restored according to the original layout, and the positions of blank cells can also be correctly marked.
[0053] (3) Through the type determination of blank cells and the range recognition of merged cells of type one and type two, the present invention can accurately identify different types of merged cells in a table. Especially when dealing with merged cells in complex financial statements, it reduces misjudgments in traditional methods. This process ensures the structural integrity of the table and the accuracy of the content by gradually improving the range of merged cells, avoiding data loss or misalignment caused by incorrect merged cell ranges. Structured table data conversion: By converting complete table data into structured table data (such as Excel, JSON, or CSV), the present invention realizes the efficient extraction and reconstruction of document content. This process not only improves the accuracy of data extraction but also provides a clear and standardized table format for subsequent data processing and analysis, facilitating further automated analysis and processing. For example, automatically converting the table data in a PDF financial report into CSV format enables the data to be directly used in applications such as financial analysis and data visualization. The present invention ensures the accurate reconstruction of the content of merged cells by performing weighted summation on the content of each cell in the table. By combining the merging status of merged cells (whether they participate in the merge), the content of merged cells is reasonably adjusted and filled during table reconstruction to ensure the restoration of the table structure. Description of the Drawings
[0054] To more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application.
[0055] Figure 1 It is a schematic flowchart of a method for identifying a table in a document according to the present invention. Detailed Embodiments
[0056] To make the objectives, technical solutions, and advantages of the present application clearer, the following will clearly and completely describe the technical solutions in the present application with reference to the drawings in the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0057] Embodiment 1
[0058] Please refer to Figure 1 , the present invention provides a method for identifying a table in a document, including the following steps:
[0059] S1. Based on the PDF financial report, obtain multiple groups of image files, identify the table areas in each group of image files, and obtain the table data. The PDF financial report (PDF Financial Report) refers to the financial report published in the form of a PDF (Portable Document Format) file.
[0060] The specific steps of S1 include:
[0061] S11. Obtain the PDF financial report in advance according to the company's finance department, convert the PDF financial report document into multiple groups of image files, and extract high-resolution images and text layer information from the multiple groups of image files. Among them, the text layer information includes but is not limited to text content, text position, font, and font size.
[0062] S12. Perform a convolution operation on the high-resolution image using an edge detection kernel (Sobel kernel) to obtain the edge intensity of each pixel point. , The acquisition method of the edge intensity is as follows:
[0063] ;
[0064] In the formula, is the pixel value of the image at the coordinate , which is obtained by reading the pixel matrix of the image; is the weight of the edge detection kernel, which is defined by the selected edge detection kernel (such as the Sobel kernel); and are the offsets of the i-th position in the edge detection kernel relative to the center position in the horizontal direction and the vertical direction respectively, which are specifically determined by the size and shape of the kernel. For example, for a 3x3 Sobel kernel, and range from ; i is an index variable representing the i-th position in the currently processed edge detection kernel. For example, for a 3x3 Sobel kernel, the value range of i is 1 - 9; n is the total number of elements in the edge detection kernel. For a k×k edge detection kernel, , for example, for a 3x3 Sobel kernel, n = 9, and for a 5x5 Sobel kernel, n = 25.
[0065] In the table area detection module, the edge detection kernel is a mathematical tool for detecting edges in an image, usually a matrix (also known as a convolution kernel or filter).
[0066] The specific steps of S1 also include:
[0067] S13. Using the Otsu algorithm, set the optimal threshold according to the maximum inter-class variance, and compare the threshold with the edge intensity of each pixel point for size comparison. If the edge intensity of the corresponding pixel point exceeds the set threshold, then regard the corresponding pixel point as a table edge point at this time, and detect the straight lines existing in the image through the Hough transform to find the horizontal and vertical lines of the table in the image. Determine the boundary lines in the table according to the intersections between the horizontal and vertical lines of the table to generate table data, where the table data includes the horizontal line positions, vertical line positions, and intersection positions of the tables existing in the images in each group of image files.
[0068] In this embodiment, the table data in the PDF financial report is efficiently extracted: This method can accurately extract the table area in the document by converting the PDF financial report into a high-resolution image file and combining the text layer information. The utilization of the text layer information can not only obtain the text content, but also further accurately determine the position, font, and font size of the text, thereby effectively improving the accuracy of table data extraction. Edge detection and table structure recognition: By using the Sobel edge detection kernel, the image can be convolved to identify the edge intensity in the image. This process can effectively capture the horizontal and vertical lines of the table, providing clear edge information for subsequent table structure recognition. The edge detection technology can find the boundaries of the table in the image, providing reliable support for subsequent table area division and data restoration. Precise table boundary detection: By combining the Otsu algorithm for threshold setting, the optimal threshold can be automatically determined according to the principle of maximum inter-class variance, which helps to distinguish the table edges from the background noise in the image. Further use the Hough transform to detect the straight lines in the image to accurately find the horizontal and vertical lines of the table. This technology ensures that the table boundary detection can be completed efficiently and accurately, especially in the case of low image quality or the presence of noise, and can effectively filter out irrelevant interference information.
[0069] Accurate recognition of multi-page tables: After identifying the horizontal and vertical lines of the table and their intersection positions, the structural data of the table can be further obtained, including the distribution of rows and columns, the content distribution of cells, etc. This method can accurately dock and identify the data of multi-page tables in PDF financial reports, ensuring that the table data between different pages can be completely and correctly merged. Enhancing the accuracy of table data reconstruction and subsequent processing: By accurately identifying and extracting the boundaries, rows, columns, and intersection positions of the table, the subsequent table data reconstruction process is further improved. In the processing of merged cells and complex table structures, the original table layout and content can be accurately restored, providing efficient support for financial analysis, data processing, and automated report generation. This method successfully realizes the recognition and data extraction of table areas in PDF financial reports through precise edge detection, threshold optimization, and line detection techniques. It not only improves the accuracy of table recognition but also effectively solves the problems of multi-page tables and merged cells, ensuring the integrity and accuracy of subsequent data processing. These advantages make this method have significant advantages in the application of automated financial report analysis and large-scale data processing.
[0070] Embodiment 2
[0071] Please refer to Figure 1 , specifically: S2. Based on the table data, analyze the similarity between each table area, and combine the similarity transitivity hypothesis to perform cross-page recognition to complete the complete table recognition task;
[0072] The specific steps of S2 include:
[0073] S21. Based on the table data, determine the number of columns, rows, column widths, and row widths of the tables in each group of image files, and combine the text layer information to construct the feature vector H of each table area. According to the order of the multiple groups of image files converted from the PDF financial report, determine the adjacent relationship between each table;
[0074] S22. According to the feature vector H of each table area, identify the similarity between adjacent tables to obtain the similarity , specifically:
[0075] ;
[0076] In the formula, is the feature vector of the j-th table area, is the feature vector of the (j + 1)-th table area, is the dot product of the feature vector and the feature vector , is the norm of the feature vector (i.e., the length of the vector).
[0077] The specific steps of S2 also include:
[0078] S23. Preset a similarity threshold, and by comparing the similarity threshold with the similarity to determine whether there is a page-crossing behavior between the corresponding adjacent tables. The specific comparison content is as follows:
[0079] If the similarity exceeds the similarity threshold, it is determined that there is a page-crossing behavior between the corresponding adjacent tables, indicating that the current adjacent tables belong to the same group of tables, and the current table is marked as a sub-table;
[0080] If the similarity does not exceed the similarity threshold, it is determined that there is no page-crossing behavior between the corresponding adjacent tables, indicating that the current adjacent tables do not belong to the same group of tables.
[0081] The specific steps of S2 also include:
[0082] S24. According to the comparison result in S23, combined with the similarity transitivity hypothesis, if table A and table B belong to the same group of tables, and table B and table C belong to the same group of tables, then table A and table C also belong to the same group of tables. Through summarization, multiple complete tables are identified in multiple groups of image files, and the complete table includes at least one group of sub-tables.
[0083] In this embodiment, based on the number of columns, number of rows, column width, row width, etc. extracted from the table data in the PDF financial report, combined with the text layer information, the system can accurately construct the feature vector of each table area. This method ensures that the key features of each table are effectively extracted, and can provide high-quality data support for subsequent similarity analysis and page-crossing recognition. Table similarity analysis and page-crossing recognition: Using the dot product and norm of the feature vectors, the system can efficiently calculate the similarity between adjacent tables, and determine whether there is a page-crossing behavior by setting a similarity threshold. The level of similarity directly determines whether adjacent tables belong to the same group, so as to accurately identify page-crossing tables. This method not only improves the accuracy of page-crossing table recognition, but also can handle complex multi-page table layouts, so that the information in the table can still maintain consistency after crossing pages. Application of the similarity transitivity hypothesis: Through the similarity transitivity hypothesis, the system can recursively identify the relationship of page-crossing tables. For example, if table A and table B belong to the same group of tables, and table B and table C also belong to the same group, then table A and table C will also be identified as belonging to the same group. This reasoning mechanism not only improves the accuracy of table group recognition, but also enhances the coherence of the table data in the entire document. Through this transitivity analysis, the system can successfully identify and reconstruct complete tables in multi-page PDF financial reports, avoiding omissions and misidentifications that may occur when processing individual tables separately.
[0084] Improving the integrity and consistency of cross-page table data: This method completes the accurate recognition of cross-page tables by combining the table features of multiple image files, ensuring that the table data on multiple pages can be successfully spliced and restored. Regardless of how many parts the table is split into, the system can ensure that these parts correctly belong to the same set of tables through the similarity of feature vectors and the similarity transfer mechanism, and merge the relevant sub-tables into a complete table. This ability to recognize cross-page tables effectively avoids data loss and ensures the integrity and consistency of table data. Automation and efficiency: Through the above automated steps, including feature vector extraction, similarity calculation, cross-page judgment, and the application of transitive assumptions, the system can efficiently and automatically process complex financial report data. This method reduces manual intervention, avoids the time and cost overhead in traditional manual processing methods, and improves the overall efficiency and accuracy of table recognition. In summary, the present invention significantly improves the recognition effect of tables in multi-page documents by combining high-precision table feature analysis, similarity calculation, cross-page table recognition, and similarity transfer mechanism, providing reliable technical support for the extraction and structured processing of table data in PDF financial reports.
[0085] Embodiment 3
[0086] Please refer to Figure 1 , specifically: S3. On the basis of S2, preliminarily determine the data related to the complete table, determine the types of blank cells according to the form types of merged cells, identify the ranges of merged cells, further improve the data related to the complete table, and perform the restoration task on the table according to the improved data related to the complete table.
[0087] The specific steps of S3 include:
[0088] S31. Splice the sub-tables in the complete table to determine the coordinates of each intersection in the complete table, and divide the complete table into several groups of cells according to the coordinates of each intersection in the complete table.
[0089] The specific steps of S3 also include:
[0090] S32. Traverse all cells in the complete table. If there are merged cells in the complete table and the form type of the merged cells is type one, then identify the merged cells through the following steps:
[0091] Among them, the merged cell with the form type of type one refers to the standard merged cell, and the characteristics are: this form of merged cell is usually the most common type. Cells are merged together through horizontal or vertical merging operations to form a region.
[0092] S321. Starting from the starting point at the upper left corner of the complete table, expand to the right until a boundary line is encountered, and record the number of horizontally merged columns Hs, specifically:
[0093] ;
[0094] In the formula, is the boundary line of the cell at the position of row and column p + r. That is to say, it refers to whether there is a boundary of the cell at the row and the (p + r)-th column; is the number of horizontally merged columns of the cell ; is to take the maximum value of r such that there is no boundary line at the position, indicating that the cell has been merged horizontally for so many rows; is that there is no boundary line at the position of the cell . Among them, the subscript is translated into Chinese as "indicating non-existence"; it means that the cell at this position is part of the merged cell. If there is no boundary line at this position, it means that it is merged with the previous cell; r is the number of columns checked to the right from the current position. For horizontally merged cells, the number of merged columns is determined by checking whether there is a boundary line in the columns to the right of the current cell. r is an increasing integer until a place where there is no boundary line on the right is found, which means that the horizontal merge has ended.
[0095] S322. Starting from the starting point at the upper left corner of the complete table, expand downwards until a boundary line is encountered, and record the number of vertically merged rows Zs, specifically:
[0096] ;
[0097] In the formula, is the boundary line of the cell at the position of row and column ; is the number of vertically merged rows of the cell ; is to take the maximum value of e such that there is no boundary line at the position, indicating that the cell has been merged vertically for so many rows; e is the number of rows checked downwards from the current position. For vertically merged cells, the number of merged rows is determined by checking whether there is a boundary line in the rows below the current cell. e is an increasing integer until a place where there is no boundary line below is found, which means that the vertical merge has ended; is for the cell There is no boundary line for the position, where the subscript is translated into Chinese and means non-existent;
[0098] S323. Based on the content of S321 and S322, determine the range F of the form type of the merged cells in the complete table as type one. The range ; When calculating the range of the merged cells, subtracting 1 is to correctly represent the index range of rows and columns. Assume = 2, = 3, indicating that the merged cells start from the second row and span three rows. We want to obtain the row range of the merged cells, that is, from the second row to the fourth row. If we do not subtract 1, the range we will get is , which means from the second row to the fifth row.
[0099] In this embodiment, precise table splicing and cell division: By splicing multiple sub-tables, it is possible to effectively identify each intersection of the complete table, ensuring the consistency of the coordinates and structure of the table when it spans pages or regions. This splicing and division method can ensure that all cells of the table are accurately positioned, avoiding the problem of data loss in cross-page tables or split tables, thus achieving the complete reconstruction of the table. Automatically identifying merged cells and their ranges: The present invention can accurately judge the positions and ranges of merged cells through the automatic identification of the number of horizontally merged columns and vertically merged rows. Using a calculation method based on boundary lines, it accurately determines the types and ranges of merged cells. Especially by checking the merged areas column by column and row by row, it can effectively distinguish which cells belong to the merged areas, avoiding inaccurate processing or missed identification of merged cells in traditional methods. Improved data processing and table restoration: When calculating the range of merged cells, the step of subtracting 1 ensures the correct row and column index ranges, avoiding problems such as misalignment of table data or inaccurate identification of merged cells caused by index errors. Through the improved complete table data, it can ensure that the finally restored table is not only complete in content and accurate in layout, but also all merged cells and blank cells can be accurately identified and processed, providing a reliable basis for subsequent data analysis and decision-making. Efficiently processing cross-page tables and complex layouts: This method can process tables with multiple pages, ensuring the continuity and consistency of cross-page tables. By performing similarity analysis on cross-page tables and combining the feature vectors of adjacent tables for cross-page identification, it can automatically splice the sub-tables in cross-page tables, improving the efficiency and accuracy of table identification and reconstruction in complex documents (especially financial statements). In summary, the method of the present invention significantly improves the automation and accuracy of table processing through precise identification of merged cells, table splicing and restoration, and automatic identification of cross-page tables, especially suitable for table data extraction in complex PDF financial reports and other similar documents, greatly improving the efficiency and reliability of data processing.
[0100] Example 4
[0101] Please refer to Figure 1 , specifically: The specific steps of S3 also include:
[0102] S33. Traverse all cells in the complete table. If there are merged cells in the complete table and the form type of the merged cells is type two, then identify the merged cells in the following way:
[0103] Among them, the merged cells with the form type of type two refer to blank merged cells. Feature: This type of merged cells usually involves blank cells. These cells may have no content themselves, but these blank cells are still part of the merged cells. Their blank state only indicates that they have no data, but they are part of the merged area.
[0104] ;
[0105] In the formula, is the judgment result, indicating whether the cell is a blank cell in the merged cell, indicates that the cell is part of the merged cell (i.e., it is blank and part of the merged area), indicates that the cell is an independent, truly blank cell and does not belong to any merged cell area; indicates that the cell does not span other columns and is an independent cell; indicates that the cell spans multiple columns, i.e., it is part of a horizontally merged cell; indicates that the cell spans multiple rows, i.e., it is part of a vertically merged cell; indicates that the cell does not span other rows and is an independent cell; indicates the merged state of the cell , indicates that the cell is not the first cell of the merged cell, indicates that the cell is the first cell of the merged cell;
[0106] , , indicates that the cell is a blank cell belonging to the merged cell;
[0107] , , indicating that the cell is a truly blank cell and does not belong to any merged area.
[0108] In this embodiment, the difference between merged cells and blank cells is accurately identified: by specifically processing the second type of merged cells (i.e., blank cells in the merged area), the present invention can accurately distinguish between blank cells in the merged area and truly blank cells. In traditional table recognition methods, blank cells are often misjudged as part of a merged cell or misidentified as independent cells, while the present invention can classify each blank cell according to the merging status of the cell and the number of rows and columns it spans, effectively avoiding these misjudgments. Dynamically determining the range of merged cells: The present invention can dynamically determine the exact range of merged cells by combining the merging status of the cell (whether it spans multiple rows or columns) and whether it is the first cell of a merged cell. By determining the blank cells in the merged cell, the actual range of the merged cell can be accurately restored, and data misalignment or omission caused by misjudgment can be prevented. Enhanced table structure recognition ability: By accurately classifying merged cells and blank cells, the present invention can better recognize the structure and data distribution of the table. Especially for tables with complex merged areas (such as financial statements, etc.), this recognition ability makes the parsing of the table more stable, can avoid structural errors caused by information loss in the merged cell part, and improves the reliability of table reconstruction. Adaptability to different table layouts: The present invention can adapt to different table layouts and forms of merged cells, especially in the case where there are a large number of horizontally and vertically merged cells in the document. By accurately identifying the blank parts of the merged cells and combining the row and column span information, it can efficiently perform cross-page recognition and restoration of merged cells in different types of tables, ensuring that the finally recognized table structure is complete and conforms to the actual situation. In traditional methods, manually checking and correcting merged cells in the table is often time-consuming and error-prone, while the present invention can automatically complete this process efficiently and accurately, improving the efficiency and accuracy of table data processing. By accurately determining the types of merged cells and blank cells, the method of the present invention improves the accuracy of table reconstruction and the integrity of cross-page table recognition, and has significant advantages especially when dealing with documents containing complex merged cells.
[0109] Embodiment 5
[0110] Please refer to Figure 1 , specifically: The specific steps of S3 further include:
[0111] S34. Based on the determination of the types of blank cells in the complete table in S33, gradually determine the range of the merged cell form type as type two in the complete table, and based on the merged cell range of type one and the merged cell range of type two, re - improve the relevant data of the complete table. The relevant data of the complete table includes the positions of each horizontal line, vertical line, dot, number of columns, number of rows, column width, row width, text layer information, and the merged cell range in the complete table; based on the relevant data of the complete table, convert the complete table involved in the PDF financial report into structured table data, where the structured table data includes but is not limited to Excel, JSON, and CSV, and during the conversion process, perform a restoration task on the table, specifically:
[0112] ;
[0113] Where S is the reconstructed table, is the content of the cell ; is the row index of the cell in the complete table, and the value range is from 1 to Q; Q is the number of rows of the complete table, is the column index of the cell in the complete table, and the value range is from 1 to P; P is the number of columns of the complete table;
[0114] If , it means that the cell participates in the merge, and the content of the cell will be added to the reconstructed table;
[0115] If , it means that the cell does not participate in the merge, and the content of the cell will not be re - edited in the reconstructed table;
[0116] The purpose of this step is to reconstruct the content and structure of the merged cells by performing a weighted sum on the merge status of the content of each cell in the table.
[0117] In this embodiment, by determining the type based on blank cells and identifying the range of merged cells of type two, the present invention can gradually and accurately determine the exact range of each merged cell. Whether it is horizontal merging or vertical merging, the merging status of each cell can be clearly identified. This method ensures accurate calculation of the range of merged cells and will not affect subsequent data reconstruction due to misjudgment. Automatic restoration and reconstruction of tabular data: Based on the accurate identification of merged cells, the present invention can automatically convert the tables in PDF financial reports into structured data formats (such as Excel, JSON, or CSV). The restoration task in this process can ensure that the data of merged cells is correctly transferred to the target location, and the table content will not be lost or misaligned due to incorrect merging areas, thereby improving the accuracy of data extraction. Through this method, the generation process of structured data is highly automated, greatly reducing manual intervention and the occurrence of errors. Efficient restoration of table content and judgment of merging status: By performing weighted summation on the merging status of the content of each cell in the table, the method of the present invention can accurately restore the content of merged cells in the original table to the reconstructed structured table. The content involved in merged cells will be correctly added to the reconstructed table, while cells that are not involved in merging will not be re-edited. This step ensures the efficiency and accuracy of the table restoration process and guarantees the integrity and reliability of financial statement data. Precise data conversion and format compatibility: The reconstruction process of the present invention supports converting financial report tables into multiple structured data formats (such as Excel, JSON, CSV), providing convenience for subsequent data analysis, storage, and sharing. This multi-format support enables table data to be easily integrated into various analysis systems, providing a convenient operation experience for enterprises or financial personnel.
[0118] Enhancing the processing ability of multi-page tables: The present invention can efficiently process multi-page tables, ensuring that the content of the tables on multiple pages can be accurately spliced and restored. This advantage is particularly applicable to processing multi-page documents such as PDF financial reports, avoiding problems such as inconsistent or lost data caused by tables being scattered on multiple pages. The present invention provides an efficient and automated solution for the table data in PDF financial reports by accurately identifying merged cells and accurately reconstructing the table structure. Through weighted summation of merged cells, automatic data conversion, and reconstruction of structured tables, the accuracy and efficiency of table data extraction and restoration are significantly improved, while reducing the risk of manual intervention and errors, ensuring the integrity and efficiency of data processing.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for identifying a table in a document, characterized in that: The following steps are involved: S1. Based on the PDF financial report, multiple groups of image files are obtained, and the table area in each group of image files is identified to obtain the table data; S2. Based on the table data, analyze the similarity between the areas of each table, and perform cross-page recognition in combination with the similarity transitivity hypothesis to complete the complete table recognition task, wherein the similarity transitivity hypothesis specifically includes: if table A and table B belong to the same group of tables, and table B and table C belong to the same group of tables, then table A and table C also belong to the same group of tables, and through aggregation, multiple groups of complete tables are identified in multiple groups of image files, and the complete table includes at least one group of sub-tables; S3. Based on S2, preliminarily determine the relevant data of the complete table, and according to the type of blank cells in the complete table, gradually determine that the form type of the merged cells in the complete table is the range of type two, and according to the merged cell range of type one and the merged cell range of type two, re-perfect the relevant data of the complete table, the relevant data of the complete table includes the horizontal line position, vertical line position and point position, number of columns, number of rows, column width, row width, text layer information and merged cell range in the complete table; based on the relevant data of the complete table, convert the complete table involved in the PDF financial report into structured table data to realize the table restoration task, wherein the form type includes type one and type two, the merged cell of form type one refers to the merged cell without blank cells, the merged cell of form type two refers to the merged cell with blank cells, and the type of blank cells refers to whether they exist in the merged cells.
2. The method for identifying a table in a document according to claim 1, characterized in that: S1 specific steps include: S11, obtaining the PDF financial report in advance from the company's finance department, converting the PDF financial report document into multiple sets of image files, and extracting high-resolution images and text layer information from the multiple sets of image files, wherein the text layer information includes text content, text position, font, and font size; S12. Use the edge detection kernel to perform convolution operation on the high-resolution image to obtain the edge strength of each pixel , the edge strength The way to obtain is as follows: ; In the formula, For the image at coordinates The pixel value at ; is the weight of the edge detection kernel, which is defined by the selected edge detection kernel; and are the horizontal and vertical offsets of the ith position in the edge detection kernel relative to the center position; i is an index variable, indicating the ith position in the edge detection kernel currently being processed; and n is the total number of elements in the edge detection kernel.
3. The method for identifying a table in a document according to claim 2, characterized in that: The specific steps of S1 also include: S13, using the Otsu algorithm, the optimal threshold is set according to the maximum inter-class variance, and the threshold is combined with the edge strength of each pixel. Compare the size of the corresponding pixel points. When the set threshold is exceeded, the corresponding pixel point is used as the edge point of the table, and the straight lines in the image are detected by Hough transform, the horizontal and vertical lines of the table in the image are found, and the boundary lines in the table are determined according to the intersections between the horizontal and vertical lines of the table to generate table data, wherein the table data includes the horizontal line positions, vertical line positions and intersection positions of the table in the image in each group of image files.
4. The method for identifying a table in a document according to claim 3, characterized in that: The specific steps of S2 include: S21, based on the table data, determining the number of columns, the number of rows, the column width and the row width of the table in each group of image files, and combining the text layer information, constructing the feature vector H of each table area, and determining the adjacent relationship between the tables according to the order of the multiple groups of image files converted from the PDF financial report; S22, identifying the similarity between adjacent tables based on the feature vector H of each table area to obtain the similarity , specifically: ; In the formula, is the eigenvector of the jth table region, is the eigenvector of the j+1th table region, is the feature vector and the eigenvector The dot product of is the feature vector The norm of .
5. The method for identifying a table in a document according to claim 4, characterized in that: The specific steps of S2 also include: S23, pre-set the similarity threshold, and Compare to determine whether there is cross-page behavior between corresponding adjacent tables. The specific comparison content is as follows: If the similarity If the similarity threshold is exceeded, it is determined that there is cross-page behavior between the corresponding adjacent tables, indicating that the current adjacent tables belong to the same group of tables, and the current table is marked as a subtable; If the similarity If the similarity threshold is not exceeded, it is determined that there is no cross-page behavior between the corresponding adjacent tables, indicating that the current adjacent tables do not belong to the same group of tables.
6. The method for identifying a table in a document according to claim 5, characterized in that: The specific steps of S3 include: S31, splicing the sub-tables in the complete table to determine the coordinates of each intersection in the complete table, and dividing the complete table into a plurality of groups of cells according to the coordinates of each intersection in the complete table.
7. A method for identifying a table in a document according to claim 6, characterized in that: The specific steps of S3 also include: S32, traversing all cells in the complete table, if there are merged cells in the complete table, and the form type of the merged cells is type 1, identifying the merged cells by the following steps: S321, starting from the starting point in the upper left corner of the complete table, expanding to the right until encountering the boundary line, and recording the number of horizontally merged columns Hs, specifically: ; In the formula, For cells In line and the boundary line at column p+r; For cells The number of columns to merge horizontally; To obtain the maximum r value, There are no boundary lines in the location; For the cell The boundary line of the position does not exist, where the subscript Translated into Chinese, it means non-existence; r is the number of columns to check to the right from the current position; S322, starting from the starting point in the upper left corner of the complete table, extending downward until encountering a boundary line, recording the number of vertically merged rows Zs, specifically: ; In the formula, For cells In line and column The boundary line of the location, For cells The number of rows to merge vertically; To obtain the maximum value of e, There are no boundary lines in the location; For the cell The boundary line of the position does not exist, where the subscript Translated into Chinese, it means non-existence; e is the number of rows to check from the current position downwards; S323: Based on the contents of S321 and S322, determine that the form type of the merged cells in the complete table is a range F of type 1, the range .
8. The method for identifying a table in a document according to claim 7, characterized in that: The specific steps of S3 also include: S33, traversing all cells in the complete table, if there are merged cells in the complete table, and the form type of the merged cells is type 2, identifying the merged cells in the following manner: ; In the formula, is the result of the judgment, indicating that the cell Is it a blank cell in the merged cell? Represents a cell is part of a merged cell, Represents a cell It is an independent, truly blank cell that does not belong to any merged cell area; Represents a cell It does not span other columns and is an independent cell; Represents a cell spans multiple columns; Represents a cell spans multiple rows; Represents a cell No other rows are crossed; Represents a cell The merge status of Represents a cell It is not the first cell of the merged cell. Represents a cell is the first cell of the merged cell; , , indicating a cell Blank cells that are part of a merged cell; , , indicating a cell Is a true blank cell that does not belong to any merged area.
9. The method for identifying a table in a document according to claim 8, characterized in that: The specific steps of S3 also include: S34. Structured table data includes Excel, JSON and CSV. During the conversion process, the table is restored, specifically: ; In the formula, S is the reconstructed table, For cells content, For cells The row index in the complete table ranges from 1 to Q; Q is the number of rows in the complete table. For cells The column index in the complete table ranges from 1 to P, where P is the number of columns in the complete table. like , which means that the cell Participated in the merge, cell Contents will be added to the reconstruction table; like , which means that the cell Not involved in the merge, cell Contents No re-editing will be done in the rebuilt form.
Citation Information
Patent Citations
Announcement document table data identification method and terminal
CN112434496A
Table standardization processing method and device, equipment and storage medium
CN113033170A