Table identification method and device, electronic equipment, storage medium and program product
By identifying the structure and content of the recognition table, determining the correspondence between the logical position of the cell and the text content, and reconstructing the table, solving the problem of poor generalization ability of table recognition in the prior art, and improving the accuracy and generalization ability of recognition.
Patent Information
- Application Number
- CN202510143093.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-23
AI Technical Summary
The existing table recognition scheme has poor generalization ability and low recognition accuracy when encountering new types of tables.
By identifying the structure and content recognition tables, the logical position and physical position of the cell, as well as the text content and physical position of the text area, determine the correspondence between the logical position of the cell and the text content, and then reconstruct the table.
It realizes the recognition of borderless and partial border tables, effectively identifying complex table structures such as cross rows and cross columns, greatly improving the generalization ability and accuracy of table recognition.
Smart Images

Figure CN120032384A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present specification relate to the field of information processing technology, and in particular, to a table recognition method, device, electronic device, storage medium, and program product. Background Art
[0002] As a way to organize and present data, tables play an indispensable role in all walks of life, helping professionals quickly understand and analyze complex information. Tables in different industries (such as financial statements, scientific papers, medical records, etc.) have different characteristics. Current table recognition solutions have poor generalization capabilities and perform poorly when encountering new types of tables, with low recognition accuracy. Summary of the invention
[0003] In view of this, one or more embodiments of this specification provide the following technical solutions:
[0004] According to a first aspect of one or more embodiments of this specification, a table recognition method is proposed, including:
[0005] Get the table to be identified;
[0006] Performing structural recognition on the table to be recognized to obtain a cell logical position and a cell physical position of each cell in the table to be recognized;
[0007] Performing content recognition on the table to be recognized to obtain the text content and physical location of each text area in the table to be recognized;
[0008] Determine the correspondence between the cell logical position and the text content according to the cell physical position and the text area physical position;
[0009] The table to be identified is reconstructed according to the corresponding relationship to obtain a reconstructed table.
[0010] According to a second aspect of one or more embodiments of this specification, a table recognition device is provided, comprising:
[0011] A table acquisition unit acquires a table to be identified;
[0012] A structure recognition unit performs structure recognition on the table to be recognized to obtain a cell logical position and a cell physical position of each cell in the table to be recognized;
[0013] A content recognition unit performs content recognition on the table to be recognized to obtain the text content and physical location of each text area in the table to be recognized;
[0014] a relationship determination unit, which determines a corresponding relationship between the logical position of the cell and the text content according to the physical position of the cell and the physical position of the text area;
[0015] The table reconstruction unit reconstructs the table to be identified according to the corresponding relationship to obtain a reconstructed table.
[0016] According to a third aspect of one or more embodiments of the present specification, an electronic device is proposed, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor implements the steps of the aforementioned method by running the executable instructions.
[0017] According to a fourth aspect of one or more embodiments of the present specification, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the aforementioned method are implemented.
[0018] According to a fifth aspect of one or more embodiments of the present specification, a computer program product is proposed, comprising a computer program / instruction, which implements the steps of the aforementioned method when executed by a processor.
[0019] It can be seen from the above description that when performing table recognition, this manual, on the one hand, performs table structure recognition on the table to be recognized, and obtains the cell physical position of the cell logical position of each cell, and on the other hand, performs table content recognition on the table to be recognized, and obtains the text content and physical position of each text area. Then, the correspondence between the cell logical position and the text content can be determined according to the cell physical position and the text area physical position, and then the table to be recognized is reconstructed according to the correspondence to obtain a reconstructed table.
[0020] It can be seen that the table recognition solution provided in this manual reconstructs the table by determining the correspondence between the logical position of the cell and the text content. It does not rely on the line detection method and can realize the recognition of borderless and partial border tables. At the same time, the table is reconstructed based on the logical position of the cell. It can effectively identify complex table structures such as those across rows and columns, which greatly improves the generalization ability and accuracy of table recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a schematic diagram of the architecture of a table recognition service system provided by an exemplary embodiment.
[0022] Figure 2 It is a flowchart of a table recognition method provided by an exemplary embodiment.
[0023] Figure 3 is a flowchart of another table recognition method provided by an exemplary embodiment.
[0024] Figure 4 It is a schematic diagram of another table recognition method provided by an exemplary embodiment.
[0025] Figure 5 It is a flowchart of a method for evaluating the accuracy of a reconstruction table provided by an exemplary embodiment.
[0026] Figure 6 It is a flowchart of a method for calculating content similarity between a reconstructed table and a table to be identified provided by an exemplary embodiment.
[0027] Figure 7 It is a structural schematic diagram of a device provided by an exemplary embodiment.
[0028] Figure 8 It is a block diagram of a table recognition device provided by an exemplary embodiment. DETAILED DESCRIPTION
[0029] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0030] As a way to organize and present data, tables play an indispensable role in all walks of life, helping professionals quickly understand and analyze complex information. Tables in different industries (such as financial statements, scientific papers, medical records, etc.) have different characteristics. Current table recognition solutions have poor generalization capabilities and perform poorly when encountering new types of tables, with low recognition accuracy.
[0031] Figure 1 FIG. 1 is a schematic diagram of the architecture of a table recognition service system provided by an exemplary embodiment. Figure 1 As shown, the system may include a server 11, a network 12, and several electronic devices, such as a PC (Personal Computer) 13, a mobile phone 14, and the like.
[0032] The server 11 may be a physical server including an independent host, or the server 11 may be a virtual server carried by a host cluster. During operation, the server 11 may run a server-side program of a certain application to implement the relevant functions of the application. For example, when the server 11 runs a program of a table recognition service, it may be implemented as a corresponding table recognition service platform.
[0033] PC13 and mobile phone 14 are only some types of electronic devices that users can use. In fact, users can obviously also use electronic devices such as the following types: tablet devices, laptops, PDAs (Personal Digital Assistants), wearable devices (such as smart glasses, smart watches, etc.), etc., and one or more embodiments of this specification do not limit this. During operation, the electronic device can run a program on the client side of a certain application to implement the relevant functions of the application. For example, when the electronic device runs a program for a table recognition service, it can be implemented as a client of the table recognition service. Among them, the application program of the client of the above-mentioned table recognition service can be started and run on the electronic device. The program on the client side can be a native application installed on the electronic device, or the program on the client side can be a small program, a quick application or other similar forms. Of course, when using web page technologies such as HTML5 or similar, the relevant functions can be implemented through the page displayed by the browser. The browser here can be an independent browser application or a browser module embedded in certain applications.
[0034] As for the network 12 for interaction between electronic devices such as PC 13 and mobile phone 14 and server 11, it is possible to select a wired or wireless network to achieve communication based on the communication mode supported by the corresponding electronic device, and this specification does not limit this. For example, PC 13 can support both wired and wireless communication, so it can use a wired or wireless network to achieve communication as needed, while mobile phone 14 usually only supports wireless communication, so it can use a wireless network to achieve communication.
[0035] Figure 2 It is a flowchart of a table recognition method provided by an exemplary embodiment.
[0036] Please refer to Figure 2 The table recognition method can be applied to the aforementioned Figure 1 The server shown in FIG. 1 can also be applied to the aforementioned Figure 1 The electronic device shown includes the following steps:
[0037] Step 202: Obtain a table to be identified.
[0038] In this embodiment, the form to be identified can be a form from various industries such as financial statements, scientific papers, medical records, insurance policies, etc. The form to be identified can be in image format or PDF format (Portable Document Format), and this manual does not impose any special restrictions on this.
[0039] In this embodiment, the user can upload a form that needs to be recognized and reconstructed. After receiving the form uploaded by the user, the form can be cropped to remove the background area to obtain a form to be recognized, so as to facilitate subsequent processing.
[0040] Step 204 , performing structural recognition on the table to be recognized, and obtaining a cell logical position and a cell physical position of each cell in the table to be recognized.
[0041] In this embodiment, the table to be identified usually includes multiple cells, and the structure of the table to be identified can be identified by an image recognition algorithm to identify the logical position and physical position of each cell in the table to be identified, which are called cell logical position and cell physical position.
[0042] The cell logical position is the abstract position of the cell in the entire table, representing the specific position of the cell in the table row and column, and can be the starting and ending row and column numbers starting from 0. For example, the cell logical position A can represent that cell A is located in the 1st row and 2nd column in the table, etc.
[0043] The physical position of the cell is the geometric position of the cell in the table image, which can be represented by the coordinates of the cell corner point, or by the coordinates of the cell center point combined with the height and width of the cell.
[0044] In this embodiment, a trained table structure recognition model may be used to recognize the cell logical position and the cell physical position of each cell in the table to be recognized.
[0045] For example, the table structure recognition model can first identify the physical position of the center point of each cell in the table to be identified through the center point detection network, and the physical position of the center point can be represented by the geometric coordinates of the center point, the height and width of the cell to which it belongs, and then predict the logical position of the cell and the physical position of the cell respectively through two branches. Among them, the branch for predicting the logical position of the cell realizes the prediction of the logical position of the cell by constraining the logical relationship within the cell and between cells.
[0046] Of course, in other examples, other methods may be used to identify the cell logical position and the cell physical position of each cell in the table to be identified, and this specification does not impose any special limitation on this.
[0047] Step 206: Perform content recognition on the table to be recognized to obtain the text content and physical location of each text area in the table to be recognized.
[0048] In this embodiment, the text content and the physical location of each text area in the table to be recognized can be recognized by using corresponding text recognition technology according to the format of the table to be recognized.
[0049] Among them, the text area is the area where the text content is located, which can be the minimum circumscribed rectangle of the text content. The physical position of the text area is the geometric position of the text area, which is similar to the physical position of the aforementioned cell. The physical position of the text area can be represented by the coordinates of the corner points of the text area, or by the coordinates of the center point of the text area combined with the height and width of the text area.
[0050] Step 208: Determine the correspondence between the cell logical position and the text content according to the cell physical position and the text area physical position.
[0051] In this embodiment, the correspondence between the cell and the text area can be determined according to the physical position of the cell and the physical position of the text area, and then the correspondence between the logical position of the cell and the text content can be determined.
[0052] Specifically, for each cell, the intersection over union (IoU) of the cell and each text area can be calculated according to the physical position of the cell and the physical position of each text area, and then multiple IoUs are obtained, and then the text area with the maximum IoU calculated is determined as the text area corresponding to the cell, and then the corresponding text area is determined for each cell, that is, the correspondence between the cell and the text area is obtained, and then the correspondence between the logical position of the cell and the text content of the corresponding text area can be established, and then the correspondence between the logical position of the cell and the text content is obtained. For example, the logical position of the cell indicates that the text content corresponding to the cell in the first row and the second column is "limited number of copies purchased", etc.
[0053] In another example, the text area can be used as a reference, and the intersection and union ratio between the text area and each cell can be calculated according to the physical position of the text area and the physical position of each cell. Then, the cell with the largest intersection and union ratio can be determined as the cell corresponding to the text area. Similarly, the corresponding relationship between the cell and the text area can be obtained.
[0054] In another example, the intersection-over-union ratio may not be calculated. For example, the distance between the cell and the text area may be calculated based on the physical position of the cell and the physical position of the text area, and then the correspondence between the cell and the text area may be determined based on the distance. This manual does not impose any special restrictions on this.
[0055] Step 210: reconstruct the table to be identified according to the corresponding relationship to obtain a reconstructed table.
[0056] Based on the aforementioned step 208, after the correspondence between the cell logical position and the text content is determined, the table to be recognized can be reconstructed based on the correspondence to obtain a reconstructed table, thereby realizing recognition of the table to be recognized.
[0057] It can be seen from the above description that when performing table recognition, this manual, on the one hand, performs table structure recognition on the table to be recognized, and obtains the cell physical position of the cell logical position of each cell, and on the other hand, performs table content recognition on the table to be recognized, and obtains the text content and physical position of each text area. Then, the correspondence between the cell logical position and the text content can be determined according to the cell physical position and the text area physical position, and then the table to be recognized is reconstructed according to the correspondence to obtain a reconstructed table.
[0058] It can be seen that the table recognition solution provided in this manual reconstructs the table by determining the correspondence between the logical position of the cell and the text content. It does not rely on the line detection method and can realize the recognition of borderless and partial border tables. At the same time, the table is reconstructed based on the logical position of the cell. It can effectively identify complex table structures such as those across rows and columns, which greatly improves the generalization ability and accuracy of table recognition.
[0059] Figure 3 is a flowchart of another table recognition method provided by an exemplary embodiment.
[0060] Please refer to Figure 3 and Figure 4 , the table recognition method may include the following steps:
[0061] Step 302: Obtain a table to be identified.
[0062] Step 304: perform structural recognition on the table to be recognized to obtain the cell logical position and the cell physical position of each cell in the table to be recognized.
[0063] In this embodiment, the implementation of step 302 and step 304 can refer to the aforementioned Figure 2 The embodiments shown in this specification will not be described in detail here.
[0064]
[0065] Table 1
[0066] Please refer to the table to be identified shown in Table 1, which includes 24 cells. By identifying the table structure, the cell logical position and cell physical position of each cell in the table to be identified can be identified, that is, the cell logical position and cell physical position of the 24 cells can be identified.
[0067] This embodiment adopts CL i To represent the logical position of the cell, CL i =(sr i ,er i , sc i ,ec i ), where sr i Represents the starting row of the ith cell (start-row), er i Represents the end-row of the ith cell, sc i Represents the start column of the ith cell (start-column), ec i Represents the end-column of the i-th cell. The value of i is 0 to 23.
[0068] Please continue to refer to Table 1, the cell logical position CL of the upper left cell i =(0,0,0,0), which means that the starting row and ending row of the cell are both row 1, and the starting column and ending column are both column 1, that is, the cell is located in row 1 and column 1. The cell logical position CL of the lower right corner cell i =(5, 8, 3, 3), which means that the starting row of the cell is the 6th row, the ending row is the 9th row, the starting column is the 4th column, and the ending column is also the 4th column, that is, the cell is a cross-row cell, occupying rows 6-9 and located in the 4th column.
[0069] This embodiment adopts CP i To represent the physical location of the cell, CP i =(x i ,y i , W i , H i ), where (x i ,y i ) represents the geometric coordinates of the center point of the i-th cell, (W i , H i ) represents the width and height of the i-th cell, and the value of i is 0 to 23.
[0070] Of course, in other examples, the geometric coordinates of the upper left corner and the lower right corner of the cell may also be used as the physical position of the cell.
[0071] Step 306: Determine the target format of the table to be identified.
[0072] In this embodiment, before performing content recognition on the table to be recognized, the format of the table to be recognized may be determined first, which is referred to as a target format for easy distinction. The target format may include PDF and images.
[0073] Step 308: query the target text recognition technology corresponding to the target format in the correspondence between the format and the text recognition technology.
[0074] In this embodiment, different text recognition technologies can be used to recognize the content of different table formats. Specifically, the correspondence between the format and the text recognition technology can be pre-stored. After determining the target format of the table to be recognized, the text recognition technology corresponding to the target format can be queried in the correspondence, which is called the target text recognition technology.
[0075] Table Format Text recognition technology picture OCR PDF Camelot
[0076] Table 2
[0077] Please refer to the correspondence between the formats and text recognition technologies shown in Table 2. If the target format of the table to be recognized is an image format, it can be found that the target text recognition technology corresponding to the image format is OCR (Optical Character Recognition) technology; if the target format of the table to be recognized is a PDF format, it can be found that the target text recognition technology corresponding to the PDF format is Camelot (a Python library designed specifically for PDF table data extraction).
[0078] It is worth noting that Table 2 is only for illustrative purposes. In other examples, such a table may not be organized, or other text recognition technologies may be set for image formats and PDF formats. This specification does not impose any special restrictions on this.
[0079] Step 310: Use the target text recognition technology to perform content recognition on the table to be recognized, and obtain the text content and physical location of each text area in the table to be recognized.
[0080] Based on the aforementioned step 308, after the target text recognition technology is found, the target text recognition technology can be used to perform content recognition on the table to be recognized. In actual implementation, the interface provided by the target text recognition technology can be called to realize content recognition of the table to be recognized.
[0081] Please continue to refer to the table to be identified shown in Table 1. By performing content recognition on the table, the text content and physical locations of 24 text areas can be identified.
[0082] In this embodiment, TP j To represent the physical location of the text area, TP j =(x j ,y j , W j , Hj ), where (x j ,y j ) represents the geometric coordinates of the center point of the jth text area, (W j , H j ) represents the width and height of the jth text area, and the value of j is 0 to 23.
[0083] The text content of the text area in the upper left corner of the table to be identified shown in Table 1 is "Plan Name", and the physical location of the text area is TP 1 .
[0084] Step 312: for each cell, calculate the intersection-over-union ratio between the cell and each text region according to the cell physical position of the cell and the physical position of each text region.
[0085] In this embodiment, after the table structure and content are identified in the table to be identified, for each cell, the intersection-over-combination ratio of the cell and each text area can be calculated according to the physical position of the cell and the physical position of each text area. 1 With TP 1 Calculate the intersection and union ratio of this cell and the first text area, according to CP 1 With TP 2 The intersection-and-union ratio of the cell and the second text region is calculated, and so on, the intersection-and-union ratio of the cell and the 24th text region is calculated.
[0086] Similarly, for each cell, its intersection-over-union ratio with each text area can be calculated.
[0087] Step 314: determine the text region with the maximum intersection-over-union ratio calculated as the text region corresponding to the cell.
[0088] In this embodiment, for each cell, the calculated IoU ratios may be arranged in descending order, and then the text region corresponding to the largest IoU ratio arranged in the first position is determined as the text region corresponding to the cell. Similarly, the corresponding text region may be determined for each cell, that is, the corresponding relationship between the cell and the text region is obtained.
[0089] Step 316: Establish a correspondence between the cell logical position of the cell and the text content of the text area corresponding to the cell.
[0090] Based on the aforementioned step 314, after the correspondence between the cell and the text area is obtained, the correspondence between the cell logical position and the text content can be established.
[0091] Cell Logical cell position Text Area Text content Cell 1 <![CDATA[CL 1 ]]> Text Area 1 Program Name Cell 2 <![CDATA[CL 2 ]]> Text Area 2 Guarantee responsibility Cell 3 <![CDATA[CL 3 ]]> Text Area 3 Sum Insured
[0092] Table 3
[0093] Please refer to the example in Table 3. Assuming that cell 1 corresponds to text area 1, the cell logical position CL of cell 1 can be established. 1 The corresponding relationship between the text content "plan name" in text area 1, similarly, the cell logical position CL can be established 2 The corresponding relationship between the text content "Guarantee Responsibility" in text area 1 can establish the cell logical position CL 3 The correspondence between the text content "insurance amount" in text area 1.
[0094] Step 318: reconstruct the table to be identified according to the corresponding relationship to obtain a reconstructed table.
[0095] In this embodiment, the table to be recognized may be reconstructed according to the correspondence between the logical position of the cell and the text content. For the convenience of distinction, the table obtained after reconstruction may be referred to as a reconstructed table.
[0096] For example, the cell logical position CL 1 The representative cell is located in the 1st row and the 1st column. According to its corresponding text content "Plan Name", it can be determined that the first cell of the reconstructed table is located in the 1st row and the 1st column, and the content is "Plan Name", thereby realizing the construction of the reconstructed table.
[0097] It can be seen that the table recognition solution provided in this manual reconstructs the table by determining the correspondence between the logical position of the cell and the text content. It does not rely on the line detection method and can realize the recognition of borderless and partial border tables. At the same time, the table is reconstructed based on the logical position of the cell. It can effectively identify complex table structures such as those across rows and columns, which greatly improves the generalization ability and accuracy of table recognition.
[0098] The present specification also provides a method for evaluating the accuracy of a reconstructed table, which can evaluate the accuracy of the reconstructed table after completing the recognition and reconstruction of the table to be recognized.
[0099] Please refer to Figure 5 , the accuracy assessment of the reconstructed table may include the following steps:
[0100] Step 502, determining the number of rows and columns to be reconstructed in the reconstructed table.
[0101] In this embodiment, the number of rows and columns of the reconstructed table can be determined, that is, the number of rows and columns of the reconstructed table, which is called the reconstructed row-column number. The reconstructed table can be an html (Hyper Text Markup Language) string. In this step, the html string representing the reconstructed table can be converted into a list, and the number of rows and columns of the reconstructed table can be obtained from the list.
[0102] Step 504, compare the reconstructed row-column number with the standard row-column number of the table to be recognized.
[0103] Based on the foregoing step 502, after determining the foregoing reconstructed row-column number, the reconstructed row-column number can be compared with the row-column number of the table to be recognized (referred to as the standard row-column number) to determine whether the reconstructed row-column number is the same as the standard row-column number.
[0104] Among them, the standard row-column number of the table to be recognized can be determined based on the logical positions of its cells. For example, it can be determined based on the logical position of the bottom-right cell. Still taking the table to be recognized shown in Table 1 as an example, the logical position CL i =(5, 8, 3, 3), indicating that the table to be recognized is a table with 9 rows and 4 columns.
[0105] If the reconstructed row-column number is the same as the standard row-column number, that is, the number of rows of the reconstructed table is the same as the number of rows of the table to be recognized, and the number of columns of the reconstructed table is the same as the number of columns of the table to be recognized, it means that the structure of the reconstructed table is consistent with the structure of the table to be recognized, then step 506 can be executed.
[0106] If the reconstructed row-column number is not the same as the standard row-column number, for example, the number of rows of the reconstructed table is not the same as the number of rows of the table to be recognized, or the number of columns of the reconstructed table is not the same as the number of columns of the table to be recognized, or both the number of rows and columns are not the same, it means that the structures of the reconstructed table and the table to be recognized are not completely consistent, and step 510 can be executed.
[0107] Step 506, when the reconstructed row-column number is the same as the standard row-column number, calculate the content similarity between the reconstructed table and the table to be recognized.
[0108] Based on the judgment result of the foregoing step 504, when the reconstructed row-column number is the same as the standard row-column number, it means that the structure of the reconstructed table is consistent with the structure of the table to be recognized, and the accuracy rate of the reconstructed table can be determined based on the content similarity between the reconstructed table and the table to be recognized.
[0109] Please refer to Figure 6 , the method for calculating the content similarity between the reconstructed table and the table to be recognized may include the following steps:
[0110] Step 5062: for each cell, compare the first type of text content in the cell in the reconstructed table with the second type of text content in the same cell in the table to be identified.
[0111] In this embodiment, when calculating the content similarity between the reconstructed table and the table to be identified, the content similarity of each cell may be calculated respectively, and then the content similarities of each cell are summarized to obtain the content similarity between the reconstructed table and the table to be identified.
[0112] In this embodiment, for the convenience of distinction, the text content of each cell in the reconstructed table can be referred to as the first type of text content, and the text content of each cell in the table to be identified can be referred to as the second type of text content. Specifically, for each cell, the first type of text content of the cell in the reconstructed table and the second type of text content of the cell in the table to be identified can be compared to see whether they are the same. When determining whether the text content is the same, any text similarity algorithm can be used, and this specification does not impose any special restrictions on this.
[0113] Step 5064: Determine the content similarity of the cell according to the comparison result between the first category text content and the second category text content.
[0114] In this embodiment, the content similarity of each cell can be determined based on the comparison result of the aforementioned step 5062. In a simpler way, if the first-category text content and the second-category text content of a cell are completely identical (e.g., the text similarity is 100%), the content similarity of the cell can be determined as 1; if the first-category text content and the second-category text content of the cell are not completely identical (e.g., the text similarity is not 100%), the content similarity of the cell can be determined as 0.
[0115] In this embodiment, the content similarity of each cell may be determined.
[0116] Step 5066, summarizing the content similarity of each cell to obtain the content similarity between the reconstructed table and the table to be identified.
[0117] In this embodiment, the average value of the content similarity of each cell may be calculated as the content similarity between the reconstructed table and the table to be identified.
[0118] Of course, in other examples, the sum of the similarities of the contents of each cell may also be calculated as the content similarity between the reconstructed table and the table to be identified, and this specification does not impose any special limitation on this.
[0119] Step 508: Determine the accuracy of the reconstructed table according to the content similarity.
[0120] In this embodiment, the accuracy of the reconstructed table may be determined according to the content similarity between the reconstructed table and the table to be identified calculated in the aforementioned step 506 .
[0121] For example, if the average value of the content similarities of each cell is taken as the content similarity between the reconstructed table and the table to be identified, the content similarity can be determined as the accuracy of the reconstructed table.
[0122] For another example, if the sum of the similarities of the contents of each cell is used as the content similarity between the reconstructed table and the table to be identified, the sum can be used as the accuracy of the reconstructed table after performing corresponding operations, and this specification does not impose any special restrictions on this.
[0123] Step 510: When the reconstructed number of rows and columns is different from the standard number of rows and columns, a tree edit distance between the reconstructed table and the table to be identified is calculated as the accuracy of the reconstructed table.
[0124] Based on the judgment result of the aforementioned step 504, when the reconstructed number of rows and columns is different from the standard number of rows and columns, it means that the structure of the reconstructed table and the structure of the table to be identified are not completely consistent. The Tree-Edit-Distance-based Similarity (TEDS) between the reconstructed table and the table to be identified can be calculated as the accuracy of the reconstructed table.
[0125] It can be seen from the above description that when calculating the accuracy of the reconstructed table, this manual first compares the number of reconstructed rows and columns of the reconstructed table with the standard number of rows and columns of the table to be identified, and when the number of reconstructed rows and columns and the standard number of rows and columns are the same, the accuracy of the reconstructed table can be determined according to the content similarity between the reconstructed table and the table to be identified. Compared with the method of directly using the tree edit distance TEDS to calculate the accuracy of the reconstructed table, the efficiency of calculating the accuracy of the reconstructed table can be greatly improved.
[0126] In this embodiment, the user can evaluate the recognition effect of the reconstructed table based on the accuracy of the reconstructed table. If the recognition effect is not satisfactory, the user can split the table to be recognized into multiple smaller small tables, and then use the table recognition solution provided in this specification to recognize them. In the case of table splitting, the table recognition solution provided in this specification can merge the multiple small reconstructed tables obtained by recognition after each small table is recognized to obtain a small reconstructed table, thereby obtaining the recognition result of the large table before splitting.
[0127] This specification also provides a table recognition solution, which can automatically split and reconstruct a larger table to be recognized, thereby improving the accuracy of table recognition.
[0128] Specifically, after obtaining the table to be recognized, the standard number of rows and columns of the table to be recognized can be determined. The standard number of rows and columns can be determined based on the logical position of the bottom-right cell using the aforementioned method. Of course, in other examples, the standard number of rows and columns of the table to be recognized can also be determined using relevant tools before performing table structure recognition.
[0129] Next, a preset threshold for the number of rows and columns can be obtained, and then it can be determined whether the standard number of rows and columns exceeds the threshold for the number of rows and columns. For example, it can be determined separately whether the standard number of rows in the standard number of rows and columns exceeds the threshold for the number of rows, and whether the standard number of columns in the standard number of rows and columns exceeds the threshold for the number of columns. If the standard number of rows exceeds the threshold for the number of rows, or the standard number of columns exceeds the threshold for the number of columns, it can be determined that the standard number of rows and columns exceeds the threshold for the number of rows and columns.
[0130] In this embodiment, when the standard number of rows and columns exceeds the threshold for the number of rows and columns, it indicates that the table to be recognized is large, and the table to be recognized can be split into multiple small tables (referred to as sub-tables to be recognized), and then the method described in the foregoing embodiment can be used to separately recognize and reconstruct each sub-table to be recognized, and a corresponding reconstructed sub-table can be recognized and reconstructed for each sub-table to be recognized.
[0131] Among them, when splitting the table to be recognized, it can be split based on rows, or based on columns, or simultaneously based on rows and columns, and specifically can be determined according to the comparison between the standard number of rows and columns and the threshold for the number of rows and columns.
[0132] For example, assume that the standard number of rows exceeds the threshold for the number of rows, and the standard number of columns does not exceed the threshold for the number of columns, then the table to be recognized can be split based on rows. Assume that the standard number of rows of the table to be recognized is 100 rows, and the threshold for the number of rows is 50 rows, then the table to be recognized can be split into two sub-tables to be recognized with 50 rows each.
[0133] For another example, assume that the standard number of columns exceeds the threshold for the number of columns, and the standard number of rows does not exceed the threshold for the number of rows, then the table to be recognized can be split based on columns. Assume that the standard number of columns of the table to be recognized is 150 columns, and the threshold for the number of columns is 50 columns, then the table to be recognized can be split into three sub-tables to be recognized with 50 columns each.
[0134] For yet another example, assume that the standard number of rows exceeds the threshold for the number of rows, and at the same time the standard number of columns exceeds the threshold for the number of columns, then the table to be recognized can be split based on rows and columns respectively. Assume that the standard number of rows of the table to be recognized is 100 rows, and the standard number of columns is 150 columns, then the table to be recognized can be split into 6 sub-tables to be recognized with 50 rows and 50 columns each.
[0135] In this embodiment, after reconstructing the sub-table corresponding to each sub-table to be identified, these sub-tables may be merged to obtain the reconstructed table corresponding to the table to be identified.
[0136] It can be seen that this embodiment can split the table to be identified into multiple small sub-tables to be identified, and then identify and reconstruct each sub-table to be identified and then merge them, which can reduce the processing volume of the device and improve the accuracy of table recognition.
[0137] Figure 7 is a schematic structural diagram of a device provided by an exemplary embodiment. Figure 7 At the hardware level, the device includes a processor 702, an internal bus 704, a network interface 706, a memory 708, and a non-volatile memory 710, and may also include hardware required for other functions. One or more embodiments of this specification may be implemented based on software, such as the processor 702 reading the corresponding computer program from the non-volatile memory 710 into the memory 708 and then running it. Of course, in addition to the software implementation, one or more embodiments of this specification do not exclude other implementations, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0138] Please refer to Figure 8 The table recognition device 800 can be applied to Figure 7 The table recognition device 800 may include:
[0139] The table acquisition unit 801 acquires the table to be identified;
[0140] The structure recognition unit 802 performs structure recognition on the table to be recognized, and obtains the cell logical position and the cell physical position of each cell in the table to be recognized;
[0141] The content recognition unit 803 performs content recognition on the table to be recognized to obtain the text content and physical location of each text area in the table to be recognized;
[0142] A relationship determination unit 804 determines a corresponding relationship between the cell logical position and the text content according to the cell physical position and the text area physical position;
[0143] The table reconstruction unit 805 reconstructs the table to be identified according to the corresponding relationship to obtain a reconstructed table.
[0144] Optionally, before the content recognition unit 803 performs content recognition on the table to be recognized, the process further includes:
[0145] Determining the target format of the table to be identified;
[0146] The content recognition unit 803 specifically searches for a target text recognition technology corresponding to the target format in the correspondence between formats and text recognition technologies; and uses the target text recognition technology to perform content recognition on the table to be recognized.
[0147] Optionally, the process in which the relationship determination unit 804 determines the corresponding relationship between the cell logical position and the text content according to the cell physical position and the text area physical position includes:
[0148] For each cell, calculating an intersection-over-union ratio between the cell and each text area according to the cell physical position of the cell and the physical position of each text area;
[0149] Determine the text area with the maximum intersection-over-union ratio calculated as the text area corresponding to the cell;
[0150] A correspondence between the cell logical position of the cell and the text content of the text area corresponding to the cell is established.
[0151] Optionally, the table recognition device 800 further includes:
[0152] The accuracy determination unit 806 determines the number of reconstructed rows and columns of the reconstructed table; compares the number of reconstructed rows and columns with the standard number of rows and columns of the table to be identified; when the number of reconstructed rows and columns is the same as the standard number of rows and columns, calculates the content similarity between the reconstructed table and the table to be identified; and determines the accuracy of the reconstructed table based on the content similarity.
[0153] Optionally, the process of calculating the content similarity between the reconstructed table and the table to be identified by the accuracy determination unit 806 includes:
[0154] For each cell, comparing the first type of text content in the cell in the reconstructed table with the second type of text content in the same cell in the to-be-identified table;
[0155] Determining the content similarity of the cell according to a comparison result between the first category of text content and the second category of text content;
[0156] The content similarities of each cell are summarized to obtain the content similarities between the reconstructed table and the table to be identified.
[0157] Optionally, the accuracy determination unit 806 further calculates the tree edit distance between the reconstructed table and the table to be identified as the accuracy of the reconstructed table when the reconstructed number of rows and columns is different from the standard number of rows and columns.
[0158] Optionally, the device further comprises:
[0159] After acquiring the table to be identified, the table splitting unit 807 determines the standard number of rows and columns of the table to be identified; if the standard number of rows and columns exceeds the row and column number threshold, the table to be identified is split into a number of sub-tables to be identified;
[0160] The table reconstruction unit 805 performs table reconstruction for each sub-table to be identified to obtain a corresponding reconstructed sub-table;
[0161] The table reconstruction unit 805 is further configured to merge the reconstructed sub-tables corresponding to the sub-tables to be identified to obtain a reconstructed table corresponding to the table to be identified.
[0162] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing processor executable instructions; wherein the processor implements the steps of the method described in any of the above embodiments by running the executable instructions.
[0163] Based on the same concept as the above method, this specification also provides a computer-readable storage medium on which computer instructions are stored. When the instructions are executed by a processor, the steps of the method described in any of the above embodiments are implemented.
[0164] Based on the same concept as the above method, this specification also provides a computer program product, including a computer program / instruction, which implements the steps of the method described in any of the above embodiments when executed by a processor.
Claims
1. A table recognition method, comprising: Get the table to be identified; Performing structural recognition on the table to be recognized to obtain a cell logical position and a cell physical position of each cell in the table to be recognized; Performing content recognition on the table to be recognized to obtain the text content and physical location of each text area in the table to be recognized; Determine the correspondence between the cell logical position and the text content according to the cell physical position and the text area physical position; The table to be identified is reconstructed according to the corresponding relationship to obtain a reconstructed table.
2. The method according to claim 1, before performing content recognition on the table to be recognized, the method further comprises: Determining the target format of the table to be identified; The process of performing content recognition on the table to be recognized includes: Searching for a target text recognition technology corresponding to the target format in the correspondence between formats and text recognition technologies; The target text recognition technology is used to perform content recognition on the table to be recognized.
3. The method according to claim 1, wherein the process of determining the correspondence between the cell logical position and the text content according to the cell physical position and the text area physical position comprises: For each cell, calculating an intersection-over-union ratio between the cell and each text area according to the cell physical position of the cell and the physical position of each text area; Determine the text area with the maximum intersection-over-union ratio calculated as the text area corresponding to the cell; A correspondence between the cell logical position of the cell and the text content of the text area corresponding to the cell is established.
4. The method according to claim 1, further comprising: Determining the number of reconstructed rows and columns of the reconstructed table; Comparing the reconstructed number of rows and columns with the standard number of rows and columns of the table to be identified; When the number of reconstructed rows and columns is the same as the number of standard rows and columns, calculating the content similarity between the reconstructed table and the table to be identified; The accuracy of the reconstructed table is determined according to the content similarity.
5. The method according to claim 4, wherein the process of calculating the content similarity between the reconstructed table and the table to be identified comprises: For each cell, comparing the first type of text content in the cell in the reconstructed table with the second type of text content in the same cell in the to-be-identified table; Determining the content similarity of the cell according to a comparison result between the first category of text content and the second category of text content; The content similarities of each cell are summarized to obtain the content similarities between the reconstructed table and the table to be identified.
6. The method according to claim 4, further comprising: In the case that the reconstructed number of rows and columns is different from the standard number of rows and columns, a tree edit distance between the reconstructed table and the table to be identified is calculated as the accuracy of the reconstructed table.
7. The method according to claim 1, further comprising: After obtaining the table to be identified, determining the standard number of rows and columns of the table to be identified; When the number of standard rows and columns exceeds the row and column number threshold, splitting the table to be identified into a plurality of sub-tables to be identified; Reconstruct each sub-table to be identified to obtain a corresponding reconstructed sub-table; The reconstructed sub-tables corresponding to the sub-tables to be identified are merged to obtain a reconstructed table corresponding to the table to be identified.
8. A table recognition device, comprising: A table acquisition unit acquires a table to be identified; A structure recognition unit performs structure recognition on the table to be recognized to obtain a cell logical position and a cell physical position of each cell in the table to be recognized; A content recognition unit performs content recognition on the table to be recognized to obtain the text content and physical location of each text area in the table to be recognized; a relationship determination unit, which determines a corresponding relationship between the logical position of the cell and the text content according to the physical position of the cell and the physical position of the text area; The table reconstruction unit reconstructs the table to be identified according to the corresponding relationship to obtain a reconstructed table.
9. The device according to claim 8, Before the content recognition unit performs content recognition on the table to be recognized, the target format of the table to be recognized is also determined; The process of the content recognition unit performing content recognition on the table to be recognized includes: Searching for a target text recognition technology corresponding to the target format in the correspondence between formats and text recognition technologies; The target text recognition technology is used to perform content recognition on the table to be recognized.
10. The apparatus according to claim 8, wherein the process of the relationship determination unit determining the corresponding relationship between the cell logical position and the text content according to the cell physical position and the text area physical position comprises: For each cell, calculating an intersection-over-union ratio between the cell and each text area according to the cell physical position of the cell and the physical position of each text area; Determine the text area with the maximum intersection-over-union ratio calculated as the text area corresponding to the cell; A correspondence between the cell logical position of the cell and the text content of the text area corresponding to the cell is established.
11. The apparatus according to claim 8, further comprising: The accuracy determination unit determines the number of reconstructed rows and columns of the reconstructed table; compares the number of reconstructed rows and columns with the standard number of rows and columns of the table to be identified; when the number of reconstructed rows and columns is the same as the standard number of rows and columns, calculates the content similarity between the reconstructed table and the table to be identified; and determines the accuracy of the reconstructed table based on the content similarity.
12. An electronic device comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method according to any one of claims 1 to 7 by executing the executable instructions.
13. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 7.
14. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.