Table Reconstruction Method, Device, Computer Equipment and Readable Storage Medium
By detecting and identifying the text box content and layout information in the table image, as well as the coordinates of the table lines, and combining row and sequence numbers for table reconstruction, the problem of inaccurate semi-wireless table reconstruction in the prior art is solved, and the accuracy of reconstruction is improved.
Patent Information
- Application Number
- CN202111417747.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-11-26
AI Technical Summary
The detection results of the prior art are inaccurate when reconstructing semi-wireless tables, resulting in the reconstructed tables that do not match the actual tables.
By obtaining the table image, detecting and identifying the content and layout information of the text box, as well as the coordinates of the row table line and list table line, and combining the row sequence number and column sequence number for table reconstruction.
Improve the accuracy of table reconstruction, especially when processing semi-wireless tables, the actual structure of the table can be accurately restored.
Smart Images

Figure CN114005126B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of table processing, and in particular, to a table reconstruction method, device, computer device and readable storage medium. Background Art
[0002] A table is a common document form. Commonly used editable table documents include Excel tables and tables inserted in Word. There are many types of tables, such as wired tables and semi-wired tables. Users often need to convert various types of text materials into editable table documents for storage.
[0003] Currently, before obtaining an editable table document, related technologies only use the method of detecting table lines to reconstruct the table. However, this method often has a good detection effect on wired tables, but for semi-wired tables (such as three-line tables), the detection result of this method is inaccurate, resulting in the reconstructed table not matching the actual table. Summary of the Invention
[0004] One of the purposes of the present invention is to provide a table reconstruction method, device, computer device and readable storage medium to solve the above technical problems. The embodiments of the present invention can be implemented as follows:
[0005] In a first aspect, the present invention provides a table reconstruction method, the method including: obtaining a table image; performing detection and recognition on the table image to obtain the text content and layout information corresponding to each of a plurality of text boxes, as well as the coordinates of row table lines and the coordinates of column table lines; wherein, the layout information includes position information, row number and column number; performing table reconstruction according to the text content, layout information, and the coordinates of row table lines and column table lines.
[0006] For the table reconstruction method provided by the above technical solution, due to the layout information corresponding to each text content determined in advance, therefore, when reconstructing a table, especially a semi-wired table, the actual structure of the table can be accurately restored, improving the accuracy of the reconstructed table.
[0007] In an optional embodiment, performing recognition on the table image to obtain the text content and layout information corresponding to each of a plurality of text boxes in the table, as well as the coordinates of row table lines and column table lines of the table, includes: performing text recognition on the table image to respectively obtain the position information and text content corresponding to each of the plurality of text boxes; determining the row number and column number corresponding to each of the plurality of text boxes according to the position information corresponding to each of the plurality of text boxes; performing straight line detection on the table image to obtain the coordinates of row table lines and column table lines.
[0008] Through the above technical solutions, the layout information of each text box, the coordinates of the row table lines, and the coordinates of the list table lines can be determined quickly and accurately, providing a reliable basis for subsequent table reconstruction and improving the accuracy of the reconstructed table.
[0009] In an alternative embodiment, according to the position information corresponding to each of the multiple text boxes, determining the row numbers and column numbers corresponding to each of the multiple text boxes includes: numbering the multiple text boxes along a first direction and a second direction respectively, and obtaining a first text box sequence and a second text box sequence in order of the numbers; wherein, the first direction and the second direction are perpendicular; inputting the position information of the multiple text boxes in the first text box sequence into a row prediction model to obtain a row label sequence corresponding to the multiple text boxes; wherein, the sequence order of the row label sequence is the same as the sequence order of the first text box sequence; inputting the position information of the multiple text boxes in the second text box sequence into a column prediction model to obtain a column label sequence corresponding to the multiple text boxes; the sequence order of the column label sequence is the same as the sequence order of the second text box sequence; respectively parsing the row label sequence and the column label sequence to determine the row numbers and column numbers corresponding to each of the multiple text boxes.
[0010] Through the row prediction model and the column prediction model in the above technical solutions, the row numbers and column numbers corresponding to each text box can be determined quickly and accurately.
[0011] In an alternative embodiment, respectively parsing the row label sequence and the column label sequence to determine the row numbers and column numbers corresponding to each of the multiple text boxes includes: respectively determining the row demarcation position and the column demarcation position from the row label sequence and the column label sequence; according to the row demarcation position, determining the row numbers corresponding to each of the multiple text boxes, and according to the column demarcation position, determining the column numbers corresponding to each of the multiple text boxes.
[0012] Through the above technical solutions, the row numbers and column numbers corresponding to each text box can be determined quickly and accurately, providing a reliable data basis for subsequent table reconstruction.
[0013] In an alternative embodiment, numbering the multiple text boxes along a first direction and a second direction respectively includes: according to the position information corresponding to each of the multiple text boxes, determining the center coordinates corresponding to each of the multiple text boxes; the center coordinates include the sub-coordinates in the first direction and the sub-coordinates in the second direction; sorting the multiple text boxes according to the magnitudes of the sub-coordinates in the first direction, and sequentially numbering the sorted multiple text boxes; sorting the multiple text boxes according to the magnitudes of the sub-coordinates in the second direction, and sequentially numbering the sorted multiple text boxes.
[0014] By numbering the text boxes in the above solution, the prediction results of each text box can be quickly located in the model output, improving the efficiency of subsequent table reconstruction.
[0015] In an alternative embodiment, line detection is performed on the table image to obtain the coordinates of the row table lines and the coordinates of the column table lines, including: performing semantic segmentation on the table image to obtain a first feature map and a second feature map, where the first feature map contains the row table lines and the second feature map contains the column table lines; respectively performing line detection on the first feature map and the second feature map to obtain the coordinates of the row table lines and the coordinates of the column table lines.
[0016] Through the above technical solution, the coordinates of the row table lines and the coordinates of the column table lines in the table can be detected quickly and accurately, improving the efficiency of table reconstruction and providing a reliable data basis for subsequent table reconstruction.
[0017] In an alternative embodiment, table reconstruction is performed according to the text content, layout information, and the coordinates of the row table lines and the coordinates of the column table lines, including: generating a text matrix according to the row numbers and column numbers, and writing the text content into the text matrix according to the row numbers and column numbers of the text boxes corresponding to the text content; determining the table positions corresponding to the border lines, row table lines, and column table lines of the table according to the position information of the text boxes existing in each row and each column of the text matrix, the coordinates of the row table lines of the table, and the coordinates of the column table lines; performing table reconstruction according to the table content, the border lines of the table, the row table lines, and the table positions corresponding to the column table lines.
[0018] Through the above technical solution, the positions of each row table line and each column table line in the table, as well as the border line information in the table, can be accurately determined. Combining this information with the pre-determined layout information, the table can be accurately reconstructed to make the reconstructed table conform to the actual table.
[0019] In an alternative embodiment, determining the table positions corresponding to the border lines, row table lines, and column table lines of the table according to the position information of the text boxes existing in each row and each column of the text matrix, the coordinates of the row table lines of the table, and the coordinates of the column table lines includes: determining the average center coordinates of each row and the average center coordinates of each column within the text matrix according to the position information of the text boxes existing in each row and each column of the text matrix; comparing the coordinates of the row table lines with the average center coordinates of each row and the average center coordinates of each column respectively to determine the row border lines of the table, the rows where each row table line is located, and the starting and ending positions of each row table line in the column direction; comparing the coordinates of the column table lines with the average center coordinates of each column and the average center coordinates of each row respectively to determine the column border lines of the table, the columns where each column table line is located, and the starting and ending positions of each column table line in the row direction.
[0020] Through the above technical solution, the positions of each row table line and each column table line in the table can be accurately determined, providing a data basis for subsequent table reconstruction, so that the reconstructed table conforms to the actual table and improves the efficiency of table reconstruction.
[0021] In an alternative embodiment, table reconstruction is performed according to the table content, the border lines of the table, and the respective corresponding table positions of the row table lines and the column table lines, including: generating an editable table according to the border lines of the table, the row table lines, and the respective corresponding table positions of the column table lines; writing the text content into the editable table in sequence to obtain the reconstructed table.
[0022] Through the above technical solution, an editable table document that conforms to the actual table can be obtained, facilitating the user's subsequent operations on the table.
[0023] In an alternative embodiment, obtaining a table image includes: obtaining an image to be recognized, where the image to be recognized contains a table; performing semantic segmentation on the image to be recognized to obtain a table region feature map; performing contour analysis on the table region feature map to determine the image region containing the table; and intercepting the image to be recognized according to the image region to obtain the table image.
[0024] Through the above technical solution, the area where the table is located can be quickly and accurately located, and then the table image can be obtained.
[0025] In an alternative embodiment, intercepting the image to be recognized according to the image region to obtain the table image includes: determining the coordinate information of the minimum bounding rectangle containing the table according to the image region; and intercepting the image to be recognized according to the coordinate information of the minimum bounding rectangle to obtain the table image.
[0026] Through the above technical solution, a complete table image can be quickly and accurately obtained.
[0027] In a second aspect, the present invention provides a table reconstruction device, including: an acquisition module for acquiring a table image; a recognition module for detecting and recognizing the table image to obtain the text content and layout information corresponding to each of a plurality of text boxes, as well as the coordinates of the row table lines and the coordinates of the column table lines; wherein the layout information includes position information, row numbers, and column numbers; and a reconstruction module for performing table reconstruction according to the text content, the layout information, and the coordinates of the row table lines and the coordinates of the column table lines.
[0028] In a third aspect, the present invention provides a computer device, including a processor and a memory, where the memory stores a computer program that can be executed by the processor, and the processor can execute the computer program to implement the table reconstruction method according to any one of the foregoing embodiments.
[0029] Fourthly, the present invention provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the table reconstruction method according to any one of the foregoing embodiments is implemented.
[0030] The table reconstruction method, device, computer device and readable storage medium provided by the embodiments of the present invention, the method includes: obtaining a table image; detecting and recognizing the table image to obtain the text content and layout information corresponding to each of a plurality of text boxes, as well as the coordinates of row table lines and the coordinates of column table lines; wherein, the layout information includes position information, row numbers and column numbers; performing table reconstruction according to the text content, layout information, and the coordinates of row table lines and column table lines. It can be seen that since the row numbers and column numbers corresponding to each text content are obtained, therefore, the table can be reconstructed by combining the row numbers, column numbers and the detected table lines. Compared with the prior art that only relies on the detected table lines to reconstruct the table, since this detection method is likely to misidentify the text content actually located in different rows in a semi-wireless table into the same row, resulting in the reconstructed table not conforming to the actual table. However, the table reconstruction method provided in this embodiment, due to the layout information corresponding to each text content determined in advance, therefore, when reconstructing the table, especially a semi-wireless table, the actual structure of the table can be accurately restored, improving the accuracy of table reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0032] Figure 1 It is an application environment diagram of a table reconstruction method provided by an embodiment of the present invention;
[0033] Figure 2 It is a schematic flowchart of a table reconstruction method provided by an embodiment of the present invention;
[0034] Figure 3 It is an image to be recognized provided by an embodiment of the present invention;
[0035] Figure 4 It is an example diagram of a table image provided by an embodiment of the present invention;
[0036] Figure 5 It is a table area feature diagram provided by an embodiment of the present invention;
[0037] Figure 6Schematic flowchart of the implementation manner of step S202 provided by an embodiment of the present invention;
[0038] Figure 7 Schematic diagram of the text recognition result of the table image provided by an embodiment of the present invention;
[0039] Figure 8 Schematic flowchart of the implementation manner of step S202-2 provided by an embodiment of the present invention;
[0040] Figure 9 Example diagram of sorting text boxes in the vertical and horizontal directions provided by an embodiment of the present invention;
[0041] Figure 10 Schematic diagram of a row prediction model provided by an embodiment of the present invention;
[0042] Figure 11 Schematic diagram of the prediction result of a row label provided by an embodiment of the present invention;
[0043] Figure 12 Schematic diagram of the prediction result of a column label provided by an embodiment of the present invention;
[0044] Figure 13 Schematic diagram of the first feature map provided by an embodiment of the present invention;
[0045] Figure 14 Schematic diagram of the second feature map provided by an embodiment of the present invention;
[0046] Figure 15 Schematic flowchart of the implementation manner of step S203 provided by an embodiment of the present invention;
[0047] Figure 16 Example diagram of a text matrix provided by an embodiment of the present invention;
[0048] Figure 17 Example diagram of a reconstructed table provided by an embodiment of the present invention;
[0049] Figure 18 Functional module diagram of the table reconstruction device provided by an embodiment of the present invention;
[0050] Figure 19 Schematic block diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manner
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0052] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, terms such as "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance.
[0053] A table is a common form of document. In daily work, there are many scanned or printed text materials such as expense lists, invoice receipts, and financial statements, which need to be converted into editable table documents and saved. Common editable table documents include Excel tables and tables inserted in Word.
[0054] Please refer to Figure 1 , Figure 1 which is an application environment diagram of a table reconstruction method provided by an embodiment of the present invention, including: a database 110, a terminal 120, a computer device 130, and a network 140.
[0055] The database 110 can be used to store text materials with table information in various forms, such as receipts, bills, insurance policies, notices, confirmation letters, application forms, etc. The formats of these text materials can include, but are not limited to, various types of pictures, screenshots, screen captures, scanned documents, PDF documents, etc., such as jpg, jpeg, ppm, bmp, png.
[0056] The terminal 120 can create or generate the above text materials in real time, and upload the text materials to the database for storage in real time, or upload the text materials to the computer device 130 for processing in real time.
[0057] The computer device 130 can be a device for processing the above text materials. Specifically, the computer device 130 can obtain the above text materials from the database 110, or the computer device 130 receives the above text materials uploaded by the terminal 120 in real time, and then executes the table reconstruction method provided by the embodiment of the present invention to achieve corresponding technical effects.
[0058] In some possible implementation manners, the above computer device 130 may be an independent physical server, or may be a server cluster or a distributed system composed of multiple physical servers. The above network 140 may include, but is not limited to: a wired network, a wireless network, where the wired network includes: a local area network, a metropolitan area network, and a wide area network, and the wireless network includes: Bluetooth, Wi-Fi, and other networks that implement wireless communication. The above terminal 120 may be, but is not limited to, a smart phone, a tablet computer, a personal computer (PC for short), a smart wearable device, and the like.
[0059] Please continue to refer to Figure 1 , for the above text materials with tabular information, in order to obtain an editable tabular document, the most common method is to process it through OCR recognition technology. Currently, most OCR recognition technologies can only recognize the text positions and their contents in the table, and cannot recognize the structural information of the table. If it is to be converted into an editable tabular document, manual intervention is required to manually restore the recognized text into an editable table, which will consume a large amount of human costs. Therefore, related technologies have proposed a method of detecting table lines to obtain an editable tabular document. However, this method often has a good detection effect on wired tables, but for semi-wireless tables (such as three-line tables), the detection results of this method are inaccurate, resulting in the reconstructed table not matching the actual table.
[0060] To solve the above technical problems, taking the application environment shown above Figure 1 as an example, the embodiments of the present invention provide a table reconstruction method. It can be understood that this table reconstruction method can be applied in a computer device 130 as shown in Figure 1 . Please refer to Figure 2 , Figure 2 is a schematic flowchart of the table reconstruction method provided by the embodiments of the present invention. The method may include the following steps:
[0061] S201, obtain a table image.
[0062] In one possible implementation manner, the above method for obtaining a table image may be: first obtain an image to be recognized, then identify the area where the table is located from the image to be recognized, and obtain the corresponding table image. For example, refer to Figure 3 , Figure 3 is an image to be recognized provided by the embodiments of the present invention. It can be seen that there is a table in the image to be recognized. By identifying the area where the table is located, and then intercepting this area, a table image can be obtained. The obtained table image is as shown in Figure 4 . Figure 4 is an example diagram of a table image provided by the embodiments of the present invention.
[0063] In another possible implementation, the method for obtaining the above table image may also be: obtaining the table image shown in Figure 4 by another electronic device with recognition ability, and then sending it to the computer device provided by the embodiment of the present invention, or directly inputting the pre-stored table image shown in Figure 4 by the user into the computer device.
[0064] S202. Detect and recognize the table image to obtain the text content and layout information corresponding to each text box, as well as the coordinates of the row table lines and the coordinates of the list table lines;
[0065] Among them, the layout information includes position information, row number, and column number. In this embodiment, the position information of the text box can be represented by the center coordinates corresponding to the text box, the row number and the column number can be determined one by one through the row prediction model and the column prediction model respectively, and the coordinates of the row table lines and the coordinates of the list table lines can be obtained by any existing straight line detection method.
[0066] S203. Perform table reconstruction according to the text content, layout information, and the coordinates of the row table lines and the coordinates of the list table lines.
[0067] In this embodiment, since the row number and column number corresponding to each text content are obtained, the table can be reconstructed by combining the row number, column number, and the detected table lines. Compared with the prior art that only relies on the detected table lines to reconstruct the table, since this detection method is likely to misidentify the text content actually located in different rows in a semi-wireless table into the same row, resulting in the reconstructed table not conforming to the actual table. However, the method for reconstructing the table provided in this embodiment can accurately restore the actual structure of the table due to the pre-determined layout information corresponding to each text content. Therefore, when reconstructing the table, especially a semi-wireless table, the accuracy of the reconstructed table is improved.
[0068] Optionally, as can be seen from the above embodiments, this embodiment can adopt various implementation manners to obtain the table image. Hereinafter, one implementation manner for obtaining the table image in this embodiment will be introduced in detail, that is, step S201 may include:
[0069] Step 1. Obtain the image to be recognized, where the image to be recognized contains a table.
[0070] It can be understood that the above image to be recognized may be the image to be recognized shown in Figure 3 or other forms of images with tables, which are not limited here.
[0071] Step 2. Perform semantic segmentation on the image to be recognized to obtain a table region feature map; the table region feature map contains a table.
[0072] In this embodiment, taking Figure 3 the to-be-recognized image shown as an example, the table area feature map obtained after semantic segmentation of the to-be-recognized image can be as Figure 5 shown, Figure 5 which is a table area feature map provided by an embodiment of the present invention. It can be seen that this table area feature map is a to-be-recognized image processed by binarization, where the white area in the image is the table area.
[0073] Step 3: Perform contour analysis on the table area feature map to determine the image area containing the table;
[0074] Step 4: Intercept the to-be-recognized image according to the image area to obtain a table image.
[0075] Specifically, in a possible implementation manner, the method of intercepting the table image from the to-be-recognized image may be: determining the coordinate information of the minimum bounding rectangle containing the table according to the image area; intercepting the to-be-recognized image according to the coordinate information of the minimum bounding rectangle to obtain a table image.
[0076] Optionally, for the above step S202, this embodiment also provides a possible implementation manner. Please refer to Figure 6 , Figure 6 which is a schematic flowchart of the implementation manner of step S202 provided by an embodiment of the present invention. Step S202 may include the following steps:
[0077] S202-1: Perform text recognition on the table image to respectively obtain the position information and text content corresponding to each text box;
[0078] In a possible implementation manner, taking Figure 4 the table image shown as an example, the table image can be input into a pre-trained text detection model to detect the text boxes in the table, and then the detected text lines can be input into a pre-trained text recognition model to recognize the text content corresponding to each text box. The text recognition efficiency and accuracy can be improved through the model, and the recognition result can be as Figure 7 shown, Figure 7 which is a schematic diagram of the text recognition result of the table image provided by an embodiment of the present invention.
[0079] In this embodiment, the position information corresponding to the above text box can be represented by the center coordinates of the text box. Specifically, obtaining the center coordinates of each text box can be calculated using the four vertex coordinates of the text box. The vertex coordinates can be sorted in clockwise or counterclockwise order, which is not limited herein. For example, when sorted in clockwise order, the vertex in the upper left corner can be used as the first vertex.
[0080] Assume: The coordinates of the four vertices are (x1 , y 1 , x 2 , y 2 , x 3 , y 3 , x 4 , y 4 ), then the center coordinates are denoted as (x c , y c ), where x c and y c are calculated as follows:
[0081]
[0082]
[0083] Taking Figure 4 and Figure 6 in "Age(years)" as an example, assuming that the four vertex coordinates corresponding to the text content are: (180, 5, 220, 5, 220, 20, 180, 20), then the center coordinates corresponding to "Age(years)" can be obtained through the above center coordinate calculation formula as: (x c = 200, y c = 12).
[0084] S202-2. Determine the row numbers and column numbers corresponding to multiple text boxes according to the position information corresponding to each of the multiple text boxes.
[0085] S202-3. Perform line detection on the table image to obtain the coordinates of the row table lines and the coordinates of the column table lines.
[0086] The above steps S202-2 and S202-3 will be introduced in detail below.
[0087] In a possible implementation manner, the above step S202-2 can be implemented by a pre-trained model, which can improve the recognition efficiency and accuracy. Therefore, the possible implementation manner of step S202-2 can be as Figure 8 shown Figure 8 is a schematic flowchart of the implementation manner of step S202-2 provided by an embodiment of the present invention. Step S202-2 may include:
[0088] S202-2-1. Number multiple text boxes along the first direction and the second direction respectively, and obtain a first text box sequence and a second text box sequence respectively according to the numbering order; wherein, the first direction and the second direction are perpendicular.
[0089] It can be understood that after obtaining the position information corresponding to each of the multiple text boxes, the text boxes can be sorted and numbered based on the magnitudes of the sub - coordinates in different directions of the position information. The purpose of numbering is to facilitate quickly locating and determining the text label corresponding to each text box from the output results of the row prediction model and the column prediction model subsequently.
[0090] In a possible implementation manner, the above - mentioned step S202 - 2 - 1 can be implemented as follows: According to the position information corresponding to each of the multiple text boxes, determine the center coordinates corresponding to each of the multiple text boxes; the center coordinates include the sub - coordinate in the first direction and the sub - coordinate in the second direction; sort the multiple text boxes according to the magnitude of the sub - coordinate in the first direction, and sequentially number the sorted multiple text boxes; sort the multiple text boxes according to the magnitude of the sub - coordinate in the second direction, and sequentially number the sorted multiple text boxes.
[0091] It should be noted that before numbering the text boxes above, the coordinate system established for the table image in this embodiment is: taking the upper - left corner of the table image as the origin, the direction from the upper - left corner to the upper - right corner along the ascending direction is the abscissa, and the direction from the upper - left corner to the lower - left corner along the ascending direction is the ordinate. Subsequently, the coordinates of the row table lines and the column table lines obtained are also calibrated in this coordinate system. Of course, the user can also establish the coordinate system in other ways. When establishing the coordinate system in other ways, the various position information mentioned in this embodiment can be adjusted accordingly.
[0092] In a possible implementation manner, the above - mentioned first direction can be the vertical direction perpendicular to the horizontal plane, and the second direction is the horizontal direction parallel to the horizontal plane. During the process of numbering the multiple text boxes, they can be sorted and numbered from top to bottom according to the magnitude of the vertical direction (y c ) in the center point coordinates, and in the second direction, they can be sorted and numbered from left to right according to the magnitude of the horizontal direction (x c ) in the center point coordinates. Therefore, taking the table image shown in Figure 7 as an example, the sorting and numbering results can be as shown in Figure 9 . Figure 9 FIG. is an example diagram provided by an embodiment of the present invention for sorting text boxes in the vertical direction and the horizontal direction. It can be seen that in different directions, the sorting results corresponding to each text box are different. Taking "Age(years)" as an example, the number corresponding to this text box in the vertical direction is 1, and the number corresponding to it in the horizontal direction is 13.
[0093] S202 - 2 - 2, input the position information of multiple text boxes in the first text box sequence into the row prediction model to obtain a sequence of row labels corresponding to the multiple text boxes; wherein, the sequence order of the sequence of row labels is the same as the sequence order of the first text box sequence.
[0094] For the convenience of understanding the first text box sequence and the second text box sequence, continue with Figure 9 as an example. First, in the vertical direction, there are 21 text boxes numbered from 1 to 21. Each text box is sequentially combined according to the number order to form a text box sequence, obtaining the first text box sequence. Similarly, in the horizontal direction, there are text boxes numbered from 1 to 21. It should be noted that the text boxes corresponding to numbers 1 to 27 in the horizontal direction are different from the text boxes corresponding to numbers 1 to 27 in the vertical direction. The second text box sequence is obtained in sequence according to the number order in the horizontal direction.
[0095] The row prediction model in this embodiment can be as Figure 10 shown. Figure 10 This is a schematic diagram of a row prediction model provided by an embodiment of the present invention. The input of the row prediction model is the position information corresponding to each text box in the first text box sequence, and the output result is a row label sequence. The sequence order of the row label sequence is the same as the sequence order of the first text box sequence, that is, the first row label corresponds to the text box numbered 1, and so on. Among them, the row label sequence contains two values, for example, distinguished by "S" and "O". Among them, "S" indicates that this text box is a row boundary point, and the text boxes after this text box are on the next row, and "O" indicates others.
[0096] For example, taking Figure 9 the first text box sequence in the vertical direction shown as an example, inputting the position information corresponding to each text box in the first text box sequence into the row prediction model, the obtained row label sequence can be seen in Figure 11 . Figure 11 This is a schematic diagram of the prediction result of a row label provided by an embodiment of the present invention. The finally obtained row label sequence can be (S, O, O, O, O, S, O, O, O, O, S, O, O, O, O, S, O, O, O, O, S). For the first "S", it corresponds to the text "Age(years" numbered 1, indicating that the texts corresponding to numbers 2 to 6 after number 1 are on the next row, which is consistent with the Figure 9 result in the vertical direction shown.
[0097] S202-2-3, input the position information of multiple text boxes in the second text box sequence into the column prediction model to obtain a column label sequence corresponding to the multiple text boxes; the sequence order of the column label sequence is the same as the sequence order of the second text box sequence.
[0098] In this embodiment, the method for obtaining the second text box sequence is similar to the method for obtaining the first text box sequence described above, which will not be elaborated here. Among them, the column prediction model has the same structure as the row prediction model. In a possible implementation, the column label sequence is similar to the row label sequence and contains two values, which are distinguished by "S" and "O" for example. Among them, "S" indicates that this text box is a column demarcation point, and the text boxes after this text box are in the next column, and "O" indicates others.
[0099] For example, Figure 9 taking the second text box sequence in the horizontal direction shown as an example, inputting the position information corresponding to each text box in the second text box sequence into the row prediction model, the obtained column label sequence can be seen in Figure 12 , Figure 12 which is a schematic diagram of the prediction result of a column label provided by an embodiment of the present invention. The finally obtained column label sequence can be (O, O, O, S, O, O, O, S, O, O, O, O, S, O, O, O, S, O, O, O, S). For the first "S", the corresponding text is "Age(years)" with number 4, indicating that the text from number 1 to number 4 is in the first column, and the text after number 4 is in the subsequent columns, which is consistent with the Figure 9 result in the horizontal direction shown.
[0100] S202-2-4. Respectively parse the row label sequence and the column label sequence to determine the row numbers and column numbers corresponding to each of the multiple text boxes.
[0101] As can be seen from the above content, the row label sequence and the column label sequence each contain the row demarcation position and the column demarcation position. Therefore, parsing can be performed based on the row demarcation position and the column demarcation position. For example, continuing with the above row label sequence (S, O, O, O, O, S, O, O, O, O, S, O, O, O, O, S, O, O, O, O, S) as an example, the finally parsed row numbers corresponding to each text box numbered 1 to 21 in the vertical direction are (0, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 4, 4, 4, 4, 4), where 0-4 respectively represent the 1st row to the 5th row. Similarly, parsing the column label sequence (O, O, O, S, O, O, O, S, O, O, O, O, S, O, O, O, S, O, O, O, S), the finally obtained column numbers corresponding to each text box numbered 1 to 21 in the horizontal direction are (0, 0, 0, 0, 1, 1, 1, 1, 2, 2, 2, 2, 2, 3, 3, 3, 3, 4, 4, 4, 4), where 0-4 respectively represent the 1st column to the 5th column, so that the row number and column number corresponding to each text box are obtained.
[0102] The following is an introduction to the above step S202-3. Among them, the above step S202-3 may include the following implementation process:
[0103] Step 1: Perform semantic segmentation on the table image to obtain a first feature map and a second feature map. Among them, the first feature map contains row table lines, and the second feature map contains list table lines.
[0104] In this embodiment, taking Figure 4 the shown table image as an example, the obtained first feature map can be as shown in Figure 13 . Figure 13 This is a schematic diagram of the first feature map provided by the embodiment of the present invention. Among them, the white straight line in the first feature map is the row table line. The second feature map can be as shown in Figure 14 . Figure 14 This is a schematic diagram of the second feature map provided by the embodiment of the present invention. Among them, the white straight line in the second feature map is the list table line.
[0105] Step 2: Respectively perform straight line detection on the first feature map and the second feature map to obtain the coordinates of the row table lines and the coordinates of the list table lines respectively.
[0106] In this embodiment, the method of straight line detection can be but is not limited to the Hough transform straight line detection method. By performing straight line detection on the first feature map and the second feature map, the endpoint coordinates of each straight line can be obtained. For example, in the first feature map shown in Figure 11 , the coordinate points of the first and second row straight lines are respectively: (x 1 = 5, y 1 = 5, x 2 = 295, y 2 = 5), (x 1 = 6, y 1 = 23, x 2 = 295, y 2 = 25); in the second feature map shown in Figure 12 , the coordinate points of the first and second column straight lines are respectively: (x 1 = 5, y 1 = 6, x 2 = 6, y 2 = 55), (x 1 = 5, y 1 = 6, x 2 = 48, y 2 = 54).
[0107] Optionally, after obtaining the position information of the text box, the line number, the column number, and the coordinates of the row table line and the list table line in the table line, all the above information can be combined to reconstruct the table and obtain an editable table document. Therefore, an implementation manner of reconstructing the table is also given below. Please refer to Figure 15 , Figure 15 which is a schematic flowchart of the implementation manner of step S203 provided by the embodiment of the present invention. Among them, step S203 may include the following steps:
[0108] S203-1, generate a text matrix according to the line number and the column number, and write the text content into the text matrix according to the line number and the column number of the text box corresponding to the text content.
[0109] In this embodiment, a text matrix can be generated according to the line number and the column number to Figure 9 The line numbers corresponding to each text box obtained as a result shown are (0, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 4, 4, 4, 4, 4), and the column numbers are (0, 0, 0, 0, 1, 1, 1, 1, 2, 2, 2, 2, 2, 3, 3, 3, 3, 4, 4, 4, 4). That is to say, a 5-row and 5-column text matrix can be generated, and then the text content is written into the positions in the text matrix one by one. Taking the text box "Age(years)" as an example, its corresponding line number and column number are "0" and "2" respectively, so it is in the 1st row and the 3rd column. By analogy, the positions of all text boxes in the text matrix can be determined. The obtained text matrix can be as Figure 16 shown. Figure 16 which is an example diagram of a text matrix provided by the embodiment of the present invention.
[0110] S203-2, determine the table positions corresponding to the border lines, row table lines, and list table lines of the table according to the position information of the text boxes existing in each row and each column in the text matrix, the coordinates of the row table line of the table, and the coordinates of the list table line.
[0111] In this embodiment, the border lines of the table may include an upper border line, a lower border line, a left border line, and a right border line. The table positions corresponding to the row table line and the list table line respectively represent the number of columns crossed by the row table line and the number of rows crossed by the list table line. It can also be understood as the starting position and the ending position of the row table line or the list table line. Here, the starting position and the ending position refer to a certain row or a certain column.
[0112] S203-3, perform table reconstruction according to the table content, the border lines of the table, and the table positions corresponding to the row table line and the list table line respectively.
[0113] In the above manner, it is possible to determine whether there are border lines in the table and the actual crossing range of each table line, so that the finally reconstructed table can better conform to the actual situation.
[0114] For the above step S203-2, an exemplary implementation manner is given in the embodiments of the present invention, that is, the above step S203-2 can be implemented as follows:
[0115] Step 1: Determine the average center coordinates of each row and each column in the text matrix according to the position information of the text boxes existing in each row and each column of the text matrix.
[0116] As Figure 16 shown, in the first row of the text matrix, there is a text box "Age(years", then the average center coordinate of the row is the center coordinate of this text box (x c =200, y c =12). In the last row, there are 4 text boxes. Assume that the obtained average center coordinate of the row is (x c =150, y c =50), and so on, to obtain the average center coordinate of each corresponding row; for each column, assume that the average value of the center point coordinates of the text boxes in the first column is (x c =35, y c =30), and the average value of the center point coordinates of the text boxes in the last column is (x c =282, y c =32).
[0117] Step 2: Compare the coordinates of the row table lines with the average center coordinates of each row and each column respectively to determine the row border lines of the table, the row where each row table line is located, and the starting and ending positions of each row table line in the column direction.
[0118] In this embodiment, before the comparison, the row table lines can be sorted according to the vertical direction y 1 coordinates. In this way, the minimum coordinate can be compared with the minimum average center coordinate of the row first, and the maximum coordinate can be compared with the maximum average center coordinate of the row, so that it can be quickly and accurately determined whether there are row border lines in the table, that is, the upper and lower border lines of the table, and then start comparing row by row according to the sorted row table lines.
[0119] For example, take the first row table line (x 1 =5, y 1 =5, x 2 =295, y 2 =5) and compare it with the average center coordinate of the first row (x c =200, y c= 12) for comparison. Since y 1 = y 2 <y c , it is determined that the first row table line is the upper side of the first row text box. That is to say, the first row straight line is the upper boundary line of the table. In this way, the first row table line does not need to be compared with the remaining row average center coordinates anymore. Then, the coordinates of the first row table line are compared with each column average center coordinate. Since the column average center coordinate of the first column (x c = 35, y c = 30) has x c >x 1 , and the column average center coordinate of the last column (x c = 282, y c = 32) has x c <x 2 , it is determined that the first row table line crosses all columns. By analogy, the judgment of all row straight lines is completed, and the row where each row table line is located and the columns it crosses are determined.
[0120] Step 3: Compare the coordinates of the list table lines with the column average center coordinates of each column and the row average center coordinates of each row respectively to determine the column boundary lines of the table, the column where each list table line is located, and the starting position and ending position of each list table line in the row direction.
[0121] In this embodiment, similar to the determination method of the row boundary line, the list table lines can also be sorted by coordinates in the horizontal direction x 1 . In this way, the minimum coordinate can be compared with the minimum column average center coordinate first, and the maximum coordinate can be compared with the maximum column average center coordinate, so as to quickly and accurately determine whether there are column boundary lines in the table, that is, the left and right boundary lines of the table, and then start comparing column by column according to the sorted list table lines.
[0122] Similar to the processing method of the above row table straight line, for example, take the first list table line (x 1 = 5, y 1 = 6, x 2 = 6, y 2 = 55) and compare it with the column average center coordinate of the first column (x c = 35, y c = 30). Since x 1 <x 2 <x c , it is determined that the first list table line is on the left side of the first column text box. That is to say, the first list table line is the left boundary line of the table; then, the first list table line is compared with the row average center coordinates of each row. Among them, for the row average center point coordinate of the first row (x c= 200, y c = 12), since y 1 < y c , for the row average center point coordinates (x c = 150, y c = 50) of the last line, there exists y 2 > y c , then it can be determined that the first list table line passes through all rows. By analogy, the judgment of all column lines is completed, and the column where each list table line is located and the rows it passes through are determined.
[0123] Through the above method, the boundary lines of the table can be determined. Of course, it can also be determined whether there are boundary lines in the table, and it can also be determined which row each row table line is located in and which rows it passes through. The same applies to the list table lines. Furthermore, based on the above information, table reconstruction can be carried out.
[0124] Optionally, in the embodiments of the present invention, the starting position and ending position of each row table line and list table line in the table can be obtained, as well as which row the row table line is located in and which column each list table line is located in. After obtaining the above information, the method of reconstructing the table can be: generating an editable table according to the table positions corresponding to the boundary lines, row table lines, and list table lines of the table; writing the text content into the editable table in sequence to obtain the reconstructed table.
[0125] In this embodiment, taking the Figure 4 shown table image as an example, the finally reconstructed table can be as shown in Figure 17 . Figure 17 is an example diagram of table reconstruction provided by the embodiments of the present invention. Combining Figure 16 and Figure 17 it can be seen that although in Figure 16 , the texts in the second row and the third row are two different rows, but in the process of determining the positions of the row table line and the list table line in the table as described above, it can be recognized that there is a line break situation between these two rows of text, but there is no row table line. This situation can be recognized through the above comparison method, so that the finally reconstructed table is consistent with the actual Figure 4 table, improving the accuracy of table reconstruction.
[0126] In order to implement the above steps in the embodiments to achieve the corresponding technical effects, the table reconstruction method provided by the embodiments of the present invention can be executed in a hardware device or in the form of software modules. When the table reconstruction method is implemented in the form of software modules, the embodiments of the present invention also provide a table reconstruction device. Please refer to Figure 18 . Figure 18 is the functional module diagram of the table reconstruction device provided by the embodiments of the present invention. The table reconstruction device 300 may include:
[0127] An acquisition module 310, configured to acquire a table image;
[0128] An identification module 320, configured to detect and identify the table image, obtain the text content and layout information corresponding to each of a plurality of text boxes, and the coordinates of row table lines and the coordinates of column table lines; wherein, the layout information includes position information, row numbers, and column numbers;
[0129] A reconstruction module 330, configured to perform table reconstruction according to the text content, layout information, and the coordinates of row table lines and column table lines.
[0130] In an optional embodiment, the identification module 320 is specifically configured to: perform text recognition on the table image to respectively obtain the position information and text content corresponding to each of a plurality of text boxes; determine the row numbers and column numbers corresponding to each of the plurality of text boxes according to the position information corresponding to each of the plurality of text boxes; perform straight line detection on the table image to obtain the coordinates of row table lines and the coordinates of column table lines.
[0131] In an optional embodiment, the identification module 320 is further specifically configured to number the plurality of text boxes respectively along a first direction and a second direction, and obtain a first text box sequence and a second text box sequence in order of the numbers; wherein, the first direction and the second direction are perpendicular; input the position information of the plurality of text boxes in the first text box sequence into a row prediction model to obtain a row label sequence corresponding to the plurality of text boxes; wherein, the sequence order of the row label sequence is the same as the sequence order of the first text box sequence; input the position information of the plurality of text boxes in the second text box sequence into a column prediction model to obtain a column label sequence corresponding to the plurality of text boxes; the sequence order of the column label sequence is the same as the sequence order of the second text box sequence; respectively parse the row label sequence and the column label sequence to determine the row numbers and column numbers corresponding to each of the plurality of text boxes.
[0132] In an optional embodiment, the identification module 320 is further specifically configured to respectively determine a row demarcation position and a column demarcation position from the row label sequence and the column label sequence; determine the row numbers corresponding to each of the plurality of text boxes according to the row demarcation position, and determine the column numbers corresponding to each of the plurality of text boxes according to the column demarcation position.
[0133] In an optional embodiment, the identification module 320 is further specifically configured to determine the center coordinates corresponding to each of the plurality of text boxes according to the position information corresponding to each of the plurality of text boxes; the center coordinates include sub-coordinates in the first direction and sub-coordinates in the second direction; sort the plurality of text boxes according to the magnitudes of the sub-coordinates in the first direction, and number the sorted plurality of text boxes in sequence; sort the plurality of text boxes according to the magnitudes of the sub-coordinates in the second direction, and number the sorted plurality of text boxes in sequence.
[0134] In an alternative embodiment, the recognition module 320 is further specifically configured to perform semantic segmentation on the table image to obtain a first feature map and a second feature map, where the first feature map contains row table lines and the second feature map contains column table lines; perform straight line detection on the first feature map and the second feature map respectively to obtain the coordinates of the row table lines and the coordinates of the column table lines respectively.
[0135] In an alternative embodiment, the reconstruction module 330 is specifically configured to: generate a text matrix according to the row numbers and column numbers, and write the text content into the text matrix according to the row numbers and column numbers of the text boxes corresponding to the text content; determine the table positions corresponding to the border lines, row table lines and column table lines of the table according to the position information of the text boxes existing in each row and each column of the text matrix, the coordinates of the row table lines of the table and the coordinates of the column table lines; perform table reconstruction according to the table content, the border lines of the table, the row table lines and the column table lines and their corresponding table positions.
[0136] In an alternative embodiment, the reconstruction module 330 is specifically configured to: determine the average center coordinates of each row and the average center coordinates of each column in the text matrix according to the position information of the text boxes existing in each row and each column of the text matrix; compare the coordinates of the row table lines with the average center coordinates of each row and the average center coordinates of each column respectively to determine the row border lines of the table, the rows where each row table line is located, and the start and end positions of each row table line in the column direction; compare the coordinates of the column table lines with the average center coordinates of each column and the average center coordinates of each row respectively to determine the column border lines of the table, the columns where each column table line is located, and the start and end positions of each column table line in the row direction.
[0137] In an alternative embodiment, the reconstruction module 330 is specifically configured to generate an editable table according to the table positions corresponding to the border lines, row table lines and column table lines of the table; write the text content into the editable table in sequence to obtain the reconstructed table.
[0138] In an alternative embodiment, the acquisition module 310 is specifically configured to: acquire a to-be-recognized image, where the to-be-recognized image contains a table; perform semantic segmentation on the to-be-recognized image to obtain a table region feature map; perform contour analysis on the table region feature map to determine the image region containing the table; intercept the to-be-recognized image according to the image region to obtain the table image.
[0139] In an alternative embodiment, the acquisition module 310 is specifically configured to determine the coordinate information of the minimum bounding rectangle containing the table according to the image region; intercept the to-be-recognized image according to the coordinate information of the minimum bounding rectangle to obtain the table image.
[0140] It should be noted that each functional module in the table reconstruction device 300 provided in the embodiment of the present invention can be stored in the memory in the form of software or firmware or fixed in the operating system (OS) of the computer device, and can be executed by the processor in the computer device. At the same time, the data and program codes required to execute the above modules can be stored in the memory. Therefore, the embodiment of the present invention also provides a computer device, which can be Figure 1 The computer device 130 shown, or other computer devices with data processing functions, is not limited in the present invention.
[0141] like Figure 19 , Figure 19 A block diagram of a computer device provided in an embodiment of the present invention. The computer device 130 includes a communication interface 131, a processor 132 and a memory 133. The processor 132, the memory 133 and the communication interface 131 are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The memory 133 can be used to store software programs and modules, such as program instructions / modules corresponding to the table reconstruction method provided in an embodiment of the present invention. The processor 132 executes various functional applications and data processing by executing the software programs and modules stored in the memory 133. The communication interface 131 can be used to communicate signaling or data with other node devices. In the present invention, the computer device 130 can have multiple communication interfaces 131.
[0142] Among them, the memory 133 can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), etc.
[0143] The processor 132 may be an integrated circuit chip with signal processing capabilities. The processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processing (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0144] An embodiment of the present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the table reconstruction method according to any one of the foregoing embodiments. The computer-readable storage medium may be, but is not limited to, various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a PROM, an EPROM, an EEPROM, a magnetic disk, or an optical disc.
[0145] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for table reconstruction, characterized in that, the method comprises: obtaining a table image; detecting and recognizing the table image to obtain the text content and layout information corresponding to each of a plurality of text boxes, as well as the coordinates of row table lines and the coordinates of column table lines; wherein, the layout information includes position information, row numbers and column numbers; performing table reconstruction according to the text content, the layout information, and the coordinates of the row table lines and the coordinates of the column table lines; detecting and recognizing the table image to obtain the text content and layout information corresponding to each of a plurality of text boxes, as well as the coordinates of row table lines and the coordinates of column table lines, includes: performing text recognition on the table image to respectively obtain the position information and the text content corresponding to each of the plurality of text boxes; determining the row numbers and column numbers corresponding to each of the plurality of text boxes according to the position information corresponding to each of the plurality of text boxes; performing line detection on the table image to obtain the coordinates of the row table lines and the coordinates of the column table lines; determining the row numbers and column numbers corresponding to each of the plurality of text boxes according to the position information corresponding to each of the plurality of text boxes, includes: numbering the plurality of text boxes respectively along a first direction and a second direction, and sequentially obtaining a first text box sequence and a second text box sequence according to the numbering order; wherein, the first direction and the second direction are perpendicular; inputting the position information of the plurality of text boxes in the first text box sequence into a row prediction model to obtain a row label sequence corresponding to the plurality of text boxes; wherein, the sequence order of the row label sequence is consistent with the sequence order of the first text box sequence; inputting the position information of the plurality of text boxes in the second text box sequence into a column prediction model to obtain a column label sequence corresponding to the plurality of text boxes; the sequence order of the column label sequence is consistent with the sequence order of the second text box sequence; respectively parsing the row label sequence and the column label sequence to determine the row numbers and column numbers corresponding to each of the plurality of text boxes.
2. The table reconstruction method according to claim 1, characterized in that, respectively parsing the row label sequence and the column label sequence to determine the row numbers and column numbers corresponding to each of the plurality of text boxes, includes: respectively determining a row demarcation position and a column demarcation position from the row label sequence and the column label sequence; determining the row numbers corresponding to each of the plurality of text boxes according to the row demarcation position, and determining the column numbers corresponding to each of the plurality of text boxes according to the column demarcation position.
3. The table reconstruction method according to claim 1, characterized in that, numbering the plurality of text boxes respectively along a first direction and a second direction, includes: determining the center coordinates corresponding to each of the plurality of text boxes according to the position information corresponding to each of the plurality of text boxes; the center coordinates include sub-coordinates in the first direction and sub-coordinates in the second direction; sorting the plurality of text boxes according to the magnitudes of the sub-coordinates in the first direction, and sequentially numbering the sorted plurality of text boxes; Sort the multiple text boxes according to the magnitudes of the sub - coordinates in the second direction, and sequentially number the sorted multiple text boxes.
4. The table reconstruction method according to claim 1, wherein, performing line detection on the table image to obtain the coordinates of the row table lines and the coordinates of the column table lines, including: performing semantic segmentation on the table image to obtain a first feature map and a second feature map, wherein the first feature map contains row table lines and the second feature map contains column table lines; performing line detection on the first feature map and the second feature map respectively to obtain the coordinates of the row table lines and the coordinates of the column table lines respectively.
5. The table reconstruction method according to claim 1, wherein, performing table reconstruction according to the text content, the layout information, and the coordinates of the row table lines and the column table lines, including: generating a text matrix according to the row numbers and the column numbers, and writing the text content into the text matrix according to the row numbers and column numbers of the text boxes corresponding to the text content; determining the table positions corresponding to the boundary lines, row table lines, and column table lines of the table according to the position information of the text boxes existing in each row and each column in the text matrix, the coordinates of the row table lines of the table, and the coordinates of the column table lines; performing table reconstruction according to the text content, the boundary lines of the table, and the table positions corresponding to the row table lines and the column table lines respectively.
6. The table reconstruction method according to claim 5, wherein, determining the table positions corresponding to the boundary lines, row table lines, and column table lines of the table according to the position information of the text boxes existing in each row and each column in the text matrix, the coordinates of the row table lines of the table, and the coordinates of the column table lines, including: determining the average row center coordinates of each row and the average column center coordinates of each column within the text matrix according to the position information of the text boxes existing in each row and each column in the text matrix; comparing the coordinates of the row table lines with the average row center coordinates of each row and the average column center coordinates of each column respectively to determine the row boundary lines of the table, the rows where each row table line is located, and the starting and ending positions of each row table line in the column direction; comparing the coordinates of the column table lines with the average column center coordinates of each column and the average row center coordinates of each row respectively to determine the column boundary lines of the table, the columns where each column table line is located, and the starting and ending positions of each column table line in the row direction.
7. The table reconstruction method according to claim 6, wherein, performing table reconstruction according to the text content, the boundary lines of the table, and the table positions corresponding to the row table lines and the column table lines respectively, including: generating an editable table according to the table positions corresponding to the boundary lines, row table lines, and column table lines of the table; writing the text content into the editable table in sequence to obtain the reconstructed table.
8. The table reconstruction method according to claim 1, wherein, Obtain a table image, including: Obtain an image to be recognized, wherein the image to be recognized contains the table; Perform semantic segmentation on the image to be recognized to obtain a table region feature map; Perform contour analysis on the table region feature map to determine the image region containing the table; Intercept the image to be recognized according to the image region to obtain the table image.
9. The table reconstruction method according to claim 8, characterized in that Intercepting the image to be recognized according to the image region to obtain the table image includes: Determine the coordinate information of the minimum bounding rectangle containing the table according to the image region; Intercept the image to be recognized according to the coordinate information of the minimum bounding rectangle to obtain the table image.
10. A table reconstruction device, characterized in that includes: An acquisition module for acquiring a table image; An identification module for detecting and identifying the table image to obtain the text content and layout information corresponding to each of the multiple text boxes, as well as the coordinates of the row table lines and the coordinates of the column table lines; wherein, the layout information includes position information, row numbers and column numbers; A reconstruction module for performing table reconstruction according to the text content, the layout information, the coordinates of the row table lines and the coordinates of the column table lines; The identification module is specifically configured to perform text recognition on the table image to respectively obtain the position information and the text content corresponding to each of the multiple text boxes; determine the row numbers and column numbers corresponding to each of the multiple text boxes according to the position information corresponding to each of the multiple text boxes; perform line detection on the table image to obtain the coordinates of the row table lines and the coordinates of the column table lines; The identification module is further specifically configured to number the multiple text boxes respectively in a first direction and a second direction, and obtain a first text box sequence and a second text box sequence respectively in the order of the numbers; wherein, the first direction and the second direction are perpendicular; input the position information of the multiple text boxes in the first text box sequence into a row prediction model to obtain a row label sequence corresponding to the multiple text boxes; wherein, the sequence order of the row label sequence is consistent with the sequence order of the first text box sequence; input the position information of the multiple text boxes in the second text box sequence into a column prediction model to obtain a column label sequence corresponding to the multiple text boxes; the sequence order of the column label sequence is consistent with the sequence order of the second text box sequence; respectively parse the row label sequence and the column label sequence to determine the row numbers and column numbers corresponding to each of the multiple text boxes.
11. A computer device, characterized in that includes a processor and a memory, the memory stores a computer program that can be executed by the processor, and the processor can execute the computer program to implement the table reconstruction method according to any one of claims 1-9.
12. A readable storage medium, on which a computer program is stored, characterized in that When the computer program is executed by a processor, it implements the table reconstruction method according to any one of claims 1-9.
Citation Information
Patent Citations
Image processing method and device
CN110837796A
Method and device for extracting and reconstructing table in receipt image
CN111079756A